Showing posts with label architecture. Show all posts
Showing posts with label architecture. Show all posts

Saturday, June 29, 2019

You're doing it wrong

Attention software developers: you're doing it wrong.

To be more specific, you need to be doing these things, and you're not.

Don't block the UI thread!


No matter what the application is, not matter what you're doing, do not block the UI! A "Please Wait" dialog or screen is trivial to code and makes the user experience so much better. You can never assume that an operation will always execute quick enough not to block. It always will under some scenario.

Catch everything!


There is no excuse (at least in the .NET world) to allow an exception to bubble up so far that the application crashes. There is always a way to catch exceptions at the highest level. Even if there's no way to recover, exit gracefully.

And don't restart the app without warning. Visual Studio and SQL Server Management Studio are big offenders here. If recycling is the only way to recover, fine, but don't just do it without telling me!


Always format numbers.


I'm so sick of having to squint at 41095809 and mentally insert the commas (or whatever the delimiter is in your culture). Having the code format this kind of output is so trivial. Software is supposed to make our lives easier. If you don't do this you have failed.

Don't confuse a pattern with a function.


Okay this one's not quite so obvious. There are times when you have to do a similar thing in lots of places. The natural inclination is to make this into a function/method/class. This is a good instinct! Don't repeat yourself! But it can be taken too far. Some things look similar but have enough differences that trying to create a generic base for them just makes the code way more complicated.

To put it another way, you wouldn't try to write one ForEach() function that handles all loop cases. It would quickly become a monster. Looping is a pattern, not a function.


Let's be clear: I have committed all these sins and will again. Learn from my fail, and we can all stop doing it wrong.

Thursday, August 16, 2018

Stop describing me!

Anyone who's ever worked in software knows how difficult it is to keep the code and the documentation in agreement. It's all too common to find errors in documentation, and not just typos, but misleading or even just false statements. This is (generally) not the result of malice, laziness, or stupidity, but is rather a natural consequence of trying to keep two unrelated things in sync with each other. Programming and natural languages are just not similar enough to make this a simple or automatic task. It's one reason experienced programmers are prone to say that all code comments are lies.

Additionally, documentation is far too often a crutch used to make up for the fact that the design is not intuitive. As Chip Camden once said, "Documentation helps, but documentation is not the solution - simplicity is." That doesn't mean that software can't be doing complex things - doing complex things is the whole point! What it means is that the interfaces to software, whether they be APIs or GUIs, have to make the complex thing simple.

Documentation then is typically a symptom of poor design and a difficult symptom to manage at that. But how then do we communicate the what and why of a software system?

The answer is self-documenting systems. "Self-documenting system" sounds like the tagline of some new tool somebody's trying to sell, but it's actually a design philosophy. It's an approach that says: make the endpoints, buttons, menu options, links, namespaces, class, method, parameter, and property names so obvious as to their purpose & function that, for the majority of them, documentation would be redundant.

The term "redundant" is used very deliberately here - it's not that documentation is never warranted, it's that we should seek to make it unnecessary, for the reasons mentioned above. There will still need to be a few bits of explanation here and there, but they'll be few, so maintaining them won't be a big burden, and they'll be very targeted, so it'll be very obvious when they need to be updated. (By targeted, I mean they'll be either high-level descriptions of the unit as a whole, or notes on some edge cases.)

Self-documenting systems lead us to the Principle of Minimum Necessary Documentation, which states:

There is an inverse relationship between how well designed something is and how much documentation it needs.

The term "inverse" is also used very deliberately. The inverse of infinity is not zero - it's darn near close to it, but it's not actually nothing. A perfect system needs no documentation, but there are no perfect systems, so there is no system that can have zero documentation, no matter how well designed it is. But for a well-designed system, a README file or some tooltips will probably suffice.

But do not assume from this that a project with no documentation except a README is well-designed, or that you can get away with just throwing together a README and taking no other thought for what's going on. Note the principle says "how much documentation it needs" not how much it has. If users have cause to say "the documentation has been skimped on", then the design is poor and no crutch was provided, which is even worse than a bad design band-aided with documentation. Intuitive, self-documenting design is the critical factor, not the documentation itself.

Now, there are a few caveats & clarifications that should be mentioned:
  • 'Self-documenting' is often used in connection with test-driven development. There is a vein of thought that says the tests should be the mechanism for documenting the code. This is a very good approach to take, but tests cannot serve as the user documentation. If the project is a public one, expecting people to read the tests is not realistic, and even with an internal project, understanding the guts of how a component works is often beyond the scope of the task at hand. Tests can and should be the developer documentation, and complete test coverage is critically important to developing robust software, but a test suite is typically too large and specialized to be the sole documentation.
  • Most systems have a fairly obvious surface. It may not be instantly obvious what a class structure or API or SDK or GUI does, but it's usually pretty easy to see what it has. There are interfaces though that don't have an easily reflectable surface, such as command-line tools. At first blush, it might seem that this is a place where more documentation is a necessary evil. Certainly there are a lot of tools that believe extensive documentation is preferable to abstraction (git, I'm looking at you). But the principle does still apply here. It is still incumbent on the developer to make the behavior intuitive (typically by following conventions), and to make the options easily discoverable. There are a myriad of ways to approach this, and it depends a lot on how flexible the tool needs to be. It requires more work, but that's our job: to make it easy to use.
It's as I've said before: deal with things directly. Don't build hedges around them, because it just takes more work and isn't as effective. Don't worry about documentation unless there truly is no other way to get the point across. Focus instead of making your code as intuitive and straightforward as possible, as this is the greatest programming good.

Tuesday, June 28, 2016

The Case for Dealing Directly

Update 8/1/2016: After giving this article some further thought, I added another reason.

Suppose you started a new job, and sat down to look over the codebase. As you dug through things, you found that, instead of using public methods and/or properties on classes, they had written a whole reflection framework to reach into objects and modify the private fields. Then, when you asked why they had done something so very peculiar, they responded by saying, "Oh, this way is much better - we don't have to care about the structure of the objects, or waste time writing accessor methods and properties, and we can spend all our time just writing actual code!"

What? No. What?

We could probably spend a whole series of articles explaining why this is a bad approach, but let's just focus on these reasons:
  1. it works against the language's paradigm
  2. it leads to contrivances that don't use the objects very effectively
  3. it makes figuring out the usage and dependencies of objects very difficult
  4. it has security risks
  5. it is more risky to change
  6. it doesn't scale, and the performance could get pretty bad without a lot of ways to optimize it
This is what dynamic SQL is like, and my reaction to it is the same.

Current or former colleagues of mine might quickly note that I have written my fair share of dynamic SQL, and astute readers might note that I have published dynamic SQL on this very blog before. Dynamic SQL, like reflection, is not a bad thing, and has good uses. But it's generally not good for line-of-business applications, for all the reasons listed above.

Dynamic SQL is code generation. When code generation is employed at design (develop) time, it is a very good thing! Intellisense is the most common form of code generation, and it's so fluent & ubiquitous that we barely even think about it anymore. One of the reasons I sing the praises of ReSharper to anyone and everyone who'll listen (and I'm starting to get there with SQL Prompt) is because of the code generation features. They just make life easier, while still giving you the flexibility to write exactly what you need, because you don't have to accept everything they give you.

But code generation at run-time is a different story. Would you trust an application that generates, compiles, and executes code on the fly? That kind of sounds like the behavior of a virus, or at least the work of someone who doesn't yet realize that they'll have to debug that thing somehow. Sure, it's clever, but it's fraught with security, maintainability, reliability, and performance pitfalls.

There are pieces of code that we have to write that are trivial and repetitive, and that's what code generation is for. But the whole reason programmers haven't all been replaced by robots is that writing good code is not a deterministic activity. It requires a fair amount of specialized knowledge. SQL is no different.

Let's consider each one of these points a little more closely.


It works against the paradigm


In other words, just because we can doesn't mean we should. 

To an application, a database is an external resource, like a file system or an API. Interactions to external resources generally need to be coarse-grained, and resources need to be well-encapsulated. This means that while tables are the basic unit of database storage, they are not the basic unit of database interactions. In an ideal application, any discrete action will incur only one database call - e.g., there's only one call when the page loads, and only one when the user clicks the button to update the page. Relational databases are built very unambiguously on the principle that interactions should be atomic so they can be consistent and isolated.

Dynamic SQL works against this paradigm because it wants to deal with the tables as individual units instead of dealing with endpoints (stored procedures) performing atomic operations. The workaround for this is typically to create a transaction in the application code, which often is a performance anti-pattern - it makes the database layer chatty, and holds locks open longer, which leads to blocking & degraded performance. (There are places where in-code transactions make sense, but they are the exception, not the rule.)

It's true that databases are often under-abstracted and have too much repetition, but the solution isn't to go around the database, it's to employ the mechanisms it does have.


It leads to contrivances


I've never met an ORM framework I really liked (and it turns out there are a lot of other engineers who feel the same way; Google it to find out more). You're constantly having to work around them to get things done properly, and the SQL they generate is almost always a classic example of what not to do in a relational database.

Writing effective and appropriate UPDATE, JOIN and WHERE clauses is not a trivial task. Structuring them in the correct manner requires addressing multiple considerations - the impact to/of indexes, the size of the tables involved, the native change tracking features, replication, etc. Dynamic SQL code generators do not provide any abstraction in this regard: they have to get very detailed in order to provide the level of control needed. In other words, they have to re-create the wheel, and in the end their output is no more optimal, and typically less optimal, than writing the SQL directly in stored procedures.


It makes usage discovery difficult


When dynamic SQL generators build statements, the table, view, column, etc. names are coming from diverse locations, and how they are used is often obfuscated. This makes it hard to know which database objects are actually in use. As with any type of code, when it's difficult to tell what's in use and how it's being used, it's difficult to move forward. It significantly hampers development and debugging efforts. Every database I've ever encountered has had many things in it that everyone is 'pretty sure aren't even used anymore', but because discovering dependencies is difficult even under the best of circumstances, nothing is ever done about it. Dynamic SQL's obfuscation of usage just exacerbates this problem, while direct SQL in stored procedures is much more clear and consolidated.


It has security risks


It is often pointed out that the security concerns of dynamic SQL are mitigated by parameterization. This is true, but it's also true that stored procedures mitigate these risks even further. When the SQL code is being constructed on the fly, it still opens up possibilities for injection (and compile errors) that simply don't exist with stored procedures. (Naturally, stored procedures that build dynamic SQL suffer from these same flaws, and in some ways more so - I include them in the ranks of 'dynamic SQL code generators.' Traitors.)

Stored procedures are also a better way to implement the principle of least privilege. When dynamic and/or inline SQL is involved, the application must be given carte blanche to act on the tables. But with stored procedures, the interactions are much more targeted, and (typically) only one object has to have permissions associated to it. It's more maintainable and granular access, two key elements of good security.

These may seem like minor concerns. They are not. Quite frankly, security in every single piece of software on this planet needs to be better - and not just better, but many orders of magnitude better. Good enough security is not good enough.

It is riskier to change


When every operation is being managed by the same code, the risk of changing that code becomes much greater. If a change is made to a dynamic SQL builder for one scenario, it runs the risk of breaking all the other scenarios - in other words, it is harder to keep changes isolated. Say for example you have code that is generating INSERT statements, and you make a change to accommodate one edge case. In doing so, you run the risk of breaking all the INSERT statements in the application. But with stored procedures, a change to one statement has zero effect on similar statements in other procedures. Changes can be kept isolated, the system is more flexible and we can be more confident in our changes.


It doesn't scale


Typically, the only way to optimize dynamic SQL is to pull it out and make it not dynamic. Yes, indexes are always a part of the optimization equation in a relational database, but they're just one tool, and quite frankly they're usually just the first layer. To really address these matters we have to get into the queries themselves, and when we can't directly control the queries, we're at a serious disadvantage.

Also, compilation is not free. Dynamic/inline SQL execution plans can be cached, and stored procedure execution plans aren't necessarily cached, but the former typically aren't and the latter typically are. It's a significant advantage for the DBAs or the DevOps to be able to look up an execution plan by the procedure name, an impossible task with dynamic and inline SQL. Reducing the number of compilations is also a important step in preventing CPU bottlenecks in SQL Server, something that can only be achieved with stored procedures.


(It is not always possible to use stored procedures - after all, some applications have to interact with databases controlled by third parties. But in these scenarios, inline SQL is preferable to dynamic SQL, because it mitigates or avoids most of these discussed pitfalls of dynamic SQL even though it lacks some of the advantages of stored procedures.)

In the end, this can all be boiled down to one simple idea: deal with the database on its own terms. In that reflection framework example in the first paragraph, the coders are effectively trying to treat their class structure as a set of global variables, or a key-value store. Trying to make something act like something it is not only leads to trouble. Relational databases are fundamentally different from imperative object code, and attempts to treat them like they're not are misguided.

Tuesday, February 3, 2015

Learn the Rules so you can figure out who's breaking them

99.99 percent [of subatomic interactions] are explainable ... Spending your time exploring each particle trail will lead you to conclude that all the particles obey known physics, and there's nothing left to discover ... Most stars are boring; advances comes from studying the weirdies - the quasars, the pulsars, the gravitational lenses - that don't seem to fit into the models that you've grown up with ... Collect raw data and throw away the expected. What remains challenges your theories.
- Clifford Stoll, in The Cuckoo's Egg
The Cuckoo's Egg is an amazing book - it's a memoir, a techno-thriller, and a manifesto all in one. It recounts how, in the mid-1980s, an astronomer-turned-sysadmin at UC Berkley stumbled across a hacker in their system, and upon tracking down the intruder, ended up catching a ring of KGB agents in their first real stabs at cyber-espionage! It's a must-read for anyone who's ever going to write code, and just a darn good read for everyone else.

The quote above (from chapter 3) is a very insightful one because it tells us, in essence, that the really useful details are in the edge cases. Stoll's first hint that something was up was a $0.75 discrepancy in an accounting system. No one suspected foul-play, and in most scenarios, the bean-counters would just write this off and be done with it. But Stoll and his co-workers decided to dig, and found a rabbit hole so deep it literally did go all the way through to the other side of the world.

It's a principle that technology professionals need to keep in mind at all times. Yes, you will encounter that case eventually. The numbers will get that big. That reference will find some way to be null. There will be someone with 25 dependents. It's not enough to make sure it works for the most typical cases. Of course that's where to start coding, and where to focus the majority of the testing. But you can't ignore the weirdo edge cases - finding and understanding them can expose hidden problems, as well as opportunities. A lot of revenue can be lost by systems that are leaky around the edges, and a lot of unrealized revenue can be tapped by looking in the same places. Don't assume there aren't any outliers, because there always are. Find them, and either bring them in line, or exploit the opportunity they present you!

And the thing about outliers is that they can always pop back up. You think you've fixed that bug, only to see some more examples of it weeks after the release. You write off a 'harmless' variation each month, only to realize it added up to a pretty big loss by year's end. Again, always assume there are outliers, and put checks and balances in your systems to find them. Sometimes this mean making sure your test suite is robust enough; other times it means having an independent audit system.

This is especially important these days. I'm not suggesting that most weird events can be explained by a hacker, but you never know. Anyone who reads the news knows you can't be too careful. Don't assume your security is good enough! If something seems suspicious, figure it out! Outliers can just be bugs, but they can also be attacks.

Very few people are actually lazy, but we can all be lulled into a false sense of security. Just because the fire alarm isn't going off doesn't mean there aren't any fire hazards. We don't need to be paranoid, or rule-bound, but we do need to be vigilant and thorough. How one maintains such an attitude is probably the subject for a whole book and in the end is different for everyone, but it's a skill that technology professionals must develop.

Sunday, October 28, 2012

Software Development Roles

In software development, there are four major stakeholders, or roles to play: the customer, the architect, the engineer, and the operator.
  • The Customer: The customer is most typically the business analyst, executive, or client that is driving the production of software. They are the people who write the checks.
  • The Architect: The architect is the person or persons responsible for designing how the system(s) work. They must take a holistic view and are responsible for guiding software to a place that is reliable, testable, efficient, and de-coupled.
  • The Engineer: The engineers are those who actually build the software, who sit down and write code based on the customer's needs and the architectural designs.
  • The Operator: Operators come in two very different classes: end-users and operations groups. If you release software for public consumption, your operator is the end-user. If however your software runs on company servers and/or workstations, then the systems personnel are the operators.
The ideal scenario, the sweet-spot where quality, sustainable software is developed, is where each of these different roles work in concert with all the others. The problem of course is that they often don't. The customer is impatient or bombastic, demanding features and deadlines that turn the architects and developers into slaves to his whims, too harried to do their jobs correctly. The architect is overbearing and tyrannical, over-planning the system to the point that it can never be completed and insisting all development pass through him. The engineers are slap-dash, churning out code without regard to system performance, readability, or bug count. The operators are too cheap, unwilling to provide enough servers to handle the load gracefully.

Hopefully no one has ever been in a position where all four of these sentences were true (or at least didn't have to work there long), but everyone's been involved in a job or project where at least one of them was. When one group wields too much say over the software development cycle, problems ensue. Each role should have autonomy in their own domain - their working environment should not be dictated to them by another role. However, for each role to enjoy such independence, there has to be a certain amount of give and take between them. For instance, if your engineers want to develop ASP.NET websites, the operators can't very well insist on Apache servers. If your operators are largely iMac owners, it would be unrealistic for the engineers to decide to write C# desktop apps. If the architect thinks it would be swell to use SharePoint, it's not his/her place to demand that the customer abandon the current CMS system they've come to know and love. Deciding how a piece of software (or software system) comes together requires negotiation and collaboration between the different groups.

One mistake commonly made is to assume that these roles have to be separate people. Naturally, in small companies, particularly start-ups, a small group of IT people will wear multiple hats. Freelancers, consultants, and one-man-IT-shops will often wear the architect, developer, and operator hats simultaneously and exclusively. There is a tendency, though, as the company grows, for these roles to become different departments. There's not necessarily anything wrong with that. But how a company sets up its reporting relationships should not determine how software development roles are filled. Software quality is improved and innovation is fostered when individuals who are engineers 90% of the time are allowed to be architects when the time is right. Architects who step into developer shoes produce more realistic systems. Operators who can be the customer sharpen system requirements. Fostering an environment where this kind of 'cross-pollination' is looked on favorably should certainly be a priority for all involved parties.

I once worked in the Information Systems Division of a Fortune 100 company. They had a lot of clumsy processes and a lot of the folks who worked there weren't standouts in their fields. But their environment was set up in such a way that each role had the autonomy they needed. An architectural group had created an overall vision for how the various systems should fit together, and the development groups were expected to follow that vision. However, inside their own domains, development teams had enormous flexibility in choosing the programming language, OS platform, and design patterns to use. The operators, the infrastructure teams that maintained the servers and terminals that ran this software, had a finite list of platforms and runtimes they would allow, but it was a long list, and anything on that list they would fully support.

My current job employs a much higher caliber of person, and has a much better development life cycle. They have, however, struggled with finding this correct balance between the stakeholders. Despite the flaws of that previous position, I find myself looking back at how they did things as a guide in this area. On paper, this sounds like a rather abstract and theoretical discussion, and in some IT shops it might be. But when this balance is off and/or these roles are not clearly defined, your workday can quickly transform into a series of turf wars. If you find yourself getting into that kind of situation, it's time to take a step back, define, and balance. We didn't, and it went badly: we thought we'd 'won' the turf war, only to have the problem come back at us sideways and make things worse. Don't let this happen to you!