Showing posts with label security. Show all posts
Showing posts with label security. Show all posts

Wednesday, February 12, 2025

Cryptocurrency

The argument for the adoption of cryptocurrency goes something like this: "Traditional currencies are based on trust in the government, and you can't trust the government. Let's put our trust instead in an algorithm-based currency, which is objective and mathematical. That'd be a better trust-based currency system."

On the surface, this is a reasonable argument. The problem is that both these assumptions are false.

Government-backed, trust-based currencies are not based on trust in public officials or even really the government itself. They're based on trust in the assets, institutions, and economy the currency operates in. That trust is not based on the efficiency or affinity of those things, but rather their longevity and viability. Take federal land in the United States - there are 640 million acres, worth hundred of billions of dollars. When you combine that with the collective value of all of the private and public land under the sovereignty of the USA, that number goes up to hundreds of trillions. And that's just the land - that doesn't count the value of the U.S.'s gold reserves ($300 billion in 2024) or any other physical resource the U.S controls. And all that pales in comparison to the economic value the U.S. creates through its political and military influence.

This isn't a patriotic thing - the same is true of the U.K. and the British pound, the Euro, and even the economies of smaller countries like Luxemburg. Within the world of NATO at least, you have dozens of countries with their valuable assets and institutions all supporting each other. So trust in practically any first-world currency is trust in a massive, global network of economic stability, not trust in any specific government and certainly not trust in any specific leader or policy.

By contrast, an algorithm - and software in general - is the work of a relatively small number of people. Large, widely-adopted open source projects can attract tens of thousand of contributors, but that's the exception rather than the rule. More significantly, even 25,000 people is a trivial number compared to the billions of people who participate in the world economy. So putting trust in an algorithm is putting trust in a tiny group of people and a trivially tiny piece of infrastructure.

It's also not true that algorithms have no bias. Algorithms incorporate the unconscious biases of their developers by default, and every algorithm is built for a purpose, so they are going to have conscious biases towards inputs and outputs that achieve their goals. That doesn't make them suspect, but it means they aren't objective. We don't need to look any further than social media to know that algorithms can be engineered and primed to exclude certain outcomes, scenarios and groups.

Cryptocurrency is an interesting concept, and it does seem like the math behind it is sound. But that doesn't tell us anything about the health of the implementation. And the adoption of it as an actual financial vehicle is based on unsound logic.

Tuesday, June 28, 2016

The Case for Dealing Directly

Update 8/1/2016: After giving this article some further thought, I added another reason.

Suppose you started a new job, and sat down to look over the codebase. As you dug through things, you found that, instead of using public methods and/or properties on classes, they had written a whole reflection framework to reach into objects and modify the private fields. Then, when you asked why they had done something so very peculiar, they responded by saying, "Oh, this way is much better - we don't have to care about the structure of the objects, or waste time writing accessor methods and properties, and we can spend all our time just writing actual code!"

What? No. What?

We could probably spend a whole series of articles explaining why this is a bad approach, but let's just focus on these reasons:
  1. it works against the language's paradigm
  2. it leads to contrivances that don't use the objects very effectively
  3. it makes figuring out the usage and dependencies of objects very difficult
  4. it has security risks
  5. it is more risky to change
  6. it doesn't scale, and the performance could get pretty bad without a lot of ways to optimize it
This is what dynamic SQL is like, and my reaction to it is the same.

Current or former colleagues of mine might quickly note that I have written my fair share of dynamic SQL, and astute readers might note that I have published dynamic SQL on this very blog before. Dynamic SQL, like reflection, is not a bad thing, and has good uses. But it's generally not good for line-of-business applications, for all the reasons listed above.

Dynamic SQL is code generation. When code generation is employed at design (develop) time, it is a very good thing! Intellisense is the most common form of code generation, and it's so fluent & ubiquitous that we barely even think about it anymore. One of the reasons I sing the praises of ReSharper to anyone and everyone who'll listen (and I'm starting to get there with SQL Prompt) is because of the code generation features. They just make life easier, while still giving you the flexibility to write exactly what you need, because you don't have to accept everything they give you.

But code generation at run-time is a different story. Would you trust an application that generates, compiles, and executes code on the fly? That kind of sounds like the behavior of a virus, or at least the work of someone who doesn't yet realize that they'll have to debug that thing somehow. Sure, it's clever, but it's fraught with security, maintainability, reliability, and performance pitfalls.

There are pieces of code that we have to write that are trivial and repetitive, and that's what code generation is for. But the whole reason programmers haven't all been replaced by robots is that writing good code is not a deterministic activity. It requires a fair amount of specialized knowledge. SQL is no different.

Let's consider each one of these points a little more closely.


It works against the paradigm


In other words, just because we can doesn't mean we should. 

To an application, a database is an external resource, like a file system or an API. Interactions to external resources generally need to be coarse-grained, and resources need to be well-encapsulated. This means that while tables are the basic unit of database storage, they are not the basic unit of database interactions. In an ideal application, any discrete action will incur only one database call - e.g., there's only one call when the page loads, and only one when the user clicks the button to update the page. Relational databases are built very unambiguously on the principle that interactions should be atomic so they can be consistent and isolated.

Dynamic SQL works against this paradigm because it wants to deal with the tables as individual units instead of dealing with endpoints (stored procedures) performing atomic operations. The workaround for this is typically to create a transaction in the application code, which often is a performance anti-pattern - it makes the database layer chatty, and holds locks open longer, which leads to blocking & degraded performance. (There are places where in-code transactions make sense, but they are the exception, not the rule.)

It's true that databases are often under-abstracted and have too much repetition, but the solution isn't to go around the database, it's to employ the mechanisms it does have.


It leads to contrivances


I've never met an ORM framework I really liked (and it turns out there are a lot of other engineers who feel the same way; Google it to find out more). You're constantly having to work around them to get things done properly, and the SQL they generate is almost always a classic example of what not to do in a relational database.

Writing effective and appropriate UPDATE, JOIN and WHERE clauses is not a trivial task. Structuring them in the correct manner requires addressing multiple considerations - the impact to/of indexes, the size of the tables involved, the native change tracking features, replication, etc. Dynamic SQL code generators do not provide any abstraction in this regard: they have to get very detailed in order to provide the level of control needed. In other words, they have to re-create the wheel, and in the end their output is no more optimal, and typically less optimal, than writing the SQL directly in stored procedures.


It makes usage discovery difficult


When dynamic SQL generators build statements, the table, view, column, etc. names are coming from diverse locations, and how they are used is often obfuscated. This makes it hard to know which database objects are actually in use. As with any type of code, when it's difficult to tell what's in use and how it's being used, it's difficult to move forward. It significantly hampers development and debugging efforts. Every database I've ever encountered has had many things in it that everyone is 'pretty sure aren't even used anymore', but because discovering dependencies is difficult even under the best of circumstances, nothing is ever done about it. Dynamic SQL's obfuscation of usage just exacerbates this problem, while direct SQL in stored procedures is much more clear and consolidated.


It has security risks


It is often pointed out that the security concerns of dynamic SQL are mitigated by parameterization. This is true, but it's also true that stored procedures mitigate these risks even further. When the SQL code is being constructed on the fly, it still opens up possibilities for injection (and compile errors) that simply don't exist with stored procedures. (Naturally, stored procedures that build dynamic SQL suffer from these same flaws, and in some ways more so - I include them in the ranks of 'dynamic SQL code generators.' Traitors.)

Stored procedures are also a better way to implement the principle of least privilege. When dynamic and/or inline SQL is involved, the application must be given carte blanche to act on the tables. But with stored procedures, the interactions are much more targeted, and (typically) only one object has to have permissions associated to it. It's more maintainable and granular access, two key elements of good security.

These may seem like minor concerns. They are not. Quite frankly, security in every single piece of software on this planet needs to be better - and not just better, but many orders of magnitude better. Good enough security is not good enough.

It is riskier to change


When every operation is being managed by the same code, the risk of changing that code becomes much greater. If a change is made to a dynamic SQL builder for one scenario, it runs the risk of breaking all the other scenarios - in other words, it is harder to keep changes isolated. Say for example you have code that is generating INSERT statements, and you make a change to accommodate one edge case. In doing so, you run the risk of breaking all the INSERT statements in the application. But with stored procedures, a change to one statement has zero effect on similar statements in other procedures. Changes can be kept isolated, the system is more flexible and we can be more confident in our changes.


It doesn't scale


Typically, the only way to optimize dynamic SQL is to pull it out and make it not dynamic. Yes, indexes are always a part of the optimization equation in a relational database, but they're just one tool, and quite frankly they're usually just the first layer. To really address these matters we have to get into the queries themselves, and when we can't directly control the queries, we're at a serious disadvantage.

Also, compilation is not free. Dynamic/inline SQL execution plans can be cached, and stored procedure execution plans aren't necessarily cached, but the former typically aren't and the latter typically are. It's a significant advantage for the DBAs or the DevOps to be able to look up an execution plan by the procedure name, an impossible task with dynamic and inline SQL. Reducing the number of compilations is also a important step in preventing CPU bottlenecks in SQL Server, something that can only be achieved with stored procedures.


(It is not always possible to use stored procedures - after all, some applications have to interact with databases controlled by third parties. But in these scenarios, inline SQL is preferable to dynamic SQL, because it mitigates or avoids most of these discussed pitfalls of dynamic SQL even though it lacks some of the advantages of stored procedures.)

In the end, this can all be boiled down to one simple idea: deal with the database on its own terms. In that reflection framework example in the first paragraph, the coders are effectively trying to treat their class structure as a set of global variables, or a key-value store. Trying to make something act like something it is not only leads to trouble. Relational databases are fundamentally different from imperative object code, and attempts to treat them like they're not are misguided.

Tuesday, February 3, 2015

Learn the Rules so you can figure out who's breaking them

99.99 percent [of subatomic interactions] are explainable ... Spending your time exploring each particle trail will lead you to conclude that all the particles obey known physics, and there's nothing left to discover ... Most stars are boring; advances comes from studying the weirdies - the quasars, the pulsars, the gravitational lenses - that don't seem to fit into the models that you've grown up with ... Collect raw data and throw away the expected. What remains challenges your theories.
- Clifford Stoll, in The Cuckoo's Egg
The Cuckoo's Egg is an amazing book - it's a memoir, a techno-thriller, and a manifesto all in one. It recounts how, in the mid-1980s, an astronomer-turned-sysadmin at UC Berkley stumbled across a hacker in their system, and upon tracking down the intruder, ended up catching a ring of KGB agents in their first real stabs at cyber-espionage! It's a must-read for anyone who's ever going to write code, and just a darn good read for everyone else.

The quote above (from chapter 3) is a very insightful one because it tells us, in essence, that the really useful details are in the edge cases. Stoll's first hint that something was up was a $0.75 discrepancy in an accounting system. No one suspected foul-play, and in most scenarios, the bean-counters would just write this off and be done with it. But Stoll and his co-workers decided to dig, and found a rabbit hole so deep it literally did go all the way through to the other side of the world.

It's a principle that technology professionals need to keep in mind at all times. Yes, you will encounter that case eventually. The numbers will get that big. That reference will find some way to be null. There will be someone with 25 dependents. It's not enough to make sure it works for the most typical cases. Of course that's where to start coding, and where to focus the majority of the testing. But you can't ignore the weirdo edge cases - finding and understanding them can expose hidden problems, as well as opportunities. A lot of revenue can be lost by systems that are leaky around the edges, and a lot of unrealized revenue can be tapped by looking in the same places. Don't assume there aren't any outliers, because there always are. Find them, and either bring them in line, or exploit the opportunity they present you!

And the thing about outliers is that they can always pop back up. You think you've fixed that bug, only to see some more examples of it weeks after the release. You write off a 'harmless' variation each month, only to realize it added up to a pretty big loss by year's end. Again, always assume there are outliers, and put checks and balances in your systems to find them. Sometimes this mean making sure your test suite is robust enough; other times it means having an independent audit system.

This is especially important these days. I'm not suggesting that most weird events can be explained by a hacker, but you never know. Anyone who reads the news knows you can't be too careful. Don't assume your security is good enough! If something seems suspicious, figure it out! Outliers can just be bugs, but they can also be attacks.

Very few people are actually lazy, but we can all be lulled into a false sense of security. Just because the fire alarm isn't going off doesn't mean there aren't any fire hazards. We don't need to be paranoid, or rule-bound, but we do need to be vigilant and thorough. How one maintains such an attitude is probably the subject for a whole book and in the end is different for everyone, but it's a skill that technology professionals must develop.

Monday, August 22, 2011

Use this data, not that

The Salt Lake Tribune has recently been running a series of articles dealing with some privacy and security concerns that a fraud probe into a Utah prenatal health care program raised. The articles are primarily immigration-themed but they are also eye-opening from a software design perspective. One of the articles focuses on the fact that the software Utah uses required them to enter a Social Security Number (SSN) for patient identification. Some of the women coming into the clinics were unable or unwilling to provide this information, so the clinic staff would issue them 'dummy' SSNs to get them into the system. This eventually caused an issue because one of the dummy numbers entered happened to match the real SSN of a man in Maine. The end result was a case of accidental identity theft.

Anyone who's ever developed software will be un-surprised by the details the article gives about how the data ended up so muddled. The system required a nine-digit ID, so the staff used SSNs. When the SSN was unavailable, they'd make one up. For years they'd prepend a "V" or something to try and distinguish the reals from the fakes, but then an upgrade forced the values to become numeric only. Under both schemas ID duplication was occurring, a fact the staff was well aware of. Changing the ID field's parameters was too expensive, so they just lived with the mess. Investigations by the U.S. Social Security Administration (SSA) only prompted the helpful advice to use a different numerical prefix that the SSA doesn't use in SSNs. The state's processes have been modified to continue doing exactly what they've been doing all along, except now they have to keep a separate (most likely paper) log to be used to sort out any difficulties.

There are some very important lessons about software development that can be learned here: first, that using government-issued numbers as IDs is a very bad practice, and second, that good software design cannot ignore the human element.

SSNs are not used as IDs as much in software anymore, but I think some designers and developers don't really understand why this is the case. We may say "People don't want to give us that information" or "We don't want to be responsible for keeping that data private." While it's good to recognize the inherent privacy concerns, these reasons miss the point, plus most organizations that would use SSN in the first place do have valid reasons to collect it. The real reason SSNs make poor IDs because they cannot be changed, and because they are intended to be a private key.

An example to illustrate: I worked for an automotive shop where the mechanics would track the vehicle work via a touch-screen terminal. The mechanics would log in to said terminals using their SSN. The software running on the terminal communicated only with the server in the back room, and the shop had valid reasons to know the SSNs of the people their customers were entrusting their vehicular safety to. It all seemed like a reasonable setup. But then Employee B found out Employee A's SSN, and began to enter work under Employee A's login. I don't remember why firing Employee B was not an option, but it wasn't. We couldn't change Employee A's SSN without screwing up the payroll system, and we didn't have the resources to redo the terminal software (it was really, really bad code). This left Employee A entirely without recourse.

Using an SSN as a private key to eliminate duplication or provide positive identification for legal purposes is a perfectly valid thing to do. But to use SSN as a username or a public ID number is just wrong. It boxes you into logistical and ethical corners that can be very expensive to get out of.

The complaint is raised that we don't want our users to have to be responsible for yet another number or ID that they have to remember. This is a valid concern, and it has a simple solution: don't do it. Look the patient up by name when they come in the clinic. Issue them a card with the ID number printed on it (and include a barcode or magnetic stripe so it can just be scanned). Issue them an ID badge with an RFID chip. Let them choose a username - these are intended to be public, so people can reuse them ad infinitum. Require SSN as a element of account creation if you must, but store it privately (and securely) and map to it by the public ID of your/their choosing.

David Platt once said, "Your user is not you," and I don't think truer words have ever been spoken. Developers tend to have a certain mental block about this; they assume that because the field says "SSN" or "Email" or "Date of Birth" then that's what the user will enter. But we forget that to a user, a field is not a discreet, re-usable piece of information - it is a post-it note where they can write stuff till they need it again. Users will put information wherever they can fit it, regardless of categorization. A balance has to be struck between making forms daunting or too permissive. Validation goes a long way to helping with this. I work for a company that receives real-time (multiple per second) data feeds from the largest retail chain on the planet. One element in these feeds is email address. We get the data just as the store associate enters it, and since the software on their side does not require any validation at all - not even a check to be sure it includes a "@"! - the email addresses are not viable. A trivial regex would allow this information, which is invaluable for our marketing purposes, to be usable instead of dross. Validation is no magic bullet though - the most rigid validation in the world won't alert you to the fact that the patient's birth-date is not 1/1/1970. Unless for some exceptional reason you can verify the person's birth certificate, there's pretty much no way to independently verify that kind of data, and it would not be worth the effort for you to try. So in that instance the software should simply be aware that this value is not ironclad and may need to be treated with kid gloves.

The bottom line is that as computers and software become more and more ubiquitous, we have to avoid creating any further pitfalls like this.