Wednesday, March 13, 2013

How to get a list of objects from all databases on a SQL Server host

SQL Server has a very complete metadata catalog, which comes in handy a lot more often than you might think. One limitation of it though, is that all the metadata is stored on a per-database basis, which makes it difficult to correlate information between databases. (The metadata is stored in system tables in each individual database.)

Last week, I needed to gather a list of stored procedures from multiple databases. I found some solutions online, but most of them involved spitting out multiple results sets, which wasn't particularly useful for me.  After poking around for a little while, I came up with the following solution:

USE [master]
GO

DECLARE @SchemaName VARCHAR(50) = 'dbo';
DECLARE @sql AS VARCHAR(MAX) = '';

SELECT 
 @sql = @sql +
 'SELECT ''['+name+'].['+@SchemaName+'].'' + name AS procedure_name ' +
 'FROM ['+name+'].sys.procedures ' + 
 'WHERE schema_id = '+ 
  '(SELECT schema_id ' +
  ' FROM ['+name+'].sys.schemas ' + 
  ' WHERE name = '''+@SchemaName+''') ' +
 'UNION ' + CHAR(10)
FROM sys.databases
WHERE [state] = 0

SET @sql = LEFT(@sql, LEN(@sql)-8)

EXEC(@sql)
GO

(The sys.databases table is available in every database, not just master. Operating out of a system database just seemed to make more logical sense since the whole point is to gather data from 'user' databases.)

Essentially, what this is doing is generating the following statement for each database:

SELECT '[db_name].[dbo].' + name AS [procedure_name]
FROM sys.procedures 
WHERE schema_id =  
 (SELECT schema_id 
  FROM sys.schemas
  WHERE name = 'dbo')

and then UNION-ing all the results together. It involves a little 'magic' in that @sql = @sql + (stuff) statement, which basically makes it so each row emitted by the SELECT statement adds to the value of @sql (which is why it must be pre-populated as '' instead of NULL.)

I want to emphasize that dynamic SQL like this is risky - like any other kind of SQL statement building, it opens you up to SQL injection attacks, even if contained in a stored procedure. As a script that you store on your machine and run ad hoc though I think it's a good solution.

Monday, March 11, 2013

Complex .NET config transformations

In my previous post, I talked about (web).config transformations: how great they are and how they can be enhanced beyond the 'base install.' Continuing on, some of those enhancements later led us to some even more useful and advanced actions.

Using multiple configuration files in a .NET project is not the simplest thing to achieve. There are different mechanisms for including/merging files, but each have their limitations. One way is to specify that a section comes from another file, like so:

<connectionStrings configSource="otherfile.config"/>

The limitation with this approach is that the entire section must come from that other file. There is no 'merging' of elements; you couldn't add any additional elements to the <connectionStrings> section in this scenario.

You can get around this by importing the other config this way instead:

<appSettings file="otherfile.config">
   <add key="PagesToHide" value="AdminPage" />
   <add key="ExternalLinksToHide" value="Utilization Report" />
</appSettings>

The limitation with this approach is that this kind of import is not supported for all types of config sections:  <system.servicemodel>, for example, can't be used this way. (Also, some code inspection tools, like ReSharper, don't know how to parse this syntax.) And both these approaches force the imported file to be only that one config section; you couldn't have an <appSettings> section and a <connectionStrings> section in an external file and merge/import them both into your config.

In some scenarios, these limitations are not a problem. Both of these approaches have served us well (on a limited basis) in the past. Recently however we had the need to include a non-trivial set of configuration values into the config of multiple applications, most of which were already using web.config transformations. I generally assume that, as a developer, if I have copied and pasted something then I have failed. I wanted to find a more elegant, maintainable solution than just making each development team copy and paste the values into their individual base and transformation configs.

The solution we eventually came up involves a multi-step transformation. Don't be daunted by that "multi" though - it's actually quite a simple implementation.

We started by renaming the web.config (and its children web.debug.config & web.release.config) to web.base.config (children: web.base.debug.config & web.base.release.config). The content of these files we left untouched. We then added the files gateway.debug.config and gateway.release.config to the project. Each one looked something like this:

<?xml version="1.0" encoding="utf-8" ?>
<configuration xmlns:xdt="http://schemas.microsoft.com/XML-Document-Transform">
  <appSettings>
    <add key="Thumbprint" 
         value="52CD92D192786742DA589FEA4C83719DA43E82C9" 
         xdt:Transform="Insert" />
    <add key="Version" 
         value="4" 
         xdt:Transform="Insert" />
    <add key="ConnectionString" 
         value="Data Source=db_host;Integrated Security=True" 
         xdt:Transform="Insert" />
  </appSettings>
  <system.serviceModel>
    <bindings>
      <basicHttpBinding>
        <binding name="BasicHttpBinding_Gateway"
                 closeTimeout="00:00:10" 
                 openTimeout="00:00:10" 
                 receiveTimeout="00:00:10" 
                 sendTimeout="00:00:10"
                 xdt:Transform="Insert">
          <security mode="Transport" />
        </binding>
      </basicHttpBinding>
      <wsHttpBinding>
        <binding name="WsHttpBinding_Gateway"
                 closeTimeout="00:00:10" 
                 openTimeout="00:00:10" 
                 receiveTimeout="00:00:10" 
                 sendTimeout="00:00:10"
                 xdt:Transform="Insert">
          <security mode="Transport">
            <transport clientCredentialType="Windows"
                       proxyCredentialType="None"
                       realm="" />
            <message clientCredentialType="Windows"
                     negotiateServiceCredential="true" />
          </security>
        </binding>
      </wsHttpBinding>
    </bindings>
    <client>
      <endpoint address="https://fake.url/GatewayService/4/a.svc/basic"
          binding="basicHttpBinding" 
          bindingConfiguration="BasicHttpBinding_Gateway"
          contract="namespace.interface"
          name="GatewayTransport_1"
          xdt:Transform="Insert" />
      <endpoint address="https://fake.url/GatewayService/4/a.svc/roles"
          binding="wsHttpBinding" 
          bindingConfiguration="WsHttpBinding_Gateway"
          contract="namespace.interface2"
          name="GatewayRoles_1" 
          xdt:Transform="Insert" />
    </client>
  </system.serviceModel>
</configuration>

The important parts to notice here are the xdt:Transform="Insert" attributes in each XML node. With XML transformations, it's possible to add (insert) elements as well as modifying and deleting them. (The xmlns attribute in the <configuration> element is also very important; without it Visual Studio doesn't know the file is a transformation document.)

The 'magic' comes in with a build step added to the project file:

<Target Name="BeforeBuild">
   <TransformXml 
       Source="web.base.config"
       Transform="Gateway.$(Configuration).config"
       Destination="obj\$(Configuration)\web.intermediate.config" />
   <TransformXml 
      Source="obj\$(Configuration)\web.intermediate.config"
      Transform="web.base.$(Configuration).config"
      Destination="web.config" />
</Target>

Now when the project builds, the compiler takes the base config and transforms it with the appropriate gateway config. Since the gateway config only has addition transformations, this effectively works like a merge of the two files. Then, the build configuration-specific transformation is performed, updating the elements that originally came from the base config to their build-appropriate values. This produces the web.config for the build configuration the solution was built in.

Like approaches discussed in the previous article, this solution is very seamless because it happens at compile-time, so you're able to validate the result at any point, and the web.config that IIS wants is always there. Plus, the gateway configs could be dropped into any project with ease because they just add on to what's already there. The one drawback is that you do have to rebuild to get IIS to pick up configuration file changes, whereas usually you can just save the file and refresh the web page. (This also makes the Visual Studio context menu item 'add config transform' not work, but then it doesn't always work anyway, and adding files to a project is a pretty trivial task.)

The example used here is for a web.config, but this should all work exactly the same with an app.config (so long as the build task is imported - see previous article.)

The possibilities presented by this technique are endless - you could do some pretty complex and powerful things by chaining transformation steps together. You don't want to have too many config files, but if you need to bring separation of concerns or just better readability to your application configuration, this is a powerful and elegant way to do it.

Wednesday, January 30, 2013

A simpler approach to .NET Config File Transformations

One of the nicer features of ASP.NET 4.0 is the concept of configuration file transformations. In case you are not familiar with it, here's a quick two-paragraph overview:

ASP.NET websites almost always need a different configuration for each deployment environment. The development/test configuration will differ in important ways from the staging configuration, and again from the production/live configuration. However, these differences are often only a smart part of the configuration file: a few URLs, server names and assorted settings. Before 4.0, the typical way to manage this was to maintain a separate, but 90% identical, web.config for each environment, and somehow make the right file the 'real' web.config at deployment. Keeping the files in sync was not a trivial task, especially with large projects. We had one application with a config that was well over a thousand lines, and the various environment 'editions' have gotten so out of sync with each other that we spent an embarrassing amount of time squashing bugs that would only manifest in one environment (which of course was usually production).

Web.config transformations brings some relief to this problem. With this feature, you have the web.config, and then it has several child files, one for each build configuration. The primary web.config contains all the information, and then each child file contains instructions for modifications. For example, the primary web.config might contain a section defining the test database connection string, and then the child file web.release.config contains an instruction to replace that section with the production database connection string. When you publish the website, it will transform the primary web.config based on the child config file that corresponds to the build configuration being published, resulting in a file that is valid for the environment. (These 'instructions' are an XML transformation language developed by Microsoft.)

As useful as this feature is, there are some difficulties.

The main difficulty is that transformations are not created at build time. This can make them a little hard to validate, because you have to go out and use a secondary tool or script to execute the transforms and check them over. You can do this from the command-line like so:

msbuild /nologo /target:TransformWebConfig /p:Configuration=Release IISHost.csproj

This will create the the transformed config (for the Release build configuration) in obj\Release\TransformWebConfig\transformed\Web.config.

In the spirit of making this a usual part of the build process, we attempted to include this as a post-build event. However, because this MSBUILD task actually compiles the project, this creates a circular dependency where the project is built, and then the post-build event kicks off, which builds the project a second time, which kicks off the post-build event, which keeps looping back on itself infinitely until your computer crashes.

Not generating the transforms at build-time also causes issues for deployment. Transforms are generated when a website is published, but the Publish mechanism is not the best or preferred way to deploy. In enterprise environments especially, the deployment team typically does not have development tools such as Visual Studio or MSBUILD installed, and generally just wants to have a folder of files they can copy to the correct directory. (Publish also does not lend itself to managing backups, versioning, or rollbacks.) So build engineers and deployment teams have to manually execute the transform, and then go through and copy and rename files just as they did before ASP.NET 4.0.

The third issue is that it's also only available for web.config files, not for app.config files. Now, there are plug-ins that add this feature for non-web projects, such as SlowCheetah. SlowCheetah is a good solution, but I was always hopeful I could find a simpler, more integrated approach. A number of articles, especially this one, told me it was possible, but their solutions were more complex than I really wanted. They did however guide me to this conclusion:

It is possible to execute the transform as a native part of the build by adding some elements to the project file. These elements cannot be added through the Visual Studio GUI (as far as I know), so you do have to manually edit the project file, but these additions are very small, simple bits of XML, so it's not too daunting.

In a website project, at the bottom of the file (just before the </Project> element), add the following:

<Target Name="AfterBuild">
    <TransformXml Source="web.config"
                  Transform="web.$(Configuration).config"
                  Destination="obj\$(Configuration)\web.config" />
</Target>

The target name is very important; that's the part that tells the compiler to run the TransformXml task as an after-compile task. (Because it is not making an external, recursive call, the afore-mentioned infinite loop condition is not created.) Once this section is added, every build will produce, in the obj folder corresponding to the build configuration just compiled, the transformed web.config! The beauty of this solution is that the transformed config is always available when you want to validate it (or when build engineers or deployment teams need to copy it) but the default, development web.config remains in place and untouched.

TransformXml is included by default for website projects. You can do the same thing for app.config files, but you have to explicitly import the task in non-web project files. You can do this by including the following line before the <Target> element:

<UsingTask 
   TaskName="TransformXml" 
   AssemblyFile="$(MSBuildExtensionsPath)\Microsoft\VisualStudio\v$(VisualStudioVersion)\Web\Microsoft.Web.Publishing.Tasks.dll" 
/>

(The DLL referenced here is included as part of the Visual Studio install, so you shouldn't have to add any packages or features. It does not need to be included in the Project References either.)

Being able to automate the configs in this manner has been very helpful for us. In my next post, I'll go into detail about how we were able to leverage this knowledge to accomplish even more advanced tasks.

EDIT: If you have other post-build tasks that rely on the transformed web.config, you should use "BeforeBuild" as the target name. There appears to be a slight delay between the build finishing and the final file being written out. To avoid any race conditions, use "BeforeBuild" instead of "AfterBuild".)

EDIT 2 (7/15/2016): I have modified the <UsingTask> path to use $(VisualStudioVersion) so the project file is more maintainable.

Sunday, October 28, 2012

Software Development Roles

In software development, there are four major stakeholders, or roles to play: the customer, the architect, the engineer, and the operator.
  • The Customer: The customer is most typically the business analyst, executive, or client that is driving the production of software. They are the people who write the checks.
  • The Architect: The architect is the person or persons responsible for designing how the system(s) work. They must take a holistic view and are responsible for guiding software to a place that is reliable, testable, efficient, and de-coupled.
  • The Engineer: The engineers are those who actually build the software, who sit down and write code based on the customer's needs and the architectural designs.
  • The Operator: Operators come in two very different classes: end-users and operations groups. If you release software for public consumption, your operator is the end-user. If however your software runs on company servers and/or workstations, then the systems personnel are the operators.
The ideal scenario, the sweet-spot where quality, sustainable software is developed, is where each of these different roles work in concert with all the others. The problem of course is that they often don't. The customer is impatient or bombastic, demanding features and deadlines that turn the architects and developers into slaves to his whims, too harried to do their jobs correctly. The architect is overbearing and tyrannical, over-planning the system to the point that it can never be completed and insisting all development pass through him. The engineers are slap-dash, churning out code without regard to system performance, readability, or bug count. The operators are too cheap, unwilling to provide enough servers to handle the load gracefully.

Hopefully no one has ever been in a position where all four of these sentences were true (or at least didn't have to work there long), but everyone's been involved in a job or project where at least one of them was. When one group wields too much say over the software development cycle, problems ensue. Each role should have autonomy in their own domain - their working environment should not be dictated to them by another role. However, for each role to enjoy such independence, there has to be a certain amount of give and take between them. For instance, if your engineers want to develop ASP.NET websites, the operators can't very well insist on Apache servers. If your operators are largely iMac owners, it would be unrealistic for the engineers to decide to write C# desktop apps. If the architect thinks it would be swell to use SharePoint, it's not his/her place to demand that the customer abandon the current CMS system they've come to know and love. Deciding how a piece of software (or software system) comes together requires negotiation and collaboration between the different groups.

One mistake commonly made is to assume that these roles have to be separate people. Naturally, in small companies, particularly start-ups, a small group of IT people will wear multiple hats. Freelancers, consultants, and one-man-IT-shops will often wear the architect, developer, and operator hats simultaneously and exclusively. There is a tendency, though, as the company grows, for these roles to become different departments. There's not necessarily anything wrong with that. But how a company sets up its reporting relationships should not determine how software development roles are filled. Software quality is improved and innovation is fostered when individuals who are engineers 90% of the time are allowed to be architects when the time is right. Architects who step into developer shoes produce more realistic systems. Operators who can be the customer sharpen system requirements. Fostering an environment where this kind of 'cross-pollination' is looked on favorably should certainly be a priority for all involved parties.

I once worked in the Information Systems Division of a Fortune 100 company. They had a lot of clumsy processes and a lot of the folks who worked there weren't standouts in their fields. But their environment was set up in such a way that each role had the autonomy they needed. An architectural group had created an overall vision for how the various systems should fit together, and the development groups were expected to follow that vision. However, inside their own domains, development teams had enormous flexibility in choosing the programming language, OS platform, and design patterns to use. The operators, the infrastructure teams that maintained the servers and terminals that ran this software, had a finite list of platforms and runtimes they would allow, but it was a long list, and anything on that list they would fully support.

My current job employs a much higher caliber of person, and has a much better development life cycle. They have, however, struggled with finding this correct balance between the stakeholders. Despite the flaws of that previous position, I find myself looking back at how they did things as a guide in this area. On paper, this sounds like a rather abstract and theoretical discussion, and in some IT shops it might be. But when this balance is off and/or these roles are not clearly defined, your workday can quickly transform into a series of turf wars. If you find yourself getting into that kind of situation, it's time to take a step back, define, and balance. We didn't, and it went badly: we thought we'd 'won' the turf war, only to have the problem come back at us sideways and make things worse. Don't let this happen to you!

Friday, October 5, 2012

MVC View Compile Workarounds, or the SDK Strikes Back

I have resisted learning/using ASP.NET MVC for a while now, for several reasons. One, I'm a cranky old codger who doesn't like to change. Two, I don't hate the viewstate or the page life cycle, so when Microsoft proclaims that MVC will save us from those two pains, I go, "what pains?" Third, when I look at MVC code, I find myself grumbling, "If I wanted to be a PHP programmer, I would have stuck with that."

But time and upper management marches on, and recently I was pulled onto a project where the primary deliverable is an MVC 4 website. So I had to finally bite the bullet and get started. After installing the necessary components and add-ins and actually getting all the projects to load, I tried to build the solution, and got a compile error. The message was of the standard "The type or namespace name '...' does not exist in the namespace '...' (are you missing an assembly reference?)" variety; the breaking files, however, were entries like this:

c:\Users\nirvin\AppData\Local\Temp\
Temporary ASP.NET Files\temp\0260edc4\5834db2c\App_Code.q2tzfscm.0.cs

What? Why was Visual Studio generating code and then blaming me that it didn't build it right? Why should I care about the temporary ASP.NET files at this stage? I was trying to build, not run. I did some digging and found that Visual Studio attempts to compile the MVC views at build-time, instead of waiting until run-time (as standard ASP.NET does).

I consulted with my co-workers and determined that the issue was partially with my machine and partially with the solution. I could create an MVC 4 project from scratch and build it just fine, but this existing solution would fail consistently. However, other developers on the team were able to build the solution just fine on their machines. After more consultation and Googling, the following solutions were proposed:
  • Install Visual Studio 2012.
  • Uninstall MVC 2, 3, and 4 from my machine and then re-install them.

Neither of these options were appealing to me. I had not yet installed Visual Studio 2012 on my machine, and I knew that such a process would be time-consuming. The base install would take forever (after I found the right ISO that is), and then several of the critical add-ins and packages I use for daily development would have to be upgraded and reconfigured (e.g., ReSharper, dotCover, NuGet). Plus, I knew some of my teammates were using Visual Studio 2010 and building just fine, so I was somewhat skeptical of the proposed solution anyway.

The uninstall/re-install option, though, just rankled me. Ever since my my bad experience with Silverlight, SDK versions that refuse to play nice with one another (or, to put it another way, aren't properly isolated from one another) are a pet peeve of mine. Just as with the Visual Studio 2012 option, I didn't want to spend hours re-configuring my machine, and I certainly was not willing to play the install dependency guessing game.

After hours of trying to resolve the error, I finally found that you can actually turn off this "compile views" step.  While I didn't like this option, I needed to get back to actual work (particularly since my assigned tasks were all down in the data layer) so I decided to give this a shot. (I realize this might seem stupid; please keep reading, an endnote addresses the paradox here.) This particular setting is not accessible through Visual Studio 2010's UI, but can be toggled by editing the project file. Inside the project file there is a line that reads:

<MvcBuildViews>true</MvcBuildViews>

I switched it to "false", reloaded the project file, and built. Success! Well, not quite. I didn't want to check this change into source control; this was a personal choice/setting, and I didn't feel right imposing it on the rest of my team. It is, after all, a hack: telling the compiler to skip certain steps is not exactly advisable. I also didn't want to leave the project file checked out all the time; eventually I would accidentally check-in the hack, or have to make a legitimate change to the project file. Plus seeing stuff in the Pending Changes window at the end of the day just bothers me.

Now, the same variables/expressions that are available in pre- and post-build events are available in the project file markup and get evaluated at build-time. At first I looked into creating some rule/exception based on my machine name, but that variable wasn't readily available. So I hit on the idea of creating a separate build configuration instead.

I created a solution build configuration called "DevNoViewChecks" (emphasis on solution; the "Create new project configurations" option was not selected). This build configuration was copied from Debug, and so upon first creation it had each project set to "Debug". I then went to the MVC project, and created a new project build configuration for it with the same name (again with the emphasis; the "Create new solution configurations" option was not checked). This gave me a solution configuration "DevNoViewChecks" with each project set to "Debug", except for the MVC project, which was set to "DevNoViewChecks". I then went into the MVC project file, and edited the afore-mentioned line to read:

<MvcBuildViews Condition="'$(Configuration)' != 'DevNoViewChecks'">true</MvcBuildViews>

I then changed my build configuration in the Visual Studio toolbar to say "DevNoViewChecks", and built. No compile errors! Success! Actual Success!

I don't love that I had to hack this, and I'm not exactly warming to MVC here. But I do love that with this approach, no one has to think about it. I keep my solution set to "DevNoViewChecks", everyone else keeps themselves set to "Debug", and everyone's solution just works. No one has to care about what the other is doing and I don't have to think about what's going on with my project file all the time.

Someday I hope to actually resolve this issue with a deterministic approach that actually makes sense. But until then, this keeps me going.

(It is important to note that the website runs perfectly in the "DevNoViewChecks" build. There seems to be some inconsistency between how Visual Studio (attempts to) compile the views, and how the ASP.NET runtime actually compiles them. This compile error feels like a false negative.)

Monday, August 22, 2011

Use this data, not that

The Salt Lake Tribune has recently been running a series of articles dealing with some privacy and security concerns that a fraud probe into a Utah prenatal health care program raised. The articles are primarily immigration-themed but they are also eye-opening from a software design perspective. One of the articles focuses on the fact that the software Utah uses required them to enter a Social Security Number (SSN) for patient identification. Some of the women coming into the clinics were unable or unwilling to provide this information, so the clinic staff would issue them 'dummy' SSNs to get them into the system. This eventually caused an issue because one of the dummy numbers entered happened to match the real SSN of a man in Maine. The end result was a case of accidental identity theft.

Anyone who's ever developed software will be un-surprised by the details the article gives about how the data ended up so muddled. The system required a nine-digit ID, so the staff used SSNs. When the SSN was unavailable, they'd make one up. For years they'd prepend a "V" or something to try and distinguish the reals from the fakes, but then an upgrade forced the values to become numeric only. Under both schemas ID duplication was occurring, a fact the staff was well aware of. Changing the ID field's parameters was too expensive, so they just lived with the mess. Investigations by the U.S. Social Security Administration (SSA) only prompted the helpful advice to use a different numerical prefix that the SSA doesn't use in SSNs. The state's processes have been modified to continue doing exactly what they've been doing all along, except now they have to keep a separate (most likely paper) log to be used to sort out any difficulties.

There are some very important lessons about software development that can be learned here: first, that using government-issued numbers as IDs is a very bad practice, and second, that good software design cannot ignore the human element.

SSNs are not used as IDs as much in software anymore, but I think some designers and developers don't really understand why this is the case. We may say "People don't want to give us that information" or "We don't want to be responsible for keeping that data private." While it's good to recognize the inherent privacy concerns, these reasons miss the point, plus most organizations that would use SSN in the first place do have valid reasons to collect it. The real reason SSNs make poor IDs because they cannot be changed, and because they are intended to be a private key.

An example to illustrate: I worked for an automotive shop where the mechanics would track the vehicle work via a touch-screen terminal. The mechanics would log in to said terminals using their SSN. The software running on the terminal communicated only with the server in the back room, and the shop had valid reasons to know the SSNs of the people their customers were entrusting their vehicular safety to. It all seemed like a reasonable setup. But then Employee B found out Employee A's SSN, and began to enter work under Employee A's login. I don't remember why firing Employee B was not an option, but it wasn't. We couldn't change Employee A's SSN without screwing up the payroll system, and we didn't have the resources to redo the terminal software (it was really, really bad code). This left Employee A entirely without recourse.

Using an SSN as a private key to eliminate duplication or provide positive identification for legal purposes is a perfectly valid thing to do. But to use SSN as a username or a public ID number is just wrong. It boxes you into logistical and ethical corners that can be very expensive to get out of.

The complaint is raised that we don't want our users to have to be responsible for yet another number or ID that they have to remember. This is a valid concern, and it has a simple solution: don't do it. Look the patient up by name when they come in the clinic. Issue them a card with the ID number printed on it (and include a barcode or magnetic stripe so it can just be scanned). Issue them an ID badge with an RFID chip. Let them choose a username - these are intended to be public, so people can reuse them ad infinitum. Require SSN as a element of account creation if you must, but store it privately (and securely) and map to it by the public ID of your/their choosing.

David Platt once said, "Your user is not you," and I don't think truer words have ever been spoken. Developers tend to have a certain mental block about this; they assume that because the field says "SSN" or "Email" or "Date of Birth" then that's what the user will enter. But we forget that to a user, a field is not a discreet, re-usable piece of information - it is a post-it note where they can write stuff till they need it again. Users will put information wherever they can fit it, regardless of categorization. A balance has to be struck between making forms daunting or too permissive. Validation goes a long way to helping with this. I work for a company that receives real-time (multiple per second) data feeds from the largest retail chain on the planet. One element in these feeds is email address. We get the data just as the store associate enters it, and since the software on their side does not require any validation at all - not even a check to be sure it includes a "@"! - the email addresses are not viable. A trivial regex would allow this information, which is invaluable for our marketing purposes, to be usable instead of dross. Validation is no magic bullet though - the most rigid validation in the world won't alert you to the fact that the patient's birth-date is not 1/1/1970. Unless for some exceptional reason you can verify the person's birth certificate, there's pretty much no way to independently verify that kind of data, and it would not be worth the effort for you to try. So in that instance the software should simply be aware that this value is not ironclad and may need to be treated with kid gloves.

The bottom line is that as computers and software become more and more ubiquitous, we have to avoid creating any further pitfalls like this.

Sunday, March 20, 2011

An evaluation of Silverlight and XAML

Up till now, I've only used this blog to post about technical issues and development patterns, without really editorializing. This entry, however, is going to be a rather subjective evaluation of technology stacks.

One of the bigger paradigm shifts in .NET development that has occurred in the last few years is the introduction of Silverlight. I'm not going to go into the reasons that Silverlight was introduced or why it has seen a significant rate of adoption (partly because the reasons for the latter are incredibly varied depending on company and application). Hand in hand with Silverlight has come XAML. The two are not one and the same: Silverlight is an application framework that's very tightly coupled to Internet solutions, while XAML is a markup language that can be used in many different types of .NET projects. This post will discuss the pros and cons of both.

Silverlight is nice in that it further simplifies web application programming. Some websites are actually not suited to a model where the user can control page flow, and some user interfaces can only be accomplished in an HTML+JavaScript environment after much trickery and tweaking third-party controls (e.g. jQuery UI). Silverlight allows you to deliver an application in a sever-client and/or web-based manner without having to dance around the stateless, static-content-oriented HTTP protocol.

Silverlight's architecture has some good ideas behind it. Applets have been a source of (real and perceived) security concerns for a long time, so the Silverlight designers decided that Silverlight would be a subset, instead of a super-set, of the .NET framework. In other words, Silverlight uses a smaller selection of .NET's functionality. This deliberate scope restriction takes away the ability of Silverlight applets to do some dangerous things. The code to access local file systems, databases, and other critical resources is not even there. It enforces a safer application ecosystem by design instead of by potentially breakable (and in the end arbitrary) access switches. It also solves (or at least mitigates) another concern of applets. Applets have to have their own execution sandbox that the client has to download, in addition to downloading the applet itself. A smaller functionality set compiles to smaller binaries, which results in smaller download and install footprints. Even with today's fast connections, dual-core processors and hundred-gig hard drives, all resources are still finite, and so creating a paradigm that rewards lighter-weight deliverables is a very smart idea.

These benefits do not come without some significant hassles, though.

One of the biggest problems I've had with Silverlight is that various versions are not compatible. You can install and run .NET 1.1, 2.0, and 4.0 all on the same machine without any problem - you can have solutions that have projects in all different versions; you can have multiple IIS app pools running different framework versions all running simultaneously. Not only is this possible, but it's extremely easy - the setup and execution is all seamless. .NET is certainly not the only application framework or product that can support this kind of parallelism, but the point is that it does it well. Silverlight does not. I know it is possible to get Silverlight 3 and 4 running on the same machine - I've seen co-workers get it done - but its a very difficult process, and even those co-workers throw up their hands in surrender when I ask them to help me re-create it. "I just un-installed and re-installed things in an apparently arbitrary order until it started working" was the answer I got from more than one of them.

On the surface this sounds like a nit-picky concern - "just use the same version of Silverlight for everything", right? But let's be realistic, it's never that simple. Various applications are developed under different constraints and requirements, and sometimes using only one version is simply not a realistic option. Some clients and environments require an older framework, and you can't change that. Plus, even if you do have the option to upgrade, development hours are limited and business users/clients aren't always willing to assume the upgrade risk. This is true for anything, not just Silverlight - there are still many .NET 2.0 applications and DLLs running in production environments that won't be upgraded for years to come for these very same reasons. Effective multi-version support is a feature I don't think enterprise software development tools can skimp on, and I feel that Silverlight not only skimped, but completely dropped the ball.

This versioning/parallelism flaw is major, but there are also some important minor annoyances. The subset mentality that Silverlight was designed in is a good idea that was clumsily executed. Instead of Silverlight being a 'true' subset of .NET, it is actually a parallel, minimized fork of .NET. It looks like .NET, it smells like .NET, but it doesn't taste like .NET. Visual Studio is always cranky when you try to add a reference to a Silverlight project in a non-Silverlight project - it'll do it, and the solution will compile, but VS will always mark it as a broken reference in Solution Explorer. Tools like ReSharper will even give you pre-compile errors in non-Silverlight code that references Silverlight code (as well as in the solution-wide analysis, which is much harder to ignore).

The path of work-arounds this particular flaw sent me down was a real comedy of errors. "Hmm, VS2010 + Resharper doesn't play nice with MSTest projects referencing Silverlight projects. Okay, let's create a Silverlight test project. Hmm, no such project type. Okay, we'll just create a Silverlight class library - the whole 'test project' definition is a somewhat arbitrary distinction anyway. Argh, okay, where can I download the MSTest for Silverlight framework? Geez that's hard to find. Okay got it! Compiles, woo-hoo! What? The MS Test runner can't run the MSTest for Silverlight tests (doesn't recognize the attributes)? Crap. Well, maybe there's a port of the test runner for Silverlight. Uh ... well there's a crappy browser-based version that's difficult to use and hard to see test failures in... Okay, screw it, let's just go with NUnit, I've never really liked MSTest anyway. NUnit port for Silverlight? Unofficially done, but existent and stable! Score one for open-source! Create Silverlight class library, add tests, reference NUnit for Silverlight DLLs ... compiles! Woo-hoo! RUNS! Woo-hoo!"

I share that bit partly to inject a little levity, partly to show that with Silverlight NUnit is nicer than MSTest, and partly to reinforce my argument that Silverlight is a second-class citizen even in the Microsoft world. MSTest doesn't like it, Visual Studio doesn't like it, it doesn't even like itself. There are just so many little 'gotchas' in trying to use Silverlight, functionality that has to be re-created, or specialized ports of existing tools that you have to employ. The whole paradigm just seems to work against code re-use, which is something that makes me rather cranky.

XAML is the markup language that Silverlight uses to create its user-interface components. As mentioned previously, though, it is not tied to Silveright. The Windows Presentation Foundation (WPF), which is intended for desktop applications, also uses XAML. In fact, Visual Studio 2010 itself is written in WPF, and therefore XAML. XAML is a big leap forward in terms of simplicity and portability of UI design - it takes everything that was great about HTML, CSS, and Web Forms, and combines them all into something even better. It further closes the gap between Windows Forms and Web Forms - these two technologies used extremely similar but inherently separate structures, but now everything is united under one roof. You can design for the desktop or the web (as long as that web is Silverlight) using one approach. XAML makes formatting pages/screens much, much more intuitive than setting up CSS stylesheets or creating application themes, and it makes the flexibility of HTML layouts available to desktop apps. Making a desktop application look pretty is no small feat regardless of technology, and WPF gives you a shorter path.

Unfortunately, XAML also takes everything that was bad about ASP.NET Data Grids and makes it the standard. The Model-View Model pattern that XAML is intended to employ encourages injecting property, method, and even class names directly in the XAML markup, or in other words, into uncompiled text. I shudder every time I see this kind of thing being done, whether it's in .config documents, vanilla XML, or in 'magic' strings inside the C#/VB code. Doing this kind of thing works against refactoring. As far as I'm aware, there exists no tool that will extend object refactors into the XAML. Given the fluid nature of the XAML data-binding model, it's a difficult task to hope to accomplish, especially considering the fact that the source object doesn't have to be bound in until run-time. Again, this may seem nit-picky, but I argue that it is not. Code is always changing, and needs to be flexible enough to accommodate rapid change. This need becomes more and more pressing each year. Members in XAML {Binding} or {StaticResources} statements are disconnected from the code in a way that discourages and complicates changes. What's even more concerning to me is that it is very easy for incomplete refactors to go unnoticed. It is very easy to change something, have the {Binding} member no longer match, and then that element no longer shows up on the screen, and no one would even notice, even with the greatest QA department in the world, because no error is thrown when said binding fails. This kind of thing has bitten us more than a few times even with the more strict binding mechanism of ASP.NET Data Grids, sometimes even in production code. I am pessimistic that such occurrences will only increase in a world that relies more heavily on XAML-based implementations.

Now, the good news is that there are ways to get around this flaw in XAML. The traditional, explicit data-binding model of giving controls names and wiring them up in the code-behind can be employed. There are also code-only ways to create {Binding}s using only C# code (no XAML) - they're not as pretty, but they work, and because they eschew magic-string based reflection they are refactor-friendly. I hope to post some examples of my own here before too long.

My current evaluation of Silverlight is that it has too many flaws to justify the somewhat dubious benefits it brings. In the end, traditional ASP.NET websites with a liberal amount of jQuery can provide all the same functionality without any of Silverlight's limitations or contrivances. And if the HTML 5 standard can ever see wide-spread adoption, then Silverlight becomes even less attractive. I would urge .NET developers to discourage the use of Silverlight in order to shorten the time till its end-of-life date.

My current evaluation of XAML (and I reserve the right to modify this in the future) is that it is better than both Windows Forms and Web Forms. While it has some non-trivial pitfalls, they are worth the risk for the benefits gained. I would urge the use of WPF for desktop development, and if it ever becomes available for non-Silverlight ASP.NET use, then it is preferable to 'pure' HTML+CSS.

Comments, questions, and corrections are more than welcome!