My blog has moved! Redirecting...

You should be automatically redirected. If not, visit http://ripper234.com and update your bookmarks.

Showing posts with label C#. Show all posts
Showing posts with label C#. Show all posts

19 September 2008

A Good Question

C# programmers - you should read this question and its many answers, I'm sure you'll learn a thing or two. For example, ThreadStaticAttribute - a clean, type-safe way to get thread local storage in C#.

09 September 2008

Unhandled Exceptions Crash .NET Threads

A little something I learned at DSM today. It appears if any thread in .NET crashes (lets a thrown exception fly through the top stack level), the process crashes. I refused to believe at first, but testing on .NET 2.0 showed it to be true:

(I should really switch to another blog platform, I didn't find a decent way to write code in Blogspot).

class ThreadCrashTest
{
static void Main()
{
new Thread(Foo).Start();
for (int i = 0; i < 10; ++i)
{
Console.WriteLine(i);
Thread.Sleep(100);
}
}

private static void Foo()
{
Console.WriteLine("Crashing");
throw new Exception("");
}
}


According to Yan, the behavior on .NET 3 is to crash the AppDomain instead of the entire process.

08 September 2008

11 Tips for Beginner C# Developers

Today I sat with a friend (let's call him Joe), who just switched from a job in QA to programming, and passed on to him some of the little tips and tricks I learned over the years. I'm sharing it here because I thought it could be useful to other people that are new to programming. The focus of this post is C#, but analogous tools and methods exist for other languages of course.

Refactorings (A.K.A Resharper)


(This is just a short introduction. I recommend the book Refactoring: Improving the Design of Existing Code as further reading)

As I've written here before, I just love Resharper. It is the best refactoring tool for C# I know of (even though it's a bit heavy sometimes - make sure to get enough RAM and a strong CPU). Joe asked me about a C# feature called partial classes. He said his class was just too big (4000 lines) and becoming unmanageable, and he wanted some way to break it to smaller pieces. He also said that at his workplace, they lock the file whenever anyone edits it, because it's very hard to avoid merge conflicts on such huge files.

I was happy he wanted to simplify and clarify his code by splitting the file, but the way he thought of doing this was the wrong way.

Tip I: Single Responsibility and Information Hiding

You should strive to minimize the information you require at place in your program. Joe had dozens of unit tests, which all derived from a single base class that contain methods required for all the tests. In a second look, we saw that some of the methods were in fact only needed by some subset of the tests, that were actually logically close.

To solve Joe's problem, we created an intermediate class, which inherited from the common test base class, and changed these classes to inherit from the intermediate class. We then used Resharper's Push Down Member refactoring, which removed the methods from the test base class into the intermediate class. No functionality was changed, but we removed code from the huge 4000 lines class! By continuing this process, we can break down the huge class into separate related classes with Single Responsibilities, and no class would have access to uneeded information (like methods it doesn't care about).

Tip II - Duplicate Code Elimination

I believe the world of software would be a better place if people were not allowed to Copy-Paste code more than a few times a day. Many coders abuse this convenient shortcut and thus create unmaintainable code. The effective counter-measure to the Copy-Paste plague is elimination of duplicate code.

I saw in Joe's long file code that looked similar to this:
alice = Configuration.Instance.GetUsers("alice");
bob = Configuration.Instance.GetUsers("bob");
charlie = Configuration.Instance.GetUsers("charlie");
diedre = Configuration.Instance.GetUsers("diedre");
...
Even though this is a rather simple example of code duplication, I strongly believe even such minor infractions should be dealt with. Every line of the above code snippet knows how to obtain users from the configuration. This knowledge has to be read and maintained by developers. Instead, why not get all the "user getting" code into one place and let us simply write what we want to do, instead of how?

To solve this, I use one of these two techniques:

Extract Method, Tiger Style
  1. Choose a single instance of duplication, and locate any parameters or code that is not the same among all the instances of the duplicated code. In our examples, the username (and the assignment variable) are the only two different things between the four lines of code.
  2. For every such parameter, use Introduce Variable. The end result of this phase should look something like this:
    string username = "alice";
    alice = Configuration.Instance.GetUsers(username);
    bob = Configuration.Instance.GetUsers("bob");
    charlie = Configuration.Instance.GetUsers("charlie");
    diedre = Configuration.Instance.GetUsers("diedre");
    ...
  3. Now, use Extract Method on this code (the first line in our example). This creates a new method that I would call GetUsers, that simply gets a string argument username and reads it from the configuration.
  4. Perform Inline Variable on the variable you created in step 1.
  5. Now, change all the other instances to use this new method and delete the redundant code.
The end result looks like this:
alice = GetUsers("alice");
bob = GetUsers("bob");
charlie = GetUsers("charlie");
diedre = GetUsers("diedre");
Another way to achieve the same refactoring is Crane Style (I'm just enjoying using kung-fu styles here because I saw Kung-Fu Panda not too long ago :). You can use immediately without creating the temporary user, but then you get a method that specifically returns the username for "alice", which is not what you wanted. Nevertheless, this method can be refactoring by applying Extract Parameter on "alice", netting us the same result.

Of course these two examples do not begin to cover the myriad of ways you can and should refactor the code. What's important is that you always keep an eye out on how you can make your code more concise, which in turns leads to readability and maintainability.

Tip III: Code Cleanup

Resharper sprinkles colors to the right of your currently open file. Every such colored line is either a compilation error (for red lines), or a cleanup suggestion. Go over such suggestions and hear what Resharper has to say (in the example below it seems nobody is using the var xasdf, so a quick alt-Enter while standing on it will remove it).



Another thing which you should do is define and run FXCop
rules to perform a deeper analysis on your entire project/solution and spot potential problems.

Source Control


Tip IV - Use Source Control

I almost left this out as this goes without saying, but properly using source control can probably save you more time and money than all the other tips (or cost you if you don't use it). The best source control tool I know of for Visual Studio is of course Team Foundation Server, as it has the best integration with the IDE. Other tools are possible, but you have to have a good reason for choosing something other than TFS (One good reason might be cross-platform development and the desire to keep all your code base in a single repository).

Automatic Unit Tests


At Delver, we use NUnit to write unit tests. In past projects I've used Visual Studio's built in test tool, but I found NUnit to be slightly better, mainly due to the integration with Resharper. Resharper adds a small green button next to every test, and allows you to run your test directly from there instead of looking for the "Test View" window. Tomer told me just last week that he uses a keyboard shortcut for running tests, but for me, this is one shortcut I don't think I'll bother learning (the brain can only hole so much).
Another benefit of NUnit is that it runs the test suite in place in your current source folder, instead of copying everything aside to a separate folder like Visual Studio's tool (this used to takes me gigs of space of old unit test sessions which were rarely if ever used).

Tip V - Write Autonomous Unit Tests

Back to Joe, he has a few tests that don't work right out of the box. Before running tests, he has to manually run a separate application used by his test suite. As his project contains hundreds of tests, I would love seeing this added as an automatic procedure in his TestInit() method. Automating a manual operation, besides saving time for developers, enables you to:

Tip VI - Use Continuous Integration

Tests that nobody runs are no good. Tests that are run once when written and then forgotten are only slightly better. By the same logic, tests that run all the time are the best. Pick and use a Build Automation System. I wrote before about our chosen solution, TeamCity, and to sum it up - we're extremely happy about it. TeamCity runs our tests on every commit, on multiple configurations and build agents, and helps us detect bugs faster.

Know thy IDE


Visual Studio is one powerful tool (not belittling java IDEs like IntelliJ and eclipse which in many cases are better). Learn how to use it's features to your advantage:

Tip VII - Edit And Continue

Suppose that while debugging, you found a bug. You can edit the code, save, and continue the debug session without losing the precious time you took getting to this point.

Tip VIII - Move Execution Location

See the little yellow marker that signifies the current location inside the program being debugged? This arrow can be moved! It took me quite a while to discover this (actually heard about it from Sagie), but if you take and drag this arrow, Visual Studio will rewind or "fast forward" your execution to the desired point. None of the code gets executed, you just skip to where you want to go. Excellent for going back after executing a critical method, and rerunning it as many times as you wish.

Tip IX- Use Conditional Breakpoints

Don't waste your time waiting for some specific value to appear in the watch for an interesting variable - set your breakpoints to stop only when the desired conditions are met (right-click on the red breakpoint circle and choose "Condition").

Tip X - Attach to Process

Got a bug that only happens on production machines? You can attach your IDE to any running process (preferably one that is compiled in debug mode), and debug away. You can also programmatically cause a debugger to attach using System.Diagnostics.Debugger.Lau

Tip XI - Use The Immediate Window

The Immediate window is a great tool for executing short code snippets solely for debugging. You can place a breakpoint inside a method you wish to debug, and then call the method directly from the immediate window (saves you from doing this through Edit And Continue)

10 August 2008

Regex Complexity

Today I got a shocker.

I tried the not-too-complicated regular expression:

href=['"](?<link>[^?'">]*\??[^'" >]*)[^>]*(?<displayed>[^>]*)</a>

I worked in the excellent RegexBuddy, and ran the above regex on a normal size HTML page (the regex aims to find all links in a page). The regex hung, and I got the following message:

The match attempt was aborted early because the regular expression is too complex.
The regex engine you plan to use it with may not be able to handle it at all and crash.
Look up "catastrophic backtracking" in the help file to learn how to avoid this situation.

I looked up “catastrophic backtracking”, and got that regexes such as “(x+x+)+y” are evil. Sure – but my regex does not contain nested repetition operations!

I then tried this regex on a short page, and it worked. This was another surprise, as I always thought most regex implementations are compiled, and then run in O(n) (I never took the time to learn all the regex flavors, I just assumed what I learned in the university was the general rule).

It turns out that one of the algorithms to implement regex uses backtracking, so a regex might work on a short string but fail on a larger one. It appears even simple expressions such as “(a|aa)*b” take exponential time in this implementation.

I looked around a bit, but failed to find a good description of the internal implementation of .NET’s regular expression engine.

BTW, the work-around I used here is modify the regex. It’s not exactly what I aimed for, but it’s close enough:

href=['"](?<link>[^'">]*)[^>]*>(?<displayed>[^>]*)</a>

28 July 2008

My Arrogance in Finding Bugs

Last night I discovered this piece of code (simplified version):



private void Foo()
{
bool b = false;
new Thread((ThreadStart)delegate { b = true;}).Start();
WaitForBool(b);
}

private void WaitForBool(bool b)
{
while (!b)
{
Thread.Sleep(1000);
}
}



I was immediately filled with disgust. How could someone write such a function (WaitForBool), which is one big bug? Of course the waiting will never be over, because bool is a value type, and no external influence can modify its value.

Later, I realized the fool is me.

I ran into this code a while back, and it did have a "ref" bool, which means it can be modified by the external thread. Resharper helpfully displayed a tip "this parameter can be declared as a value" , which I took without thinking too much about the consequences (I trust Resharper too much sometimes, it appears). So I deleted the "ref" with Resharper's help and created the bug myself without noticing.


(As a side note, of course this method of waiting for a bool should never be used - use ManualResetEvent instead).

Update
Fixed in the next Resharper build (Resharper 4.0.1, build 913), within a few hours of reporting it. Cool.

30 May 2008

13 Reasons Java is Here to Stay

Over here (and here is the Google cached version after the website fell because of the Slashdot effect). This article discusses why the language family of java,C, C# and C++ are here to stay, and won't be replaced any time soon by new contenders such as Ruby and Haskell.

I agree. While as Eli likes to say, it's important to know functional languages / paradigms, I really believe that most programming tasks should be done in either the C# or java flavors of Java# (A future hybrid of the Java and C#? It's really sad IMO that two almost identical languages don't join forces and user bases). C# is my current favorite, but I can see benefits to Java as well, not the least of which is the huge open community and plethora of available tools for Java. I do like the fact the functional programming is included in C# 3.0 and will be supported in Resharper 4.0 - a strong code analysis and refactoring tool is essential for developing and maintaining large projects.

To be fair, I'll admit that most of my programming experience, and what I find most interesting, is classical server-side programming (that's what I currently do in Delver). Also, besides a brief encounter with LISP and Prolog in university courses, I haven't had the time/motivation to truly learn families outside of the C family tree.

14 April 2008

TimeoutStream

What do you do when you want to read all the data from a Stream object in .NET? You use StreamReader.ReadToEnd().


Stream stream = ...;
StreamReader reader = new StreamReader(stream);
string data = reader.ReadToEnd();


Apparently there is no sane way to put a timeout on the above logic. If you call it and your stream doesn't have a timeout, you're doomed. Even if you stream has an internal timeout, you could still be doomed - for example, say you are reading from a website that sends you the letter 'A' every second. The timeout on the HttpResponseStream could be used, but still (assuming it's higher than one second), you can't set a timeout to the entire ReadToEnd() operation.

I wrote a small proxy class TimeoutStream that wraps any stream with a total timeout since the moment of its creation. Any blocking operation performed on it will fail if the stream has been created too far in the past.

I could use Asynchronous IO to sort of guarantee completion in this timeout without relying on a timeout on the stream itself - however, it appears to be impossible to do correctly in .NET - there is no good way to kill that IO operation if it does not complete within the allotted timeout (see here)- so far now this is good enough for our needs because we can indeed set a timeout on the HttpWebRequest object itself (in addition to using TimeoutStream).

The code is available here.

03 September 2007

Practical Database in C#

Today I took Eli's challenge (whether he intended it as such or not not :).

Among other things, he mentioned on his blog his opinion that LISP is a more powerful programming language than C#. He offered Practical as an example of LISP's power. Practical is a very simple database built using very few lines of codes.

The challenge I took on myself is to code Practical in C# and compare the implementation with the one given in LISP.

It took me 2-3 hours to read the link on Practical and code it in C#, and I wrote a great deal more lines of code than what I saw in the LISP code. So in a simple contest, C# (or just me as a coder) lost.

However, here are my thoughts:


  1. About 80% of the code I wrote is general purpose. It simulates abilities already existing in LISP. General-purpose classes I wrote are:

    • ConsoleReader - Reading any class from the Console.
    • DefaultFormatter> and ListFormatter - Pretty-print any object or list.
    • Functors - Functors in C# are called delegates. I didn't find (nor looked too hard) for functions to combine, negate and manipulate delegates - so I wrote some.
    • MoreXmlSerializer - As a naming convention, I like to name MoreFoo any class that contains methods that I believe should have been in Foo from the start, where Foo is some class supplied by the framework. Starting at C# 3.0, the ability to add methods to existing classes is supported, so I could adjust the naming convention accordingly.
    • StringConverter - Convert a string to to a given (dynamic) type.


    Besides this code, the actual Practical C# code is very small and concise.
  2. The C# code is type safe. I'm not aware of type-safety in LISP, though as I said I'm not a LISP pro.
  3. The code I wrote is in C# 2.0. Some new features in C# 3.0 should help make it a bit more concise (especially lambda expressions and LINQ).


I think this experiment didn't change my opinion greatly. Yes, functional programming is good. It exists in C# as well as in LISP. Yes, LISP handles lists very well. But, not everything is life (or in programming) should be represented as a list.

I'm attaching my code here. If you have suggestions or comments, please let me know. Also, once I install Visual Studio 2008 (I'll probably wait for the final release), I might try coding this in C# 3.0 and see how it helps the cause.