My blog has moved! Redirecting...

You should be automatically redirected. If not, visit http://ripper234.com and update your bookmarks.

Showing posts with label Programming. Show all posts
Showing posts with label Programming. Show all posts

18 September 2008

Stackoverflow.com

A few days ago, stackoverflow.com launched in public beta. I've played with it a bit and I think it has the potential to become a great tool for programmers.

It's basically a combination of blog/wiki/digg/forum. Seems like a most questions get answers (on various degrees of relevance) within minutes-hours. Here are some of the questions I've asked, in case you're interested.

Here's a review of the beta site (about a month old), but instead of reading it I encourage you to just try it next time you need a programming question answered and Google fails you (don't forget to subscribe to your question's RSS feed).

08 September 2008

11 Tips for Beginner C# Developers

Today I sat with a friend (let's call him Joe), who just switched from a job in QA to programming, and passed on to him some of the little tips and tricks I learned over the years. I'm sharing it here because I thought it could be useful to other people that are new to programming. The focus of this post is C#, but analogous tools and methods exist for other languages of course.

Refactorings (A.K.A Resharper)


(This is just a short introduction. I recommend the book Refactoring: Improving the Design of Existing Code as further reading)

As I've written here before, I just love Resharper. It is the best refactoring tool for C# I know of (even though it's a bit heavy sometimes - make sure to get enough RAM and a strong CPU). Joe asked me about a C# feature called partial classes. He said his class was just too big (4000 lines) and becoming unmanageable, and he wanted some way to break it to smaller pieces. He also said that at his workplace, they lock the file whenever anyone edits it, because it's very hard to avoid merge conflicts on such huge files.

I was happy he wanted to simplify and clarify his code by splitting the file, but the way he thought of doing this was the wrong way.

Tip I: Single Responsibility and Information Hiding

You should strive to minimize the information you require at place in your program. Joe had dozens of unit tests, which all derived from a single base class that contain methods required for all the tests. In a second look, we saw that some of the methods were in fact only needed by some subset of the tests, that were actually logically close.

To solve Joe's problem, we created an intermediate class, which inherited from the common test base class, and changed these classes to inherit from the intermediate class. We then used Resharper's Push Down Member refactoring, which removed the methods from the test base class into the intermediate class. No functionality was changed, but we removed code from the huge 4000 lines class! By continuing this process, we can break down the huge class into separate related classes with Single Responsibilities, and no class would have access to uneeded information (like methods it doesn't care about).

Tip II - Duplicate Code Elimination

I believe the world of software would be a better place if people were not allowed to Copy-Paste code more than a few times a day. Many coders abuse this convenient shortcut and thus create unmaintainable code. The effective counter-measure to the Copy-Paste plague is elimination of duplicate code.

I saw in Joe's long file code that looked similar to this:
alice = Configuration.Instance.GetUsers("alice");
bob = Configuration.Instance.GetUsers("bob");
charlie = Configuration.Instance.GetUsers("charlie");
diedre = Configuration.Instance.GetUsers("diedre");
...
Even though this is a rather simple example of code duplication, I strongly believe even such minor infractions should be dealt with. Every line of the above code snippet knows how to obtain users from the configuration. This knowledge has to be read and maintained by developers. Instead, why not get all the "user getting" code into one place and let us simply write what we want to do, instead of how?

To solve this, I use one of these two techniques:

Extract Method, Tiger Style
  1. Choose a single instance of duplication, and locate any parameters or code that is not the same among all the instances of the duplicated code. In our examples, the username (and the assignment variable) are the only two different things between the four lines of code.
  2. For every such parameter, use Introduce Variable. The end result of this phase should look something like this:
    string username = "alice";
    alice = Configuration.Instance.GetUsers(username);
    bob = Configuration.Instance.GetUsers("bob");
    charlie = Configuration.Instance.GetUsers("charlie");
    diedre = Configuration.Instance.GetUsers("diedre");
    ...
  3. Now, use Extract Method on this code (the first line in our example). This creates a new method that I would call GetUsers, that simply gets a string argument username and reads it from the configuration.
  4. Perform Inline Variable on the variable you created in step 1.
  5. Now, change all the other instances to use this new method and delete the redundant code.
The end result looks like this:
alice = GetUsers("alice");
bob = GetUsers("bob");
charlie = GetUsers("charlie");
diedre = GetUsers("diedre");
Another way to achieve the same refactoring is Crane Style (I'm just enjoying using kung-fu styles here because I saw Kung-Fu Panda not too long ago :). You can use immediately without creating the temporary user, but then you get a method that specifically returns the username for "alice", which is not what you wanted. Nevertheless, this method can be refactoring by applying Extract Parameter on "alice", netting us the same result.

Of course these two examples do not begin to cover the myriad of ways you can and should refactor the code. What's important is that you always keep an eye out on how you can make your code more concise, which in turns leads to readability and maintainability.

Tip III: Code Cleanup

Resharper sprinkles colors to the right of your currently open file. Every such colored line is either a compilation error (for red lines), or a cleanup suggestion. Go over such suggestions and hear what Resharper has to say (in the example below it seems nobody is using the var xasdf, so a quick alt-Enter while standing on it will remove it).



Another thing which you should do is define and run FXCop
rules to perform a deeper analysis on your entire project/solution and spot potential problems.

Source Control


Tip IV - Use Source Control

I almost left this out as this goes without saying, but properly using source control can probably save you more time and money than all the other tips (or cost you if you don't use it). The best source control tool I know of for Visual Studio is of course Team Foundation Server, as it has the best integration with the IDE. Other tools are possible, but you have to have a good reason for choosing something other than TFS (One good reason might be cross-platform development and the desire to keep all your code base in a single repository).

Automatic Unit Tests


At Delver, we use NUnit to write unit tests. In past projects I've used Visual Studio's built in test tool, but I found NUnit to be slightly better, mainly due to the integration with Resharper. Resharper adds a small green button next to every test, and allows you to run your test directly from there instead of looking for the "Test View" window. Tomer told me just last week that he uses a keyboard shortcut for running tests, but for me, this is one shortcut I don't think I'll bother learning (the brain can only hole so much).
Another benefit of NUnit is that it runs the test suite in place in your current source folder, instead of copying everything aside to a separate folder like Visual Studio's tool (this used to takes me gigs of space of old unit test sessions which were rarely if ever used).

Tip V - Write Autonomous Unit Tests

Back to Joe, he has a few tests that don't work right out of the box. Before running tests, he has to manually run a separate application used by his test suite. As his project contains hundreds of tests, I would love seeing this added as an automatic procedure in his TestInit() method. Automating a manual operation, besides saving time for developers, enables you to:

Tip VI - Use Continuous Integration

Tests that nobody runs are no good. Tests that are run once when written and then forgotten are only slightly better. By the same logic, tests that run all the time are the best. Pick and use a Build Automation System. I wrote before about our chosen solution, TeamCity, and to sum it up - we're extremely happy about it. TeamCity runs our tests on every commit, on multiple configurations and build agents, and helps us detect bugs faster.

Know thy IDE


Visual Studio is one powerful tool (not belittling java IDEs like IntelliJ and eclipse which in many cases are better). Learn how to use it's features to your advantage:

Tip VII - Edit And Continue

Suppose that while debugging, you found a bug. You can edit the code, save, and continue the debug session without losing the precious time you took getting to this point.

Tip VIII - Move Execution Location

See the little yellow marker that signifies the current location inside the program being debugged? This arrow can be moved! It took me quite a while to discover this (actually heard about it from Sagie), but if you take and drag this arrow, Visual Studio will rewind or "fast forward" your execution to the desired point. None of the code gets executed, you just skip to where you want to go. Excellent for going back after executing a critical method, and rerunning it as many times as you wish.

Tip IX- Use Conditional Breakpoints

Don't waste your time waiting for some specific value to appear in the watch for an interesting variable - set your breakpoints to stop only when the desired conditions are met (right-click on the red breakpoint circle and choose "Condition").

Tip X - Attach to Process

Got a bug that only happens on production machines? You can attach your IDE to any running process (preferably one that is compiled in debug mode), and debug away. You can also programmatically cause a debugger to attach using System.Diagnostics.Debugger.Lau

Tip XI - Use The Immediate Window

The Immediate window is a great tool for executing short code snippets solely for debugging. You can place a breakpoint inside a method you wish to debug, and then call the method directly from the immediate window (saves you from doing this through Edit And Continue)

10 August 2008

Regex Complexity

Today I got a shocker.

I tried the not-too-complicated regular expression:

href=['"](?<link>[^?'">]*\??[^'" >]*)[^>]*(?<displayed>[^>]*)</a>

I worked in the excellent RegexBuddy, and ran the above regex on a normal size HTML page (the regex aims to find all links in a page). The regex hung, and I got the following message:

The match attempt was aborted early because the regular expression is too complex.
The regex engine you plan to use it with may not be able to handle it at all and crash.
Look up "catastrophic backtracking" in the help file to learn how to avoid this situation.

I looked up “catastrophic backtracking”, and got that regexes such as “(x+x+)+y” are evil. Sure – but my regex does not contain nested repetition operations!

I then tried this regex on a short page, and it worked. This was another surprise, as I always thought most regex implementations are compiled, and then run in O(n) (I never took the time to learn all the regex flavors, I just assumed what I learned in the university was the general rule).

It turns out that one of the algorithms to implement regex uses backtracking, so a regex might work on a short string but fail on a larger one. It appears even simple expressions such as “(a|aa)*b” take exponential time in this implementation.

I looked around a bit, but failed to find a good description of the internal implementation of .NET’s regular expression engine.

BTW, the work-around I used here is modify the regex. It’s not exactly what I aimed for, but it’s close enough:

href=['"](?<link>[^'">]*)[^>]*>(?<displayed>[^>]*)</a>

30 May 2008

The Tao of Programming

(It appears I'm having troubles sleeping today). I just came across The Tao of Programming, and found it to be much enjoyable.

"
A master was explaining the nature of Tao of to one of his novices. ``The Tao is embodied in all software - regardless of how insignificant,'' said the master.

''Is the Tao in a hand-held calculator?'' asked the novice.
''It is,'' came the reply.
''Is the Tao in a video game?'' continued the novice.
'It is even in a video game,'' said the master.
''And is the Tao in the DOS for a personal computer?''

The master coughed and shifted his position slightly. ''The lesson is over for today,'' he said.
"

13 Reasons Java is Here to Stay

Over here (and here is the Google cached version after the website fell because of the Slashdot effect). This article discusses why the language family of java,C, C# and C++ are here to stay, and won't be replaced any time soon by new contenders such as Ruby and Haskell.

I agree. While as Eli likes to say, it's important to know functional languages / paradigms, I really believe that most programming tasks should be done in either the C# or java flavors of Java# (A future hybrid of the Java and C#? It's really sad IMO that two almost identical languages don't join forces and user bases). C# is my current favorite, but I can see benefits to Java as well, not the least of which is the huge open community and plethora of available tools for Java. I do like the fact the functional programming is included in C# 3.0 and will be supported in Resharper 4.0 - a strong code analysis and refactoring tool is essential for developing and maintaining large projects.

To be fair, I'll admit that most of my programming experience, and what I find most interesting, is classical server-side programming (that's what I currently do in Delver). Also, besides a brief encounter with LISP and Prolog in university courses, I haven't had the time/motivation to truly learn families outside of the C family tree.

27 April 2008

Knuth and The City (TeamCity!)

Funny - just today I read that the great Donald Knuth, in a recent interview, dissed unit-testing:

"As to your real question, the idea of immediate compilation and "unit tests" appeals to me only rarely, when I’m feeling my way in a totally unknown environment and need feedback about what works and what doesn’t. Otherwise, lots of time is wasted on activities that I simply never need to perform or even think about. Nothing needs to be "mocked up."


(In this interview he also talked against multi-core paradigm (which may be the only way to keep up with Moor's law due to our CPUs getting hotter and hotter), but that's a different story).

Why is this funny? Because just today at work I installed JetBrains' TeamCity - a very fine unit test and automated build suite. Sadly I didn't have time to finish configuring it today, but I already see the great benefit this will bring to all our development team. It goes beyond a simple automated build system (which is important enough, making sure no member of the team messes something up by accident, opening up new venues for brave code refactoring).

TeamCity aims to take all resource-intensive activities off the developer's computer. Just some of its features are (I didn't have the chance to test them, but I believe JetBrains - these are the guys that make Resharper after all):


  • Code Analysis (finding code and style errors)

  • Code Duplication - automatically find the copy-pastes I hate so much, with configurable granularity (java only, coming soon for .NET)

  • Pre-commit Build - Suppose it is now 8PM, you're about to leave for home, and unfortunately you have a rather large piece of uncommited (checked out) code. What do you do? You do not commit the code. Why? Because there's a good chance it will screw up the build and you don't want to leave the common integration area messy. To the rescue comes the Pre-commit Build feature - you can commit the code, have TeamCity catch your commit before it actually happens, test it automatically on a build agent (dedicated build machine), and then only if the code compiles and all tests pass it will check in the code for you. You do not have to wait for tests anymore! This frees up developer's time for actual development, while preserving the high benefits of unit-testing (that Knuth doesn't see, for some reason).

  • I'm sure it has load of other features I haven't dug up yet - among them the ability to run tests without any commit process "just for fun", it's ability to smoothly handle multiple Agents (each running a configurable number of builds concurrently), supporting multiple source controls, IDEs and development languages, ICQ/RSS/Jabber support, and last but not least a button "Run this test on the fastest available test machine".



So far the systems I've used for unit-testing were:

  1. Manually running unit-tests a couple of times, then forgetting about them :)

  2. A patched-up system that automatically runs tests for you on checkin, with a build status indicator that Gal Golan in my development crew in the army wrote himself

  3. CruiseControl.NET - an open source project that does something similar to the above, only slightly more configurable



As you can see by the length of this post, I'm excited to try on a professional solution to this problem for once :) And for those of you that got this far, let me add that it's free or charge - for up to 3 Build Agents and 20 developers. Huzza!

14 April 2008

TimeoutStream

What do you do when you want to read all the data from a Stream object in .NET? You use StreamReader.ReadToEnd().


Stream stream = ...;
StreamReader reader = new StreamReader(stream);
string data = reader.ReadToEnd();


Apparently there is no sane way to put a timeout on the above logic. If you call it and your stream doesn't have a timeout, you're doomed. Even if you stream has an internal timeout, you could still be doomed - for example, say you are reading from a website that sends you the letter 'A' every second. The timeout on the HttpResponseStream could be used, but still (assuming it's higher than one second), you can't set a timeout to the entire ReadToEnd() operation.

I wrote a small proxy class TimeoutStream that wraps any stream with a total timeout since the moment of its creation. Any blocking operation performed on it will fail if the stream has been created too far in the past.

I could use Asynchronous IO to sort of guarantee completion in this timeout without relying on a timeout on the stream itself - however, it appears to be impossible to do correctly in .NET - there is no good way to kill that IO operation if it does not complete within the allotted timeout (see here)- so far now this is good enough for our needs because we can indeed set a timeout on the HttpWebRequest object itself (in addition to using TimeoutStream).

The code is available here.

09 April 2008

Bold Programming

This concept has been revisited in many Extreme Programming articles so I will not blab about it much, but I do want to raise the point - mainly because it is still not a universal practice.

A key feature of a good programmer is courage. A "cowardly" programmer will see a non-critical problem or an ugly hack begging to be refactored and will shy away from it on the grounds that it his not his concern at the moment and it will probably break the world and introduce bugs to the system. A "courageous" programmer (backup up with properly written unit tests, of course), will refactor the ugliness away, thus incrementally beautifying his code.

This overall process, reiterated over the product lifetime, will either produce a convulated codebase with loops, noodles and baggage (for the "cowardly" programmer), or will result in a clean, easy to use and modular code for the "courageous" one.

While refactoring is a time consuming activity at times, it is well worth the momentarily increased coding time, because the overall coding (including maintenance!) time for him and his entire team will be decreased significantly.

The existence of well written unit tests is crucial to this, for without it you really can't know if you introduce bugs or not. With unit tests, you get a significant "courage boost", allowing you to do major changes that affect many files - for the better - because you are reasonably certain that any change will indeed be detected.



On the other feature of a good programmer - laziness - in another time.

05 March 2008

Shared Items Feed & More

Hello guys, today we have several topics:

Shared Items

I finally really settled on Google Reader instead of a desktop feed reader. The advantage of being able to read RSS everywhere without any hassle outweigh the downsides. Also I get the benefit of easily exporting a Feed of Shared Items (both RSS and Email.

I think I will stop/reduce posting links to interesting items that I find on my RSS and instead just mark them as shared, so if you want to keep using my information filtering services be sure to register :)
In addition, here is a link to all previously shared items.

Google Notebook
If you've recently Googled you may have seen the "Note this" added to every link.



It's a useful new very useful. Upon clicking it copies the current content of said web page into your Google Notebook, a cool service that organizes your web clippings.
It opens up right on your search result page and has a full page interface as well.





Unfuddle free SVN hosting
If you're doing any non-trivial software assignment with partners, you should consider using source control. So far, I've used source control for large projects of course, but never in an assignment from Technion - back when I was an undergraduate student I was largely oblivious to source control and I didn't take any programming courses in my 2nd degree - until now. Now I actually have a few non-trivial homeworks at Managing Data on the WWW (writing an http proxy is one of them), and so far I didn't take up the trouble of setting up a source control. Well, it appears it is no hassle at all, at Unfuddle you can setup a project page, SVN server, RSS and email updates on checkins, project management and more in less than 5 minutes and no cost. Unlike some alternatives, putting your code there doesn't mean it's now open source and free to the world, you get control over who accesses your code.

12 February 2008

What You Need to Know to Work For Me

If you are ever interviewed by me and I happen to ask you, say, to write a function that returns the nth element in the Fibonacci Series, you had better not give me the recursive solution.

It appears that about 90% of the people asked this give this solution straight away, without asking the examiner anything about efficiency. I find this appalling.

Even if efficiency is not always the number one criteria (or even in the top 5), in this instance the recursive solution is much worse (exponential VS linear complexity) than the non-recursive solution AND is not much simpler to code. The difference is about 3 lines of code.

I expect programmers to anticipate and consider such concerns, and even if not asked specifically to give an efficient solution, in such a case they should (unless of course they ask the examiner if efficiency is a concern and get a negative answer).

P.S.

There are some acceptable recursive solutions. For example, from Wikipedia:

public void run(int n)
{
if (n <= 0)
{
return;
}

run(n,1,0);

}

private void run(int n, int eax, int ebx)
{
n--;

if (n == 0)
{
System.out.println(eax+ebx);
return;
}

run(n,ebx,eax+ebx);

}

I don't object to the actual recursion, just to the exponential inefficiency.

13 November 2007

Whys is Open Source so hard?

Today I came with with an idea for a cool feature for RSS clients. What I want is intelligent filtering like StumbleUpon - users would click "I like it" or "I like it not" on RSS items, and the client could learn a user's "taste" and prioritize future items. Google revealed someone just thought of this about 2 weeks ago. Still, I don't want to pay 20$ for an RSS reader, and I don't want a reader that specializes in this, I want this as just another feature of a standard RSS reader, so I went on to try and add this to some existing open source RSS client.

I immediately found NClassifier, an open source .NET text classifier. Looks promising.

For an RSS client I mainly found Rssbandit. A bit ugly, but I could work with that.

Then, crash. I spent several good hours just trying to get the blasted code. At first I naively thought perhaps the code would be attached to some version of the installer. Nope. I went to the CVS, installed TortoiseCVS, and failed to login with a blank password. At this point I decided to try accessing the SVN version, since I'm already familiar with SVN. I managed to get the code, but it did not build due to a LicenseException :(

I figured the SVN version might be different from the CVS, so I went back to my efforts to use TortoiseCVS, and succeeded ... to connect. In SVN, I could just get the entire latest version. In CVS, I had to choose "packages" to download. Trying some packages, I found that TortoiseSVN does download, only it deletes the downloaded file just at the end of the download. ARG!

I really don't understand why this entire process has to be so complicated. This entire process should have taken half an hour, including downloading the code, modifying and initial testing. Why almost every time you download open source code, it never compiles on the first build. Why the SVN and CVS versions are different.

Bottom Line
I still like my idea of intelligent, learning RSS reader. If anyone has more knowledge in the mysterious ways of Sourceforge and wants to help, call.

03 September 2007

Practical Database in C#

Today I took Eli's challenge (whether he intended it as such or not not :).

Among other things, he mentioned on his blog his opinion that LISP is a more powerful programming language than C#. He offered Practical as an example of LISP's power. Practical is a very simple database built using very few lines of codes.

The challenge I took on myself is to code Practical in C# and compare the implementation with the one given in LISP.

It took me 2-3 hours to read the link on Practical and code it in C#, and I wrote a great deal more lines of code than what I saw in the LISP code. So in a simple contest, C# (or just me as a coder) lost.

However, here are my thoughts:


  1. About 80% of the code I wrote is general purpose. It simulates abilities already existing in LISP. General-purpose classes I wrote are:

    • ConsoleReader - Reading any class from the Console.
    • DefaultFormatter> and ListFormatter - Pretty-print any object or list.
    • Functors - Functors in C# are called delegates. I didn't find (nor looked too hard) for functions to combine, negate and manipulate delegates - so I wrote some.
    • MoreXmlSerializer - As a naming convention, I like to name MoreFoo any class that contains methods that I believe should have been in Foo from the start, where Foo is some class supplied by the framework. Starting at C# 3.0, the ability to add methods to existing classes is supported, so I could adjust the naming convention accordingly.
    • StringConverter - Convert a string to to a given (dynamic) type.


    Besides this code, the actual Practical C# code is very small and concise.
  2. The C# code is type safe. I'm not aware of type-safety in LISP, though as I said I'm not a LISP pro.
  3. The code I wrote is in C# 2.0. Some new features in C# 3.0 should help make it a bit more concise (especially lambda expressions and LINQ).


I think this experiment didn't change my opinion greatly. Yes, functional programming is good. It exists in C# as well as in LISP. Yes, LISP handles lists very well. But, not everything is life (or in programming) should be represented as a list.

I'm attaching my code here. If you have suggestions or comments, please let me know. Also, once I install Visual Studio 2008 (I'll probably wait for the final release), I might try coding this in C# 3.0 and see how it helps the cause.

02 September 2007

Good Programmers

See a discussion on the selection of programming languages on Eli's Blog, and my note about general qualities of good programmers.

09 August 2007

Dotnet Web Crawler Speedup

I'm writing a web crawler in C#, and getting it to perform well was really annoying.
I tried simply using ThreadPool.QueueUserWorkItem() to queue up my requests to multiple threads. Each thread just ran WebClient.DownloadString().

While the threads did run in parallel, it turned out WebClient had an inherent lock.
I tried messing with the ConnectionManagementSection, but that turned out read-only.
After some Google, I found that the configuration can only be changed by modifying the machine.config or user.config files! Seems pretty stupid to me.

After doing that simply didn't work either, I found this code that helped me through. I still don't know exactly why WebClient.DownloadString() doesn't work, but after some tweaking I got to about 2.5 pages pre second. Still not top speed, but way better than the 0.5 pages/second I started with.

31 July 2007

RentASanta

http://funny.karmark.org/Funny%20pictures/images/csc_jpg.jpg

This is the logical progression from RentACoder.