Showing posts with label tools. Show all posts
Showing posts with label tools. Show all posts

November 23, 2011

Lunch-N-Learn Ideas? Use Sonar Hotspots

Are you eager to get your team together for a lunch-n-learn with free pizza (and beer), but can't come up with any ideas on which to present?

Or maybe you just want to get heads down and make some high impact, quality improvements to your code base, but don't know where to start.

I've written about one idea for the latter in Risk Homing Metrics. But lately, I've been using another technique to find real, practical ideas for lunch-n-learns and quality improvements: Sonar's Hotspots.

I like Sonar because it aggregates multiple metrics into single views. It also allows you to view various levels of detail: high, overall project stuff for manager types all the way down to lines of code for developers. By using the Hotspots feature, you can quickly find the top five problem areas in various categories = real, applicable ideas.

If you've never used Sonar, take a look at the Hotspots for ActiveMQ. You can also view many of your favorite open source projects' Sonar results on Nemo.

As always with metrics, you need to weed out the false positives and tweak Sonar so that it only picks up real issues. Otherwise, you risk people getting discouraged and not using the tool.

Now, go schedule that lunch-n-learn and order some pizza.

November 22, 2010

Risk Homing Metrics

Originally published 28 May 2009

Recently, I attended a talk by Neal Ford.  He was talking about a couple of metrics you can combine to identify areas for refactoring: cyclomatic complexity and afferent coupling.  He used the ckjm tool to determine what classes were both complex and used by lots of other classes.  Start refactoring these was his recommendation.

I immediately thought of Crap4j; another tool that combines a set of metrics for identifying the riskiest areas of the code base to maintain.  Crap4j implements the CRAP metric, which combines cyclomatic complexity and test coverage, but at a method level.  If a method is both complex and not very well tested, then it's risky to change.

This all led me to a new, more ultimate set of metrics that could be combined to home in on the riskiest areas of a code base:
  1. Code coverage
  2. Cyclomatic complexity
  3. Code execution frequency in the real world
Complex code, executed very often, with low test coverage.

For practical purposes, I like sticking with the granularity of a method.  I can use tools like Cobertura to find the test coverage and JavaNCSS to find the cyclomatic complexity.  (Isn't cyclomatic complexity best applied to the method level anyway?)

That just leaves me with the which-methods-execute-the-most-in-production problem.  This is hard because I can run the other two as part of a continuous build, but I won't be able to identify the hot methods until I get to production and measure true usage.  So the static and dynamic metrics will always be out of sync at some level even if I could get estimated usage through continuous functional and higher level testing.

So what can I do?  I want this to run as part of a continuous build to get feedback as soon as possible that a method is getting a little risky (or with a legacy code base, is already risky).  So, I'll fall back on afferent coupling for practicality.  But afferent coupling is typically measured at a package level.  The finest granularity that I'm aware of with current tools is measuring at a class level with ckjm.  That's a good starting point for identifying highly used code.

So here's my plan.  Use the CRAP metric to find the methods and then factor in the afferent coupling of the those methods' classes to give a prioritized list of methods to go clean up.  I'll see how this goes and consider factoring in method execution frequency from higher level testing runs.

Git And Continuous Integration

Originally published 13 Nov 2008

Subversion is the de facto version control system of the day, but Git is the rising star.  More and more people are using Git, but I'm a bit concerned about its effects on Continuous Integration (CI).

First some background...

Subversion follows a centralized repository model.  So does CVS and many others.  There's one server that everybody commits to.  Git can follow suit or be configured in a distributed fashion.  In a fully distributed model, each developer has a private copy of the repository.  Mercurial is an example of a distributed version control system.

My concern is really with the distributed model.  One of the appeals of this model, and particularly with having your own private repository, is the ability to experiment and check-in/commit at a much finer-grained rate than with a centralized model.  With a centralized model, you need to be more careful of your commits.  Otherwise, you could break the build and everybody else.  So consequently, you commit less frequently here.

But what's better in the context of CI?  I have to be honest and admit that I've never used the distributed model on a project, so I'm being theoretical here.  My gut feel is that with a distributed model, integration will be less frequent.  I believe people will check in to private repositories more often, but will push those changes to the main integration branch less frequently than people committing straight to the main branch in a centralized model.  In a centralized model, integration is in your face.  You can't commit without thinking about it.  In a distributed model, you can get carried away in your own little world.  And we've learned in the agile community that we should be integrating early and often, haven't we?

I compare the distributed model with multiple repositories to a centralized model with multiple branches.  If you can accept this analogy, you can probably see how integration would be less frequent.  I fear with Git that people will adopt more of a distributed model and thus, CI will suffer.  This is pure speculation on my part, and I'm interested to see how things play out.

So my point here is to be mindful of using lots of repositories with Git and its potential negative consequences on CI.

For some more comments on using Git and CI, particularly on larger teams, see this.

Hudson CI Game Plugin

Originally published 17 Apr 2008

redsolo has implemented a version of The Continuous Integration Build Game.  It's a Hudson plugin and is described here .  Way to go redsolo!

Parallel JUnit Ant Task

Originally published 16 Oct 2007

Over time, how do you maintain a maximum 10 minute build? This is an important agile practice for continuous integration.

It's inevitable. As more and more features are added to an application, the code base grows. The build has more to do: more compiling, more tests and metrics to run, more reports to generate. We've done numerous things to keep the build time down. In this entry, I'd like to describe one of them: running standard JUnit tests concurrently.

A colleague of mine came up with this idea. He first did a preliminary search to see if anything already existed.  Pretty much everything he found was intrusive, e.g., requiring extending a specialized TestCase.  He then attempted to combine Ant's Parallel and JUnit tasks. That is, he could simply nest N number of JUnit tasks inside a Parallel task. The problem here was coming up with a good way to divide the tests up into equal parts beforehand. My colleague quickly scrapped this idea in favor of a custom Ant task that wrapped the JUnit task.

The thought was to leverage the JUnit task as much as possible, but fronting it with the ability to run tests in parallel. The conceptual design looks like this:


There are N number of TestExecutors. A TestExecutor runs in a single JVM. Thus, there are N JVMs. Say we have a dedicated, 4 processor integration machine running builds. We may choose to configure our custom task with 5 TestExecutors (we'll add one assuming we're not 100% compute bound). Note: if you're familiar with the JUnit task, you may be wondering how its fork/forkmode attributes fit in. The answer is that they are eliminated in favor of this jvm-count attribute.

There's a single TestProvider. His job is to gather all the tests and hand them out one at a time to whichever TestExecutor is ready. A simple protocol exists to let a TestExecutor tell the TestProvider that he's done and ready for another test.  This should maximize parallelism.

The biggest negative with this approach is that since all tests run in a fixed number of JVMs, there's some loss of isolation here. Class level/global state set from a previous test could affect later tests. We're willing to trade this for increased performance.

The other possible downside is the use of a non-standard Ant task.

Using this approach we've successfully reduced our build time to usable levels.

Bad PMD Rules

Originally published 9 Mar 2007

PMD is an excellent tool for finding potential bugs and improving code quality, but it can generate a lot of false positives. Here's my list of the top 10 rules I turn off immediately, in alphabetical order, with a short comment explaining why. Descriptions for the rules can be found here.
  1. AtLeastOneConstructor: Why? Code is smaller without an empty, no-args contructor.
  2. AvoidInstantiatingObjectsInLoops: This one just seems to generate too many false positives. In addition, for short-lived objects, garbage collection is essentially free.
  3. CallSuperInConstructor: super() is called implicitly and requiring it just adds more lines of code.
  4. JUnitAssertionsShouldIncludeMessage: Most of the methods in Assert already generate good messages. I do include a message for those that don't (assertTrue(), assertFalse(), fail()).
  5. LocalVariableCouldBeFinal: final for variables is usually overkill and it adds line length.
  6. LongVariable: Clarity rules!
  7. MethodArgumentCouldBeFinal: Same as LocalVariableCouldBeFinal.
  8. OnlyOneReturn: I think it's clearer to exit early. It reduces the amount of things to think about below the exit.
  9. PositionLiteralsFirstInComparisons: This really isn't that bad, but myString.equals("x") just reads better than "x".equals(myString).
  10. ShortVariable: There are just too many cases where it's acceptable to have short variables in small methods.
 These two just missed the cut:
  • ShortMethodName: I try to make my names expressive. I've never seen this rule fire.
  • SignatureDeclareThrowsException: This is an okay rule, but for JUnit tests, who cares?

Behavior Driven Development with JUnit 4

Originally published 14 Dec 2006

JUnit 4 makes Behavior Driven Development (BDD) style testing (or specification) easier.

Let's quickly look at an example, the idea stolen from one of David Chelimsky's blog entries about specifying the behavior of a stack. We're focused here on the specification of an empty stack:

// notice the class name specifies the context
public class EmptyStack {
    private Stack stack = null;
    
    @Before
    public void setUp() {
        // set up the context
        stack = new Stack();
    }

// notice the name focuses on the context
    @Test
    public void shouldBeEmpty() {
        assertTrue("not empty", stack.isEmpty());
    }

    @Test(expected=EmptyStackException.class)
    public void shouldComplainOnPeek() {
        stack.peek();
    }
    
    
    // more specification focused methods
}

The point here is that you can get many of the benefits of BDD (a focus on specification rather than testing) using the familiar JUnit framework.

Now, if you're a hardcore BDD'er, then you might complain that you still have to use a test-centric vocabulary. You still need those Test annotations and method calls like assertEquals rather than Dave Astels' preferred shouldEquals calls.

Also, from the legacy side of the fence, you lose the convention of method names starting with testSomething and class names ending with Test. It's sometimes hard to let go of that if you've been writing tests for a long time and it's super clear to spot those test methods and test classes if you're following a naming convention. Furthermore, Ruby on Rails has taught us that convention is a good thing. The test method naming convention doesn't really bother me so much. The @Test annotation makes things clear enough and shouldSomething really gets us focused on specification. However, I haven't let go of ending the class name with Test. Maybe it's more because Ant can pick those tests up easier if there's a standard naming convention.

The last negative I can think of is that of consistency. Having a mix of old style test centric tests and BDD style tests could bother some.

So, let's be honest. What am I really doing? Well, I am incorporating more and more BDD into my work. However, I still shy away from creating new classes. The BDD style really lends itself to many test classes (contexts) per class under test. In the case of Stack, as David Chelimsky points out, you'd also need AlmostEmptyStack, AlmostFullStack, and FullStack classes to fully specify the behavior. I just can't commit myself to writing all those. But I am focusing more on the set up of a context and the specification methods of that context. I just may cheat a little and combine contexts into a single class. So perhaps in the Stack example, I'd combine the empty and almost empty stack contexts under one test class. You know, I'd set up an empty Stack in setUp and for the almost empty context, I'd do a little more set up (push something) in the appropriate shouldSomething methods to make it almost empty.

I also revert back to legacy style testing sometimes. This is usually when I'm modifying an existing, old style test. It's just simpler and faster, and typically the context isn't set up too well for the BDD style.

So my advice is to use the right tool for the job. Use BDD style when that makes things clearer and you want to focus on specification. Use test centric style when that's easier. Try to do a better job of focusing more on specification and less on verification, especially when you're adding new behavior.

A Common Set of Test Refactorings

Originally published  19 Feb 2006

Today, I want to blog about what I think might be my most common set of low level Eclipse refactorings used together when I’m developing unit tests. I use this set to parameterize tests as discussed in Tabular Tests. Lately, I’ve been trying to use the Eclipse refactorings as much as possible, thinking that the automation would be less likely to break things than if I made the changes manually. It’s also interesting to make code changes without doing any typing. The set consists of:

  1. Extract Local Variable (possibly multiple times).
  2. Extract Method
  3. Inline the local variables created in step 1.
So let’s suppose I’m about to start work on a new method. I like to start with a brain dump of most, if not all of the test cases. This is to get me thinking about the desired behavior and also because it clears my head, so that when I start writing the code, I can focus. Once I’m ready, I start with a simple case.

Let’s take the Range.equals() example from the previous blog. My brain dump would have resulted in a table like this:

Min Max Equals?
same same true
different same false
same different false

Yes, I would have jotted down other, exceptional cases for things like null and a non-Range object, but that’s out of scope here. For this blog, I’m just concentrated on comparing Range objects. I start with the positive case and come up with something like:

   private static final int MIN = 1;
   private static final int MAX = 7;

   private Range range = new Range(MIN, MAX);


   public void testEqualsAnotherRange() {
      Range anotherRange = new Range(MIN, MAX);
      assertEquals(range, anotherRange);
   }



To be honest, I probably would have had the MIN and MAX constants hardcoded in there, but at some point, I would have performed Extract Constant to come up with the above.

After I get that test to pass (by returning true), I’m ready to write the next test. I want to pick something simple again (of course in this example they’re all simple). But before I write the next test, I do something that is probably considered cheating by the hardcore test-driven development crowd. I’m not too concerned because I know from experience that I’m going to end up with tabular tests. In fact, the table is specified above!  So here goes the set of refactorings:

I first Extract Local Variable on MIN and MAX:

   public void testEqualsAnotherRange() {
      int testMin = MIN;
      int testMax = MAX;

      Range anotherRange = new Range(testMin, testMax);
      assertEquals(range, anotherRange);
   }



Now I know from my table that I need a boolean for my expected equals. There's no simple Eclipse refactoring so I manually make the change:


   public void testEqualsAnotherRange() {
      int testMin = MIN;
      int testMax = MAX;
      boolean expectedEquals = true;
      Range anotherRange = new Range(testMin, testMax);
      assertEquals(expectedEquals, range.equals(anotherRange));
   }



I highlight the last two lines and perform Extract Method:

   public void testEqualsAnotherRange() {
       int testMin = MIN;
       int testMax = MAX;
       boolean expectedEquals = true;
       testEqualsAnotherRange(testMin, testMax, expectedEquals);
   }

   private void testEqualsAnotherRange(int testMin, int testMax,
           boolean expectedEquals) {
       Range anotherRange = new Range(testMin, testMax);
       assertEquals(expectedEquals, range.equals(anotherRange));
   }



I want to get back to what it looks like in the table, so I Inline each local variable:

   public void testEqualsAnotherRange() {
      testEqualsAnotherRange(MIN, MAX, true);
   }

   private void testEqualsAnotherRange(int testMin, int testMax,
         boolean expectedEquals) {
      Range anotherRange = new Range(testMin, testMax);
      assertEquals(expectedEquals, range.equals(anotherRange));
   }



I just created the first row of the table by applying the set of refactorings that is the topic of this blog. At this point, I usually look at the parameterized method and make sure I’m happy with the signature. I’ve already done most of this when I extracted the method, but Eclipse doesn’t let me make the method static, if that’s appropriate (in this case it’s not). Eclipse also lists all the checked exceptions, which for tests, I’d rather generalize to Exception.  This doesn't apply in this case either.

During this set of refactorings, I would have ran the tests after a change or two just to make sure I didn’t break anything. I also do that because it feels good to get that green bar.

This is a pretty trivial example. Most of the time much more is going on in the parameterized method, but this still shows the removal of duplication.

I could have made these changes manually, but like I mentioned above, I like the automation; I’m less likely to break things.

Okay, I can write the next test now:
   public void testEqualsAnotherRange() {
      testEqualsAnotherRange(MIN,     MAX, true);
      testEqualsAnotherRange(MIN - 1, MAX, false);
   }


I’ve got the red bar. Time to make it green.