Showing posts with label Testing Myths. Show all posts
Showing posts with label Testing Myths. Show all posts

Wednesday, March 13, 2013

Novopay - the tale of a compelling event in the New Zealand IT industry

Open almost any book on testing, and it will start with a "cautionary tale" about software testing, where something was released into production, and it went wrong with devastating results.  They are the ghoulish "tales around the campfire" of our IT industry.

One problem I have found within the New Zealand industry is that as testers we love to focus on anything that goes "bang" and crashes.  In 2011, I was in one a talk about software testing to a group of new developers as part of the Summer of Tech, where our introductory speaker spoke of an Ariane 5 rocket where the parts had been individually tested, but put together into a new rocket.  As you can imagine, this rocket barely cleared the launchpad before it exploded. Another cautionary tale of "you didn't test enough".


What's interested is what happened next.  One of the audience interrupted with "but we're creating applications ... not rockets.  Our stuff is hardly critical like that".

To me, this reinforced how vital it is to have stories and tales of failure which are relevant to the audience.  I thought the tale was a wonderful one, but looking around the room, I saw it failed, because the audience simply did not believe it applied to them.

New Zealand is a small and very pragmatic country.  A lot of processes here are still fairly manual compared to my home country of England, and overall that's not a bad thing.  Here in New Zealand if something isn't broken, New Zealand attitude is "why try and fix it", so,

  • In the UK for trains we have automated ticket vending machines (usually vandalised), automated turnstiles to get onto platforms etc.  We used to get told our ticket prices would have to go up above inflation to cover this automation and it's continual repair.  In New Zealand they just have an old fashioned conductor on the train who checks and clips tickets, and can sell you one for cash if you need one (without charging a fee).  Simple, but y'know it works.
  • In the UK there were plans to build a so called "chip and bin" system for collecting rubbish.  The UK rubbish collection vehicles would scan your bin, which would then be weighed as it was disposed of.  Back at the main office this data would be collected, and an itemised bill created (your bill would have to be more to cover all this new technology which would have to be developed and maintained).  Again being super pragmatic, New Zealand councils just sell you council marked rubbish bags which have a charge on them for collection.  Only rubbish in council bags is disposed of.  The more you throw away, the more bags you need.  Elegantly simple yes?


An unfortunate by-product of this pragmatic way of doing things is the attitude of "she'll be right", which means if something goes wrong, don't worry, we'll be able to fix it.  This means on the whole New Zealanders can be a little bit more risk takers than their European or American counterparts.

In a testing consultancy I previously worked for, my test manager would talk about how New Zealand would have to face an inevitable "compelling event" for software testing.  Most other countries have had one, but so far although there had been failed there hadn't been one in New Zealand.

A "compelling event" is something that forces you to take action - a wake up call the the importance of the value of testing.  A compelling event isn't just software going bad when it hits production, it's software causing a very large and public pain that it unnerves the local IT industry as a motivation to not repeat that mistake, often with a gasp of "that so easily could have been us".  It's not a rocket blowing up half way around the world, it's software that fails but feels too close for comfort compared to what you're doing right now.

It's a compelling event because once it's happened it compels you to take a good hard look at your strategies around testing and quality and ask "are we (and not just testers here) doing enough?".

Sadly, we've finally had our big compelling event here - the Novopay saga.  Novopay was an online system for managing the payment of teacher and school staff salaries.  But sadly it has gone into production to go horribly wrong, with payments to these people missing in the system.  There have been schools where teachers are missing months of pay (and teachers aren't exactly in the most affluence of careers).  More than just being "missed payments", there have also been erroneous payments gone out from schools - some teachers who've never worked for a school, are receiving payments from that school - and schools are having to bring in extra staff to go through the Novopay payments with a fine tooth comb to work out what's payments are good, what payments are errors and what payments are missing.

The media have been in a frenzy over it, hounding senior staff at the Ministry Of Education (who are behind Novopay) saying they have found a report of "200 known bugs" in the production system when it went live.  It is terrifying to me as a senior tester is how bad that sounds when the media put it like that.

Of course when I worked on an avionic computer, it'd surprise people to know it flew under the strictest of safety conditions, but still there were 1,500 known bugs in the software.  The important message though isn't the bug count, but the severity.  But this is a very difficult dialogue to have with a set of press hounding for a "big story" and "potential coverup".

In an unusual move, the Ministry of Education has released the test plans for the Novopay project into the public domain.  I took a look through, and they show a fairly thorough planned approach was taken to testing, not the slap-dash "rush into production" that many in the press are claiming.  That was pretty unnerving - most of our "testing horror stories" like the Ariane rocket involve people simply not testing, thinking it would be "alright mate".

The documentation they have also shows a decent approach to several forms of functional testing, indeed going beyond what I would have initially thought to do,

http://www.minedu.govt.nz/theMinistry/NovopayProject/NovopayTestPlans.aspx

However for all it's success in testing, something has gone very wrong in production, there can be no doubt about that.  And it is causing real grief to people whose life it should be easing.

What went wrong?  Each of us will have our opinions about this, and no doubt all of us will be right to some extent, whilst at the same time knowing in our hearts how easily something like this could happen to us.

Testing is about identifying and helping to remove risks, but it does not guarantee a bug-free product.  We can test in the many different ways we expect our software to be used, we can check for all the problems we can imagine could happen.  But that will never cover everything, there will always be issues beyond our imagining.

But the imagination of any individual has it's limits.  We have to be careful not to find ourselves screaming "inconceivable" at every bug found in production.  In my opinion, this is why testing and quality is not just a "test team" ticket.  We need to be engaged with developers, business analysts, market managers, usability experts, end users to work out a whole scope of things to test (from areas of functionality to ways of using) that is beyond the imagination we as testers can summon when looking at requirements.  But by then our scope of testing has probably increased by several orders of magnitude, and we don't have unlimited time.

That's when we need to do some risk based analysis, and cover as much of it as possible, touching on all the "most likely" cases, then touching on samples in other areas.  If you find problems, then odds are there'll be others, so keep looking in that area.

For myself then , if there is a compelling lesson to be had from Novopay it is that we need to draw up our test plans as we always have done, and then ask of others "what could I have missed".

Sunday, December 23, 2012

Project Christmas: The Elf Who Learned How To Test


I have had a busy year, no doubt about it, with a lot published in magazines like Testing Planet, Testing Circus and Tea Time With Testers.  There has also been my book, The Software Minefield.

I've recently put together another much shorter book, which is free to download called The Elf Who Learned How To Test, of which I'm particularly proud (great I now sound like Q from James Bond).

The idea started a long time ago with a conversation with Rosie Sherry about the idea of "Imagine there's no testing" back in August.  And that's just what I did with the leap of imagination required.  I imagined Santa's Workshop where no-one tested, and children received substandard presents, and an elf who discovered how testing could add value.

I've been told by others in the testing community I'm a great storyteller.  Indeed in The Software Minefield, I mentioned I'm always telling war stories or parables.  What pleases me about The Elf Who Learned How To Test is that it's a tale not just for testers, but for their children as well.

I think as testers we sometimes are great at joining together as a community and sharing our stories.  But perhaps where we fail is sharing what we do, not just with others who aren't testing, but especially our children.  And my book works as essentially as a Testing Fairytale which can be shared with children, with some thinking activities at the back which I feel the children are likely to score as well in as the adults!

This year has seen me peeling back the mystique around testing for my 14 year old son, who has come in to see what we do.  I keep trying to talk to him about what I do and why I do it.  I try and develop him a sense of analytical thinking, especially in our common area-of-interest which is history.  Unsurprisingly he did well with the activities in the back (which do not have any "right answers", but as more about seeing how you can expand on the story, and how you interpret some things which are not said).

You can download the book here,

https://leanpub.com/TheElfWhoLearnedHowToTest

In addition, I did a video of me reading it for YouTube, but it turned out too big, so I've decided to put it up as a podcast, which can be accessed below,

http://testsheepnz.podbean.com/2012/12/22/the-elf-who-learned-how-to-test/


This is all aimed at encouraging donations to a very worthwhile charity, Starship, which supports sick New Zealand children and their families.  I have been helping to support this charity through work, and if you'd like to support them as well, please give a one-off-donation below,

https://www.starship.org.nz/foundation/how-i-can-help/making-a-donation-now/


Wednesday, October 12, 2011

Those darn test estimates …



How long is a piece of string?  I'm tempted to be a wise-ass and say that to project managers when they ask me “how long will it take to test my project?”.

That's actually unfair, experience gives us as testers an idea, based on similar projects, of how long it'll take to test.  But one of the problems is, testing is an activity which has a complex relationship with other factors in a project – we can keep testing and testing, but if development don't start fixing some bugs, we're going to be here forever!

So yes, we can look through the designed features and estimate how long it will take to script, and how long to execute those scripts.

But how long until the product is finished testing?

How long is a piece of string?

Having worked in a test consultancy, there is no doubt about the importance of estimates.  They need, when a project manager looks at them, to be attractive but also realistic, with some contingency.

What we introduced was a list of estimates for testing tasks for “best case”, “probable case” and “worst case”.


                BEST   PROB   WORST
Test Plan         1 2 4
Test Conditions   2      3      5
Test Scripting    4      7     10
Pre-Testing       2      4      6
UAT Execution     4      5      8
Retesting         2      5     10

This gives the test manager some leeway, usually if things go okay it should follow the Probable estimates.  If they book the Best case estimates, be very worried.

What I'm finding is my project managers are taking my estimates, adding the Probable case figures together and times it by an hourly rate to get a budget. I don't know why I'm so surprised ... makes sense, but I'm used to working against time and not $$$.

Unfortunately at the end of the last two projects we've been considerably over that budget.  There seems to be several factors at play which determine which of those estimate paths our testing is going to follow, and it's important to understand and recognise them.

Software Delivered Late

You book in a test contractor to help you test for 6 weeks.  They arrive on week 24 to start analysis and scripting, with some test execution happening in week 26 for 4 weeks.

Then your chief developer tells you there's going to be a 2 week delay getting the build together, it won't be available until week 30 now.

You've made a commitment to your test contractor, and are so obliged to pay him, and possibly find them other work.  If you can't get them to assist elsewhere, then by week 30, you're 4 weeks in and not testing yet.  You've blown over half your budget, and Lord help you if there's any more delays!

We all know developers can often deliver late.  Late software is going to burn up budget.  You need to work with your Project Manager to make them aware of their duty to get software to you as scheduled in order for your budget to be met.

Software Delivered Is Of Poor Quality

Kind of the flip side of late delivered software.  Your vendor has promised that the software delivered has been unit and system tested, and no bugs were found.

You wrote your test plan for acceptance testing, expecting the software to have been extensively tested beforehand, with a certain level of quality.  Your project manager and you are expecting what's delivered to be a candidate for release.  You turn it on, and immediately notice a dozen problems, not able to finish basic use cases.

The developers under duress delivered what they had available to schedule instead of flagging any delays.  Little if any testing has happened, and basic bugs are being discovered only now.  Testing 101 says "more bugs = more fixes = more builds = more retesting".

One thing I try and do with vendors is ask for a release note and end report for testing, detailing what defects were found and what were fixed.  This is a bit of a game of bluff.  If I receive an end report which says “everything was tested, and no defects were raised” I get suspicious.  Very suspicious.

I've also had vendors on conference calls inform me “we're running a build up now, you'll have the install delivered in an hour”.  I pull my project manager to one side when this happens and warn them that maybe that will mean no testing whatsoever has been done …

The Delivery Chain

If you have developers on-site who you can give defects to, they fix, build, test it's possible to get a build almost every day.

If they're off site, only receive defects daily, have to courier builds, you'll be hard pressed to get a build weekly.

If you have two weeks to test, and have a daily build, you'll have 10 opportunities to get it right.

If you have weekly builds it's not likely to happen.  Your second build will have to be perfect – and it usually takes about 3-4 even with an initially high quality piece of software (there are always tweaks needed).

Time erosion

It's so easy to happen.  You have,

  • a daily half hour team meeting
  • a one hour weekly project progress meeting
  • a one hour weekly project technical meeting
  • a daily 15 minute end-of-day defect wrap up meeting
  • each day you spend half an hour writing a progress report for the concerned business owner


Oh you're giggling there, but we've all been there.  Did you add it all up?  Yes, you're losing about a day a week.  Look at your estimates, did you plan on there being so much leakage?

I'm finding we're increasingly working on projects where there are a large amount of meetings to keep track of progress.  This needn't be a bad thing, and small meetings daily can help set the direction and key priorities of the day/week.  But it's easy for reporting to actually delay any progress being made, and become a sizable and unknown overhead in itself.

And some projects need test management – how do you budget for that?  It's not a solid “task” again more an ongoing overhead.

Requirements?

Requirements?  We didn't have time to write down everything we asked for!

Due to constraints a project has been only broadly defined, but you're required to perform specific testing against it, to a limited timeframe.  Oh and business analysts are too busy to answer your questions, so just get on with it and you know, test!

This is a nightmare position to be in.  You press a button, a message is displayed.  But you have no idea if it's the right message or not.  There are some things you can do – you can check the application didn't die when you pressed the button, and the message made sense in the context of the button.

But if you have vague requirements, you can only vaguely test.  Such projects really feel like they're setting up the test project to fail.  And take the blame.

Another variant of this is you raise about 10 defects against requirements, there's a review, and a business analyst says “oh yes these aren't defects, I asked for these changes by phone from our vendor”.  If things aren't documented, how can anyone keep track of these changes?




Take it easy.  Take it nice and slow.  That's no way to go.  Does your PM know?

Thankfully there's usually a place for these factors in a test plan under risks and assumptions.  But I can't emphasise enough to you the importance to talking them over time and again with your project managers before you embark on any test estimates, so they can understand and more effectively evaluate the risks and the impact to budget.

Friday, October 7, 2011

Testing's Men in Black


At a time when the world is watching the Rugby World Cup in awe of a certain set of men in black, it's interesting to see how this story has been doing the rounds on the internet – thank's to the Testing Club's Rob Lambert

http://www.t3.org/tangledwebs/07/tw0706.html

It's a no doubt apocryphal urban legend about IBM - but a lot of fun to read anyway!  In the 1960s the world of computing had a different emphasis – programming was a much slower business, and there was one shot at delivering software, no patching, it had to be right on release.

IBM supposedly found that programmers who wrote the code were blind to any faults in their software when it came to testing.  Some people though showed a natural aptitude, and thus one of the world's first test teams - “the Black Team” - came into being.

The Black Team were made up of the best-of-the-best when it came to breaking software.  They became a kind of bogey-man to terrify young developers, able to break any software they came across.  This tale here is just a brilliant parable of the supposed lengths they'd go to in order to test software ...

http://www.penzba.co.uk/GreybeardStories/TheBlackTeam.html

Tales of the Black Team go further, telling how team members started to form an identity together, wearing all-black to the white collar IBM offices.  Some even growing Dali-esque moustaches they would twirl sinisterly as they tested.


I'm very dubious – but I absolutely love how software testing, which feels sometimes like a very recent discipline (in New Zealand sometimes it feels not quite respected as a profession at all some days), has managed to pick up this urban legend.

But it also takes me back to one of my first posts here, on “what is software testing”.  The tale of the Black Team is all about a team who go out of their way to break things.  The story where they work out the resonant frequency of a large tape reader so it rocks itself over whilst reading a file is bang-on-the-money for a lot of people who see testers as people who just go out of their way to destroy.


Today in a management meeting it came out that the business owners see testing as a problem.  “The project was all going well until testing got involved”.  As if testers are responsible for the defects they encountered.  I think the reality is much closer to “we managed to delude ourselves that everything was fine until testing gave us a wake up call”.  If testing is done well there's no hiding the truth of where a project lies.

But no- testing is not about breaking things.  It's about proving quality.  And that can be a bogey man of it's own to a complacent project.

Friday, July 22, 2011

Testing and the Cassandra syndrome



What a stressful week!

We’re now a couple of days from formal User Acceptance Testing for our new product, and we have a limited 2 week window.

In an ideal world we’d have had an early version of the software to run preliminary tests on.  But there was only one machine available, and the priority was both to give this first to marketing and then to training because “testing wasn’t due to formally begin until a week later”.

This made for quite a frustrating experience, as I tried to explain to management the importance of testing getting an early look at it.  But I wasn’t really listened too – I was told as our window was so tight, it would “just have to work first time”.

It actually made me quite angry – this was project management by desperation, and flew in the face of everything I knew.  There would be defects I said, as in my experience projects always had defects.  It’s one of the fundamentals of testing “test early” but management were insisting on a rigid waterfall interpretation of “test at the end”.

My feeling of this is much like Sun Tzu who said that,


“a victorious army first obtains the conditions for victory, then seeks to do battle.”

Basically I interpret this as before you formally test, you test informally anyway to know the overall quality of your product and when if it’s ready.

In software there’s often a “can do” attitude of we can do anything.  Unfortunately this sometimes becomes as the above comment “we’re so stretched for time, it has to work first time”.  Everyone is optimistic it can be done.

This is where the Cassandra curse comes in.  Cassandra was an Oracle blessed with the power of divination.  But after a fall out with Apollo, she was cursed that her prophesies would never be believed.  And so she could see the future but was powerless to prevent the disaster she knew was coming.

This is something I think too many testers feel.  When many managers and developers feel “hey it’ll all work first time”, testers are all too experienced that often it doesn’t.  They know this because it’s their job to deal with things when they fail, and they’re expert at working the problems.  In all my development experience, I’ve only ever had one thing work first time (ironically it was also the most complicated thing I wrote, a search algorithm).

Problem is it’s hugely demotivating to be ignored or disbelieved when you try and warn a project there are problem ahead.



We finally managed our preliminary tests today – a couple of big issues, but a whole host of mediums one as well.  Perhaps too many to address in the time left.

But hey, that’s why they call me Cassandra ….