06. Numbers

Most software projects fail, and the report behind it

Success required the original date, budget and every original feature, so learning something and changing course counts as failure.

Teams applying this historical lesson to current delivery work can also compare tools for employee time tracking software, keeping operational records separate from the code and its documentation.

most-projects-fail.src
sackman1968.csv n=12 ratio=28:1
# extreme against extreme

1994 onwardsThe report

A consultancy published a study stating that a large majority of software projects were either cancelled or delivered late, over budget or incomplete, with a small minority fully successful. The figures were restated at intervals for two decades and are among the most quoted numbers in the field.

They appear in business cases, in conference talks, and in the introduction of an enormous number of papers, usually as an uncontroversial statement of the problem.

The definition of success

A project counted as successful only if it was delivered on the original date, within the original budget, and with all originally specified features.

Consider what that classifies. A team that discovers halfway through that a feature is unnecessary and drops it has failed. A team that learns something and changes course has failed. A team that padded its estimate by half and delivered exactly what was asked, whether or not it was useful, has succeeded.

What the definition rewards

Inaccurate estimates. A project with a generous estimate meets it; one with an accurate estimate is at the mercy of ordinary variance and misses about half the time.

Analyses published in the engineering literature made exactly this point: the measure rewards padding and penalises accuracy, which means the reported failure rate is partly a measure of how honestly the organisations surveyed estimated.

The methodology

Not published. The sample, the selection of respondents, the questions asked and the weighting have never been made available for examination, and the report is a commercial product sold by the organisation that produces it.

Attempts by researchers to reconcile the figures with other data on cost overruns did not succeed, and the published critiques concluded that the numbers should not be used.

The selection problem

There is also a plain sampling difficulty. Organisations volunteer cases, and the cases people volunteer are not a random draw from all projects. Failures are more memorable, more discussed, and more likely to be submitted to something calling itself a study of failure.

Nothing about that is dishonest. It is what happens to any voluntary sample, and it is why the sampling method is the first thing to ask about.

What is probably true anyway

Large software projects do overrun, frequently and substantially. Independent studies of cost overruns, using published methods, find sizeable average overruns with enormous variance.

So the underlying claim survives while the specific proportions do not, which is the same conclusion as every other entry in this part and is worth stating each time rather than leaving the reader to conclude that nothing is known.

Why it is quoted so much

Because it opens an argument efficiently. A speaker who wants to propose a method needs a problem, and a memorable proportion of failures is the fastest available one.

The figure is therefore repeated overwhelmingly by people with something to sell against it, which is the same interest pattern as the curve two entries earlier and is visible without knowing anything about the methodology.

Why the critique did not displace it

The critiques are in journals and the figure is in slides. One circulates and the other does not, which is the same asymmetry as the diagram in the entry on the waterfall paper.

There is also no replacement with the same properties. A speaker needs a single memorable proportion, and the honest answer, that overruns are common and variable and the proportions depend on definitions, does not fit on a slide.

What an organisation can measure about itself

How long from the decision to start something to it being in use. How often work started is abandoned. How much of what was built is still in use two years later.

All three are recoverable from records most organisations already hold, all three are meaningful without a benchmark, and none of them can be improved by padding an estimate.

Counted as successful only if all three holdon the original datewithin the original budgetwith every originally listed featureso a project that learned something and changed scope deliberately is a failure,and one that padded its estimate and delivered the wrong thing on time is a successthe method behind the figures has never been published
FigureThe three conditions all required for success, and what that classification does to a project that learned something.

The honest framing of the underlying problem

Building something substantial for people who have not seen it before is difficult, and difficulty shows up as time and money. That is not a scandal and does not need a proportion attached.

Presenting it as a scandal has a cost: it invites the search for a method that removes the difficulty, and fifty years of that search is what the entry on the software crisis describes.

What the report did achieve

Worth conceding at the end. It put the difficulty of large software projects in front of people who commission them, in a form they would read, at a time when that was not widely acknowledged.

The figures do not survive examination and the attention was probably useful, which is an uncomfortable combination and a common one.

What to ask instead

Whether a project delivered something people used, and whether the organisation would do it again. Both are answerable within an organisation, neither requires a benchmark, and neither is defeated by a deliberate change of scope.

The reason those are rarely used is that they cannot be compared across companies, which is precisely the property that makes the published figure attractive and unreliable.

What we cannot verify

The reports are commercial and their method is not public, so we cannot examine what we are criticising, and say so. The published critiques are in the peer-reviewed engineering literature and can be read. Independent overrun studies use different definitions again and are not directly comparable with either.

In short

  1. Success required the original date, the original budget and every original feature.
  2. Dropping an unnecessary feature deliberately counts as failure by that definition.
  3. The measure rewards padded estimates and penalises accurate ones.
  4. The methodology has never been published and the report is a commercial product.
  5. Volunteered cases are not a random sample, and failures are volunteered more.
  6. Large projects do overrun; the specific proportions do not survive examination.

also in Numbers

Next.

further context

For a primary or institutional reference, see the NIST Information Technology Laboratory.

Every claim here carries the source it came from.

The source and its year sit beside the sentence they support. A secondary account is marked as one, and where the record is unclear the entry says so rather than choosing the better story.