statementofwork.org
New · Commercial methodology

Reference Class Forecasting (RCF)
for short term services via SoW

Build your own knowledge portal of scope and statement of work data. Interrogate past delivery against new requirements. De-duplicate effort and accurately forecast future costs against historical data. Works with the new machine readable statement of work standard.

METHODBuilt on four decades of forecasting researchWorks with all four engagement types
91.5%
of 16,000+ projects miss on cost, time, or both
73%
average cost overrun across 5,392 IT projects
447%
average overrun for the worst 18% of IT projects
0.5%
deliver on cost, time, and benefits together
Sources: Flyvbjerg and Gardner, How Big Things Get Done (2023); Flyvbjerg, Budzier, Lee, Keil, Lunn and Bester, Journal of Management Information Systems 39(3), 2022.
The problem

Why SoW estimates run over.

A SoW is estimated by the people closest to the work, and usually by the people who want it approved. Behavioural economics calls this the inside view: build the forecast from the specifics of the case in front of you, and it inherits every assumption, omission, and incentive in the room. Kahneman and Tversky named the result the planning fallacy in 1979. Four decades of outcome data show what it costs.

The error has two engines, and most organisations argue about which one they have. The evidence says it does not matter.

Honest error
Optimism bias
Delivery teams believe this engagement will be the one that runs clean: no rework, no slipped dependencies, no supplier churn. Nobody is lying. The estimate is still wrong, and always wrong in the same direction.
Deliberate shaping
Strategic misrepresentation
Estimates shaped to win approval: scope understated to fit a budget, timelines squeezed to hit a funding window, services framed to pass a headcount freeze. Flyvbjerg's study of 258 public projects concluded the underestimation "cannot be explained by error".
Using distributional information from similar ventures is the single most important piece of advice regarding how to increase accuracy in forecasting.
Daniel Kahneman, Nobel Memorial Prize in Economic Sciences, Thinking, Fast and Slow (2011)

The correction does not care which engine you have. It replaces the estimate's promise with the record of comparable completed work, so honest optimism and deliberate shaping are corrected in the same move.

The method

Reference class forecasting in three steps.

First applied by Bent Flyvbjerg for the UK Department for Transport in 2004, reference class forecasting corrects a forecast using the outcome distribution of comparable completed projects. It is now mandatory in UK and Danish government appraisal.

01
Choose the reference class
Past, similar, completed engagements. Broad enough to be statistically meaningful, narrow enough to be genuinely comparable.
02
Establish the distribution
Chart the spread of outcomes for the class: cost overrun, schedule slip, scope change. Include the tail, not just the average.
03
Position and uplift
Place the new engagement in the distribution and apply the uplift that matches your risk appetite. Accept a 20% chance of overrun, apply the 80th percentile.
The gap

The standard turns completed SoWs into usable data.

Reference class forecasting needs comparable, completed records. Traditional SoWs cannot supply them: every document is structured differently, scope lives in prose, pricing and dates sit in inconsistent formats, and outcomes are never written down. There is no distribution to consult, so every new SoW starts from zero.

The standard removes that constraint. Eleven sections, 51 elements, machine readable by design. Every SoW written on the standard captures the same fields in the same structure, and every completed engagement becomes a comparable data point.

COVER
Engagement Type
The reference class key: Time and Materials, Fixed Price, Outcome-Based, or Managed Services.
03
Time
Planned dates against actual delivery: the schedule slip distribution.
04
Pricing
Contracted price against final cost: the cost overrun distribution.
05
Scope
Original scope against change orders: the scope creep distribution.
06
People
Named roles and suppliers: the variables that explain variance between classes.
07
Performance
Service levels and acceptance criteria against results: the benefit shortfall distribution.
The approach

How to build yours.

Five stages. The first two are operating discipline, the last three are arithmetic. You do not need a data science team to start; you need consistent records and the patience to let them accumulate.

01
Standardise the record
Adopt the standard for every new SoW. From signature onwards, each engagement captures the same 51 elements: engagement type, scope, pricing, dates, people, and performance measures, in the same structure every time.
The standard at work
Sections 1 to 11 define the fields. The free Legal Standard is enough to begin.
02
Close the loop at completion
When an engagement ends, record five outcomes against the original SoW: final cost, final dates, change order count and value, acceptance result, and supplier. One page, thirty minutes, written at closure while the facts are fresh.
The standard at work
The 51 elements are the baseline the closure record measures against.
03
Assemble reference classes
Group completed engagements by engagement type first, then by scope category and size. Comparable beats big: a class of thirty genuinely similar engagements outperforms a class of three hundred mixed ones.
The standard at work
Engagement Type on the Cover Page is the first classifier. Tier 3, The Scope Reference Class Forecast, sharpens the second.
04
Establish the distributions
For each class, chart cost overrun, schedule slip, and scope change across completed engagements. Report the spread and the tail, not just the average: in IT projects, the worst sixth costs more than the rest combined.
The standard at work
Machine readable structure makes this a spreadsheet exercise, not a project.
05
Uplift before signature
Position every new SoW in its reference class before it is signed. Price contingency at the percentile that matches your risk appetite, and make the uplift visible in the pricing section rather than hidden in padded estimates.
The standard at work
Certified Managing trains the closure discipline. Certified Reviewing trains the challenge questions.
What a fat-tailed class looks like
Distribution of cost overrun across completed engagements in one reference class
P50 P80 uplift Most engagements land near the estimate The tail decides the portfolio outcome In IT: 18% of projects, 447% average overrun
Scroll right for the tail
Illustrative shape. Tail figures: Flyvbjerg, Budzier, Lee, Keil, Lunn and Bester, Journal of Management Information Systems 39(3), 2022, n=5,392 IT projects.
The precedent

Governments already publish their uplifts.

Since 2003, HM Treasury's Green Book has required every UK business case to apply a published optimism bias uplift until better local evidence exists. The categories and upper bound uplifts are public, and they fall as the project demonstrates better information.

A company running the approach on this page is doing the same thing at enterprise level, with one advantage: its uplifts come from its own completed SoWs, not a national average.

The payoff

What a reference class gives you.

Budgets that hold
Contingency is priced from your own outcome distribution, not negotiated from hope. As in the Green Book, the uplift falls as your evidence improves.
Early warning on fat tails
Classes with heavy tails are visible before signature. You know which engagement types can blow up, and you govern them differently from day one.
One fix for both biases
The forecast is corrected whether the estimate was honestly optimistic or shaped to win approval. No blame conversation required.
Supplier evidence
Promised against delivered, by supplier, by class. Renewal and selection conversations move from opinion to record.
A defensible number
No theoretical optimum required. The reference class is the counterfactual: what work like this actually costs when your company buys it.
A compounding asset
Every completed SoW sharpens the next forecast. A competitor cannot buy your distribution. It only accrues to companies that keep the record.
Questions

Frequently asked.

Is this just benchmarking?
No. Benchmarks compare your prices with other companies' prices. A reference class forecast compares your new estimate with your own completed outcomes. The reference class is internal, the comparison is estimate against actual, and the output is a concrete uplift you can price into the SoW before signature.
How many completed SoWs do we need?
Aim for 20 to 30 per class before leaning on the numbers, and keep classes comparable rather than big. Below that threshold, borrow the Green Book pattern: apply a published default uplift and replace it as your own record grows.
Does this replace estimating?
No. The bottom-up estimate remains the anchor for scope and price. The forecast corrects it for the bias that four decades of outcome evidence says is in it. Estimate first, then position the estimate in the reference class and apply the uplift.
Our old SoWs are not on the standard. Can we still start?
Yes. Backfill closure records for engagements completed in the last two years: engagement type, contracted price, final cost, planned and actual dates, change orders. Most of it is recoverable from any SoW and its invoices, and it seeds your first classes while new SoWs accumulate on the standard.
What do we need to buy?
Nothing to start. The Legal Standard is free with an account and defines the full data structure. The Tier 3 Scope Reference Class Forecast pack sharpens reference classes, and the Certified Managing track trains the closure discipline, but neither is a precondition.
Team reviewing engagement records
Start now
Start with your next SoW.

Write it on the standard, record the outcome at closure, and your first reference class follows. The standard is free to download.