The Outside View

The details explain a project. Its reference class predicts it.

Title card reading 'The Outside View' in green type on a cream circle, surrounded by geometric shapes in muted blue, sage, terracotta, mustard, and charcoal

A team estimates a migration at two weeks because the implementation looks straightforward. The schema is understood, the code path is narrow, and the owner has done this kind of work before. Each detail supports the estimate, even though the team's last five migrations took between five and eight weeks.

Both descriptions can be accurate. One is the inside view, built from the particulars of the plan in front of you. The other is the outside view, built from the outcomes of comparable work. When they disagree, the history usually deserves more weight than it gets.

A plan explains more than it predicts

The inside view is attractive because it is specific. A plan names the tasks, owners, dependencies, and expected sequence. It gives every week a reason to exist.

That specificity can create confidence without improving the forecast. The plan describes how the work could proceed if its assumptions hold. It says less about how often assumptions like these have held for this team, in this system, with this kind of dependency.

Past projects include the costs that plans routinely omit: review queues, unclear ownership, customer feedback, integration surprises, and decisions that have to be reopened. Throughput makes a related point about measuring completed flow instead of planned effort. The outside view applies that discipline before the work begins.

Choose the reference class before the estimate

An outside view starts by choosing a reference class: a set of completed efforts similar enough to inform the current one. The choice needs to happen before the team settles on a date. Otherwise it is easy to search history until one example supports the answer everyone already prefers.

Useful comparisons are based on the source of difficulty, not just the label on the project. A billing migration may belong with other changes that crossed customer data, finance review, and third-party systems. It may have little in common with an internal database migration that shared its technical name but none of its coordination cost.

Ask four questions:

  1. Boundary: which systems, teams, or customers did the comparable work cross?
  2. Uncertainty: which important facts were still unknown when it started?
  3. Commitment: who had to review, approve, adopt, or communicate the result?
  4. Outcome: how long did it take, what changed in scope, and what had to be redone?

The goal is not to find a perfect twin. It is to find several completed efforts exposed to the same kinds of friction.

Let disagreement reveal the hidden assumption

Suppose the plan says two weeks and the reference class says six. Averaging them into four weeks hides the useful part of the disagreement.

Instead, ask what is materially different this time. Perhaps a reusable migration tool removes work that every prior project had to do manually. Perhaps the same external review still exists, which means a faster implementation will wait in the same queue. A claimed difference should point to an observable change in the work, not confidence in the people doing it.

This turns the outside view into a diagnostic tool. The base rate identifies where the plan is unusually optimistic. The team can then name the mechanism that earns a different outcome and watch whether that mechanism appears.

An agent needs outcomes, not only artifacts

An agent can read a detailed specification and produce an even more detailed plan. That makes the inside view cheaper to generate, but it does not make it more accurate.

To take the outside view, the agent needs completed outcomes connected to the conditions that produced them. A closed ticket alone is weak evidence. It needs the original expectation, the actual elapsed time, the dependencies encountered, the scope changes, and the decision that defined success.

This is where State Estimate matters. A pile of project records is not yet a coherent account of what happened. The records need current status, identity across systems, and enough provenance to distinguish a finished project from an abandoned one.

Without that structure, an agent will retrieve examples that sound similar. With it, the agent can compare work that behaved similarly.

Keep the forecast and the explanation separate

A useful planning review should preserve both views. The reference class supplies the starting forecast. The plan explains which concrete differences may move the result away from that base rate.

Record those differences as assumptions. If the reusable tool was supposed to remove two weeks, check whether it did. If an approval queue was expected to be shorter, record when the approval arrived. The comparison then improves the current decision and the next reference class.

This also prevents Decision Debt from hiding inside estimates. When the date changes, the team can see whether the world changed, an assumption failed, or the original forecast ignored evidence that was already available.

Before accepting a confident plan, ask one question: what happened the last five times we tried work with the same sources of difficulty?

Frequently asked questions

Is the outside view just using an average? No. The important step is choosing a defensible reference class. A precise average from irrelevant projects is less useful than a range from a small set exposed to the same constraints.

What if the project is genuinely new? Compare its risky parts separately. A new product may still reuse familiar review, migration, integration, or customer-adoption patterns. Where no comparison exists, keep the uncertainty visible instead of replacing it with precision.

Should the outside view replace a detailed plan? No. The reference class supplies a forecast, while the plan explains execution and identifies reasons the outcome may differ. Keeping both makes their disagreement inspectable.

How many past projects are enough? There is no universal minimum. Start with several relevant outcomes, show the range, and state why they are comparable. A small honest reference class is better than a large convenient one.

GET TLDR FROM:
← Back to Blog