Intent-as-Source

The one question

What is this requirement for — and if we satisfy the requirement but miss the "for," how would we ever know?

source

40 of 41 outcomes missed

The one loop

The mechanism

Intent
Compile
Verify
Lock
Execute
  • Human judgment intent and escalation
  • Automated / AI compile and verify
  • Both split by the gate

Three domains, one shape

The failure shape

Language models

SPEC MET — OUTCOME MISSED

30.4% reward-hacking rate

  • a suite of AI research-engineering tasks
  • Telling the model not to reward-hack made almost no difference
source

95% even when told to do it as intended

  • the highest measured
  • on the most-hacked task
source

Emergency care

SPEC MET — OUTCOME MISSED
Two decades earlier, English emergency departments worked under a target: no patient waits more than four hours in A&E. Bevan and Hood documented hospitals meeting it by holding patients in ambulances outside the department — the four-hour clock had not yet started — and they point to a fatal case.
source

Aviation

SPEC MET — OUTCOME MISSED

346 people died

  • met every requirement it was certified to meet
source

Forty failures, one success

The wall of failures

skip to the worked example →
Read this catalog for shape, not frequency. Forty failures and one success were chosen to show the forms this pattern takes across domains — including one sighting of the opposite shape. A curated collection cannot say how often specifications succeed or fail, and it cannot prove Intent-as-Source's thesis by itself.

source

40 of 41 outcomes missed, one apart
  • 01 · Boeing 737 MAX / MCAS (2018–2019) — component-vs-system, automation-over-trust — a single sensor with no cross-check, certified as a Speed Trim add-on.

    SPEC MET — OUTCOME MISSED source
  • 02 · Mars Climate Orbiter (1999) — interface-nonconformance — a correct interface spec, simply not followed.

    SPEC MET — OUTCOME MISSED source
  • 03 · Ariane 5 Flight 501 (1996) — component-vs-system, under-specification — code correct for Ariane 4, fatal 40 seconds into Ariane 5.

    SPEC MET — OUTCOME MISSED source
  • 04 · Space Shuttle Columbia (2003) — normalization-of-deviance — foam strikes reclassified from near-misses to routine maintenance.

    SPEC MET — OUTCOME MISSED source
  • 05 · Space Shuttle Challenger (1986) — normalization-of-deviance — a Criticality-1 waiver signed, then re-signed, for every flight after.

    SPEC MET — OUTCOME MISSED source
  • 06 · Post Office Horizon (1999–2015 and after) — model-as-truth, automation-over-trust — a legal presumption of reliability that put the burden of proof on the accused.

    SPEC MET — OUTCOME MISSED source
  • 07 · Nimrod XV230 (2006) — wrong-metric, under-specification — a signed £400k Safety Case that missed the exact hazard that killed the crew.

    SPEC MET — OUTCOME MISSED source
  • 08 · Genesis (2004) — component-vs-system, under-specification — three review gates, one inverted switch, none caught it.

    SPEC MET — OUTCOME MISSED source
  • 09 · Schiaparelli (2016) — literal-rule-harm, under-specification — a saturated sensor told the computer it was already on the ground.

    SPEC MET — OUTCOME MISSED source
  • 10 · Mars Polar Lander (1999) — under-specification — no requirement existed to clear a spurious touchdown signal.

    SPEC MET — OUTCOME MISSED source
  • 11 · Starliner OFT-1 (2019) — component-vs-system — every component passed; nothing tested the mission as a whole.

    SPEC MET — OUTCOME MISSED source
  • 12 · OceanGate Titan (2023) — INVERSE — inverse-no-spec — certification refused, judgment unconstrained by any external structure.

    NO SPEC — OUTCOME MISSED source
  • 13 · Therac-25 (1985–1987) — under-specification, automation-over-trust — hardware interlocks removed on trust in software with a race condition.

    SPEC MET — OUTCOME MISSED source
  • 14 · Mid Staffordshire NHS Foundation Trust (2005–2009) — goodhart, wrong-metric — targets and finances met while mortality signals were explained away.

    SPEC MET — OUTCOME MISSED source
  • 15 · Bristol Royal Infirmary (1988–1995) — INVERSE — inverse-no-spec — no outcome-monitoring system existed at all.

    NO SPEC — OUTCOME MISSED source
  • 16 · NHS A&E Four-Hour Target (2000s) — goodhart, spec-gaming — ambulances queued outside to stop the clock.

    SPEC MET — OUTCOME MISSED source
  • 17 · WHO Surgical Safety Checklist, Ontario (2008–2010) — wrong-metric — high compliance, no significant change in mortality or complications.

    SPEC MET — OUTCOME MISSED source
  • 18 · Knight Capital (2012) — under-specification, component-vs-system — a compliance review that never asked whether the trading system itself could malfunction.

    SPEC MET — OUTCOME MISSED source
  • 19 · 2008 Financial Crisis — AAA Mortgage Securities — goodhart, model-as-truth — 83% of 2006's triple-A securities eventually downgraded.

    SPEC MET — OUTCOME MISSED source
  • 20 · JPMorgan "London Whale" (2012) — goodhart, spec-gaming — a new VaR model brought the number back within limits; the exposure kept growing.

    SPEC MET — OUTCOME MISSED source
  • 21 · 2010 Flash Crash — literal-rule-harm, under-specification — a sell algorithm executed its 9%-of-volume target exactly, without regard to price or time.

    SPEC MET — OUTCOME MISSED source
  • 22 · Long-Term Capital Management (1998) — model-as-truth — spread relationships treated as stable until Russia's default confounded them all at once.

    SPEC MET — OUTCOME MISSED source
  • 23 · Ofqual 2020 A-Level Grading — wrong-metric — a school-level statistical brief met; 39.1% of grades downgraded individually.

    SPEC MET — OUTCOME MISSED source
  • 24 · Robodebt (2015–2019) — automation-over-trust, wrong-metric — an income-averaging rule run after being told internally it did not accord with legislation.

    SPEC MET — OUTCOME MISSED source
  • 25 · SyRI (2020) — under-specification — a fraud-risk score no one outside the system could actually see.

    SPEC MET — OUTCOME MISSED source
  • 26 · Dutch Childcare Benefits Scandal — Toeslagenaffaire (2013–2019) — literal-rule-harm — an all-or-nothing clawback rule with no proportionality and no override.

    SPEC MET — OUTCOME MISSED source
  • 27 · FBI Virtual Case File (2000–2005) — under-specification — scope grew 80% against an intent that was never re-baselined.

    SPEC MET — OUTCOME MISSED source
  • 28 · Denver International Airport Baggage System (1994–1995) — under-specification — a spec too complex for anyone to hold in judgment while writing it.

    SPEC MET — OUTCOME MISSED source
  • 29 · London Ambulance Service CAD (1992) — automation-over-trust — dispatcher judgment removed, nothing put back in its place.

    SPEC MET — OUTCOME MISSED source
  • 30 · Healthcare.gov Launch (2013) — wrong-metric — a completed readiness checklist standing in for tested performance under real load.

    SPEC MET — OUTCOME MISSED source
  • 31 · Zillow Offers (2021) — goodhart — a predicted price that became the buy signal it was meant only to inform.

    SPEC MET — OUTCOME MISSED source
  • 32 · Amazon Recruiting Tool (reported 2018) — goodhart, model-as-truth — a decade of hiring data treated as intent nobody had actually authored.

    SPEC MET — OUTCOME MISSED source
  • 33 · Facebook Meaningful Social Interactions (2018) — goodhart — a single engagement score that rewarded the content it should have suppressed.

    SPEC MET — OUTCOME MISSED source
  • 34 · Tesla Autopilot Recall (2023) — under-specification, automation-over-trust — a Level 2 spec with no boundary for foreseeable driver misuse.

    SPEC MET — OUTCOME MISSED source
  • 35 · Microsoft Tay (2016) — under-specification — a learning objective with no boundary, trained on the same stream it was deployed into.

    SPEC MET — OUTCOME MISSED source
  • 36 · Grenfell Tower (2017 fire; 2024 Phase 2 report) — spec-gaming — guidance mistaken for the safety outcome it was meant to protect, on top of manipulated test data.

    SPEC MET — OUTCOME MISSED source
  • 37 · Deepwater Horizon (2010) — wrong-metric — an exemplary personal-injury record in the same year process safety failed.

    SPEC MET — OUTCOME MISSED source
  • 38 · BP Texas City Refinery Explosion (2005) — wrong-metric — the same substitution diagnosed here recurred five years later at Deepwater Horizon.

    SPEC MET — OUTCOME MISSED source
  • 39 · Fukushima Daiichi (2011) — capture — a regulator that had lost independence from the operator it was meant to check.

    SPEC MET — OUTCOME MISSED source
  • 40 · Piper Alpha (1988) — normalization-of-deviance — a permit-to-work procedure habitually departed from, never reconciled.

    SPEC MET — OUTCOME MISSED source
  • 41 · Bun Zig-to-Rust Rewrite (2026) — SUCCESS POLE — success-pole — a test oracle held apart from the rewrite, enforced implementer/reviewer separation, and a human holding the merge seam — with one unexpired standing waiver left over.

    OUTCOME MET source

Read this catalog for shape, not frequency. Forty failures and one success were chosen to show the forms this pattern takes across domains — including one sighting of the opposite shape. A curated collection cannot say how often specifications succeed or fail, and it cannot prove Intent-as-Source's thesis by itself.

The same ticket, two ways

A worked example

Built to the ticket — customer charged twice

Shipped to the ticket

A payment already declined with a 402 gets retried anyway, and a customer who already saw one declined charge sees a second attempt land against the same card.

In full
Before. Literal implementation: retry any failed webhook call, three times, no exclusions. It ships. A payment already declined with a 402 gets retried anyway, and a customer who already saw one declined charge sees a second attempt land against the same card. The ticket was satisfied. The customer is not better off.

The numbers are illustrative — this is a constructed example, not a measured incident

source

Two sentences first — the retry bug never ships

Two sentences, said first

a customer shouldn't lose their order over a blip.

In full
Intent: a customer shouldn't lose their order over a blip. Boundary: never retry on a 4xx response; retries require an idempotency key.

never retry on a 4xx response; retries require an idempotency key.

sha-256 of this sentence 2a4617b17eab1572d3946f4d258f70a83fb35808ee6c028f9a6422ccc52181fd

The numbers are illustrative — this is a constructed example, not a measured incident

source source

Two minutes, one caught defect

The result

Two extra minutes on the ticket. One caught defect.

In full
Two extra minutes on the ticket. One caught defect. The numbers are illustrative — this is a constructed example, not a measured incident — but that's the shape of the trade this step is asking you to make, on one ticket, this week.

sha-256 of this sentence 99345fd1e52c0e57dd434c5a02434f0adf17828cf6f98a4af29324d8629125b6

The numbers are illustrative — this is a constructed example, not a measured incident

source

The mechanism, in full

The loop

skip the walkthrough →
Humans edit the top layer only. Everything below it is compiled, verified, accepted, and executed — and every change travels around the loop, never sideways into the derived text.
source

Five stages and one feedback edge, top to bottom.

  1. 1 of 6 Intent
  2. 2 of 6 Compile
  3. 3 of 6 Verify
  4. 4 of 6 Lock
  5. 5 of 6 Execute
👤 human judgment · 🤖 automated · 👤🤖 both, split by the judgment gate.
source

The lock ledger

An append-only record of accepted locks — a constructed illustration.

Change the intent and a new lock supersedes the old.

Each lock carries the adjudication — the recorded human ruling.

  • Lock
  • Accepted requirement
  • Digest
  • Adjudication
  • Dissent
LOCK No.001 SUPERSEDED — NEVER EDITED replaced by a newer accepted version
Derived requirement content is not manually edited in normal operation; break-glass edits are detected and reconciled.

sha-256 of this sentence 5979ec62daf032aa26fd711ec33399dc24a9f802f19b7290c94e9ad379b3792a

source
LOCK No.002 ACCEPTED
Hand-edits to anything below the top layer are break-glass: detected, flagged, and reconciled — not forbidden, just always visible.

sha-256 of this sentence e41864eba4352aed357aa22d8a374860e0e0f91320882e39375919972e6b145f

source a dissent, recorded and reconciled

Why the cost changed

Why now

Verification of compiled requirements replaces authoring and maintenance as the recurring human cost: the load moves, it does not vanish.
The paper's abstract in full
Across aviation, medicine, finance, public administration, and AI systems, one failure shape recurs: the specification is satisfied and the outcome fails anyway — compliant failure, with the executor blameless by the specification's own lights. This paper traces the shape to specifications substituting for the judgment they were meant to carry, and proposes an inversion: concentrate human editing into a small per-scope intent layer, compile requirement documents from it, and never hand-maintain the derived layer — a break-glass edit is detected drift, reconciled rather than absorbed. The diagnosis's components are measured — by others; the synthesis is ours. The prescription is a proposal. Verification of compiled requirements replaces authoring and maintenance as the recurring human cost: the load moves, it does not vanish. The primary prediction — C_compiled / C_manual < 1.0 inside a stated envelope — is untested, and the pilot protocol for testing it is published. Shipped with the paper: a specification, templates, worked examples, a 41-case library, agent skills, and a pre-registered small-n drift-detection demo.
source

The old cost

3–8 person-months to derive one requirements spec

  • the collapse is the bet this paper makes, not a result it can cite
  • one small-system comparison and directional field data
source

What is measured, and what is not

The evidence

How strong each kind of evidence is

  • Measured — the diagnosis's components
  • Measured — small-n, ours
  • Hypothesis — the prescription
  • Open — cannot yet verify
source

The drift-detection demo

Judge A, a fresh context-isolated Claude subagent per case at Haiku scale, detected 29/30 (96.7%).
  • the seeds were overt single-criterion mutations
  • upper bound for this seed class
Judge B, a GPT model driven through the Codex CLI, detected 30/30 (100.0%) by the protocol's binary count.
  • The stricter number matters more
The measured passage in full
Measured: small-n, ours. The one measurement this project has produced is the drift-detection demo ([experiments/drift-demo/results.md](../experiments/drift-demo/results.md)), pre-registered in the narrow, checkable sense: protocol, seeds, and the answer-key manifest were committed before any detection output exists in git history, results in a separate later commit — a private repo's history being rewritable, that ordering means something to a third party only from publication onward. Thirty single-mutation variants of the compiled spec for the refunds experiment fixture — §6's intent content, in the simplified pre-template form the experiment registered — ten each of boundary-weakening, scope-widening, scope-narrowing — were shown to two independent judges, each seeing only the fixture intent page, ten derived anchor cases, and one candidate spec at a time, never the manifest. One mediation must be stated: the judges never saw the raw intent alone — the ten anchors were derived from the intent by the experimenters before mutation seeding (protocol.md), so the measured rate is for intent-plus-anchors detection, an easier task than raw-intent detection. Judge A, a fresh context-isolated Claude subagent per case at Haiku scale, detected 29/30 (96.7%). Judge B, a GPT model driven through the Codex CLI, detected 30/30 (100.0%) by the protocol's binary count. The stricter number matters more: scoring a detection only when the judge's stated reason names the criterion actually mutated, neither judge caught mutation-28 (a scope-narrowing — Judge B flagged that seed for an unrelated standing complaint), and combined correct-attribution detection is 96.7%, not 100%. Both numbers are reported; the results file picks neither. False positives, from five interleaved runs each against the unmutated spec: Judge A 1/5 (20.0%), Judge B 4/5 (80.0%), combined-OR — the same rule behind the combined detection number — 4/5 (80.0%): the combining rule's cost, stated beside its benefit. High enough that a lone DIVERGES verdict is not a strong drift signal on this spec. And Judge B's false positives are the demo's second finding: three distinct complaints, each pointing at a real underspecification in the unmutated baseline — a refund exactly equal to the order total covered by neither clause; per-identity abuse flags never routed to the digest the intent promises finance; a cluster alert firing "past a defined threshold" with no threshold defined. Detection doubles as a spec-quality audit, and "false positive" mismeasures the event when the baseline itself has defects. The limits travel with the numbers: one intent page, one domain, thirty mutations, one judge pair — and the seeds were overt single-criterion mutations, so 96.7% reads as closer to an upper bound for this seed class than a floor; real and adversarial drift is subtler ([ASSUMPTIONS.md](../ASSUMPTIONS.md), entry 11).
source

False positives

20–80% false-alarm rate on the unmutated spec

  • depending on the judge
source
Combined detection reaches 100% only because the two judges miss on different cases for different reasons (see below) — neither judge alone gets there, and the combined FP rate inherits Judge B's much higher false-alarm rate rather than averaging it down.
  • neither judge alone gets there
source
Sample size is small (30 mutations, 10 per class, 5 FP runs per judge) — this is a demo run, not a large-n benchmark; per-class rates in particular should be read as indicative, not precise.
  • this is a demo run, not a large-n benchmark
The stated limits in full
On this fixture, both judges flagged 10/10 seeded boundary-weakening and 10/10 seeded scope-widening mutations, independently. Scope-narrowing is where the one clean miss happened (Judge A, mutation-28) — narrowing "who/what counts" without touching the boundary language itself is the hardest class here, on n=10 per class. The false-positive rate is not small enough to treat a lone DIVERGES verdict as a strong drift signal on this spec, particularly from Judge B. A single sufficiently literal judge will flag ~1 in 5 to ~4 in 5 unmutated specs, depending on how strictly it reads the anchors against spec wording gaps that have nothing to do with intentional drift. Sample size is small (30 mutations, 10 per class, 5 FP runs per judge) — this is a demo run, not a large-n benchmark; per-class rates in particular should be read as indicative, not precise.
source

The stance

We're asking you to test this, not believe it.
The decision rule in full
Pre-register the decision rule before you start, not after you see the numbers. We're asking you to test this, not believe it. Run the two-week comparison — and if it shows no difference, drop the framework. That result counts.
source

The success pole

The Bun run

The run in figures

  • 535,496 lines of Zig ported
  • 6,502 commits in the port
  • 11 days, end to end
  • 64 concurrent Claude agents
  • 60,624 tests, on Debian alone
  • 1,386,826 expect() calls, one platform
  • 128 pre-existing bugs, since fixed

a first-party vendor account — self-reported figures, no independent record. The author discloses "Bun was acquired by Anthropic in December 2025"; this repo's skills target that same vendor's agent harness. Read the numbers, and the telling, as a participant's report rather than an inquiry finding.

source source

How the outcome held

Once 100% of Bun's test suite passed in CI on all platforms (and I manually verified the tests were in fact running and not being skipped), I ran a bunch of commands locally to test things - and then I pressed the merge button.
source
Departures from the old behavior were tracked rather than silent: nineteen known regressions, each followed to a fix, and a running count of old bugs deliberately fixed rather than carried over.
The full source passage
Nobody on the Bun team was applying Intent-as-Source. Under an 11-day deadline, they converged on its machinery anyway. The test suite stood apart from the rewrite — "Bun's own test suite is written in TypeScript which means it doesn't depend on the runtime's programming language" — and its integrity was verified by a human, not assumed. Roles were separated by construction: "The implementer doesn't review. The reviewer doesn't implement," with reviewers in separate context windows given only the diff, because "The Claude that wrote the code wants the code to get accepted. The Claude that reviews wants to find issues in the code." A human designed the workflows, monitored their output, and held the acceptance seam personally. Departures from the old behavior were tracked rather than silent: nineteen known regressions, each followed to a fix, and a running count of old bugs deliberately fixed rather than carried over.
source

Citation status

a first-party vendor account — self-reported figures, no independent record. The author discloses "Bun was acquired by Anthropic in December 2025"; this repo's skills target that same vendor's agent harness. Read the numbers, and the telling, as a participant's report rather than an inquiry finding.
The full source passage
*citation-status: a first-party vendor account — self-reported figures, no independent record. The author discloses "Bun was acquired by Anthropic in December 2025"; this repo's skills target that same vendor's agent harness. Read the numbers, and the telling, as a participant's report rather than an inquiry finding.
  • no independent record
source

The one missing control

The port's declared strategy — mechanical, transpilation-shaped Rust first, idiomatic refactoring deferred until after release — is a standing waiver with no expiry attached.
Waiver, no expiry
The full source passage
One piece of the machinery was not reinvented. The port's declared strategy — mechanical, transpilation-shaped Rust first, idiomatic refactoring deferred until after release — is a standing waiver with no expiry attached. The failure cases in this library show what unexpired waivers age into: Challenger's Criticality-1 waiver re-signed for every flight (05), Piper Alpha's permit-to-work departures habitual and never reconciled (40). Nothing in the Bun account suggests that trajectory; the point is only that the one control the team did not independently arrive at is the one whose absence takes years to show.
source

The path — how far you go

The adoption path

It is a path, not a rollout plan: each step is optional, each step has its own kill criterion, and most teams should expect to stop at step 2 permanently. That is the design working, not a failure to graduate further.
  • each step has its own kill criterion
In full
This is the practitioner's path from doing nothing to running Intent-as-Source natively. It is a path, not a rollout plan: each step is optional, each step has its own kill criterion, and most teams should expect to stop at step 2 permanently. That is the design working, not a failure to graduate further.
source
  1. Step 1 The audit
    Most requirements never need more than this — the step exists to sort the ones that do from the ones that don't.
    In full
    Step 1 produces no artifact. It doesn't touch a ticket, a template, or a repo. Most requirements never need more than this — the step exists to sort the ones that do from the ones that don't.
    source
  2. Step 2 Spreads by pull
    It spreads by pull, not push: one engineer uses it, one review gets sharper, someone else asks how. If it isn't spreading that way, it isn't ready to spread at all.
    source
  3. Step 3 The graduated envelope
    Human-executed work and anything coupling-dense stays hand-authored the longest — for some teams, permanently.
    In full
    Human-executed work and anything coupling-dense stays hand-authored the longest — for some teams, permanently. A team running step 2 on most of its requirements and step 3 on a handful of agent-executed, reversible ones is not mid-migration to something else. That mixed state is the deployable form.
    source
  4. Step 4 Native mode
    Native mode is a forward-looking architecture. Its economics are a prediction. Another team's good result is their result, not yours — the pilot is how you test the prediction on your own numbers.
    source
Step 2 Kill criterion

30% Boundary fields left empty means stop

  • a proposed, unvalidated default
source
Step 2 Already running evals?
Your goldens are unprovenanced locks; give them a source and a provenance and you are running step 2 on coverage you already have.
  • on coverage you already have
source

The first step — Monday morning

Start Monday

Ask the one question

What is this requirement for — and if we satisfy the requirement but miss the "for," how would we ever know?
source
Ask it out loud, on any ticket, in a review, in under ten seconds.
In full
Ask it out loud, on any ticket, in a review, in under ten seconds. It costs nothing and requires no artifact. Most of the time the answer is obvious and the conversation moves on. When it isn't obvious — when nobody in the room can answer it — that's the signal the rest of this document is for.
source

Install the two skills

cp -r skills/jcr-audit skills/jcr-consume /path/to/your/project/.claude/skills/ Copied
source

Try the two-week comparison

the two-week comparison, in ADOPTION

Pick your 5 most contested open requirements — the ones people actually argue about in review.
  • the ones people actually argue about in review
In full
Pick your 5 most contested open requirements — the ones people actually argue about in review. Write Intent and Boundary for each. Run them for two weeks against a matched control set of comparably contested requirements handled the old way.
source

What ships today, what doesn't

Specified, not shipped: the native-mode controls — lock index, drift detection, oracle separation, the L0 lint
  • No tool implementing them exists in this repo.
In full
Usable now: steps 1–2 of [ADOPTION.md](ADOPTION.md) — the audit questions and the Intent/Boundary overlay — plus two of the three agent skills, jcr-audit and jcr-consume (intent-compile is native-mode; see below). Conversation practice; no infrastructure required. Specified, not shipped: the native-mode controls — lock index, drift detection, oracle separation, the L0 lint ([lint/RULES.md](lint/RULES.md) is a concept spec). No tool implementing them exists in this repo.
source

The delta

The system does not prevent the first failure of a kind. It is designed to make the second one cheap.
  • it is a design claim, with the pilot as its test, not an observed outcome
In full
The delta. What adoption buys is not "failures prevented," and this paper closes by declining that claim one last time. Misses surface when reality produces them, sometimes after first harm; the first failure of each kind still happens. The claim is narrower, and worth having — and it is a design claim, with the pilot as its test, not an observed outcome: the system is designed to make drift visible earlier — at the diff, the DRIFTED flag, the clustered deviations — instead of at the inquiry; to keep judgment authorized continuously — a named channel with real override power, whose health is itself monitored — instead of exercised heroically at personal cost; and to surface misses instead of burying them, because the ledger that records them is structurally useless for blame. The step-4 pilot measures whether any of that survives contact with a real team. The system does not prevent the first failure of a kind. It is designed to make the second one cheap.
source