VALUE.md - Keep a Value
You ship work. It does not land. You ship more. It does not land. Or: you ship work, the work lands, and somehow nobody on the recipient side notices it landed. Both shapes are the build trap - the first is a clarity failure (you cannot say what was supposed to change), the second is a falsifier failure (you shipped without a way to prove a stranger could check the change happened). The instinct in either case is to build faster. The instinct is wrong.
The standard is a one-page contract you write at the project root, named VALUE.md. It asks three questions before you build: who the work is for, what changes for that person, and how you would know it landed. A stranger should be able to verify the answer to the third one without asking you.
If you cannot answer the three questions cleanly, you are not ready to build.
A note on scope before you go further. The build trap is sometimes a clarity problem and sometimes a structural problem. The standard helps with the clarity problem - the moments you cannot say what you are building or who it is for. It does not protect you from the structural ones: an investor demanding velocity, a stakeholder rewarding shipping over landing, a public roadmap that has political cost to change. A VALUE.md can name those forces honestly. It cannot dissolve them. Read the standard knowing which half of your build trap you are in.
A note on exploration vs execution. The standard expects you to have done at least one round of observation - sat with the recipient, watched their behavior, ridden along for a real workflow - before you write the three sentences. If you have not, the sentences will be a fiction you can polish to passing every gate. The standard does not work for exploration-stage projects where you are searching for the recipient; Twitter was not a podcasting tool with a clear Q1. The standard is for the second project, after you have learned who the work is for from running the first one. If you are at the first one, ship the broken version, watch what happens, then come back and write the VALUE.md against what you observed.
Value Statement
Keep a Value gives the builder stuck in the build trap a one-page contract that names the value being transacted; the builder gives back the discipline of writing it honestly before they build.
This sentence is the standard's lede. Q1, Q2, and Q3 elaborate it. If you cannot say the Value Statement of your own project out loud cleanly, you are not ready to build - and the three questions below diagnose where the leak is.
Q0: The grounding question (before Q1)
What is the most recent specific moment you observed your recipient doing the thing your project is meant to change? A date, a place, behavior you watched - not a persona, not what they said in a survey, not what you assume.
For Keep a Value, the recipient is the builder stuck in the build trap. The grounding moment: a specific evening in 2026, an unnamed founder describing four launches in a row that did not move the needle, asked to name who their last launch was for, gave the answer "everyone who needs it" without recognizing the problem. The standard exists because that moment is common and the builder did not see it.
If you cannot answer Q0 for your own project, the three questions below will be a fiction you can polish to passing every gate. Go observe before you write.
Q1: Who it's for?
The builder who has shipped multiple times, knows something is off, and cannot articulate who their last launch was for. Not the builder who knows who it was for and cannot stop building it anyway - that builder has a stopping-power deficit the standard does not solve. The standard solves the clarity-deficit half of the build trap. The stopping-power half is a different problem; do not pretend a contract addresses it. If you find you cannot slow down long enough to complete this gate honestly, that is a stopping-power problem this standard cannot solve - name it before reaching for the skeleton.
About platform and multi-recipient projects. Q1 asks for one specific person. Internal platforms, infrastructure, public goods, and many-recipient projects often have no single decider. The honest move is to name the class of recipient (the on-call engineer; the citizen using a public service) plus the observable proxy you can watch in lieu of a single decider's verdict (median time-to-cause across all on-call engineers; service uptime as a stand-in for citizen access). The recipient-decides axiom still holds - you are just acknowledging that "the recipient" is a population, and the proxy is the population's behavior in aggregate.
The proxy is for the verdict, not for the grounding. Even when the recipient is a population, the Q0 grounding observation still requires you to have watched at least one specific individual in that population hit the friction your project is meant to change. A platform team that names "the engineering org" as the recipient and "p99 deploy time" as the proxy can pass Q1 - but Q0 still requires you to have sat with one named engineer at her desk while she fought the deploy that took 45 minutes. The population is for measuring whether the change landed; the individual is for proving the change was needed. If you can name the proxy but cannot name the individual, your VALUE.md is gameable and you have not done the work the standard asks for.
Q2: What changes for them?
Before: you build, it doesn't land, so you build again - stuck in the build trap. It's a Jenga tower; nothing's going to change except which way it falls.
After: your process has a circuit breaker - three questions to answer before you build, and a gate that asks every next move what value does this earn? before it ships. If you use them, you stop, you think, you cut. The mechanism: writing for a named stranger makes vagueness visible in a way that thinking about it does not. The standard doesn't do the thinking for you; it just puts the questions in your path. And it gives you a check - between work and "done" - where the named recipient, not you, decides whether the change happened.
Q3: How will you know the change happened for them?
Give two builders the same problem. One gets a VALUE.md - who it's for, what changes for them, how you'd know. One gets a description in plain text. The claim (Status: Proposed, see below) is that the description-only builder would get pulled into the build trap and the VALUE.md builder would not - that the VALUE.md builder would reach a result the named recipient accepts faster and you could point at what changed for the recipient. That is the hypothesis. It is not yet established. The paired-run experiment that would test it is the standard's own Q3 falsifier and the v1.0 to v1.1 gate.
"Reaches a result" means: the named recipient is shown the artifact and says it does the named change for them - not the builder's self-report, not "it compiles," not "it looks done."
Breaks if: the VALUE.md builder takes the same or more total time to a result the named recipient accepts. Or: the acceptance instrument's recorded verdicts (kept under audits/q3-acceptance/YYYY-Q.md in the using team's repo) show no greater rate of "the change happened for me" responses for the VALUE.md condition than for the description-only condition over a rolling 90-day window.
Check: Demonstrate / Judge - paired runs, evaluator-blind via artifact-only judging (participant blinding is not feasible since the treatment must be disclosed to the treatment-group builder), preregistered rubric, prompts and scoring open so anyone can re-run. (The paired-run rubric is to be written when we run it.)
Status: Proposed. We haven't run it yet.
Last verified: not yet · Re-verify by: empty (status is not yet Proven)
Promises
P1 - You can name the trap you're in
Status: Proposed · Recipient: the builder stuck in the build trap (often without knowing it).
Last verified: not yet · Re-verify by: empty (status is not yet Proven)
You read the page once and can say "I'm in this - here's how" - before the next missed deal or burned year.
Breaks if: builders who later recognize they were in the trap report that the page failed to name their condition at the time they first read it.
Check: Field - survey readers 30 days after first read; "did you recognize yourself, and how soon?"
P2 - You reach a result the recipient accepts faster than without
Status: Proposed · Recipient: the builder who's just missed - deal lost, demo broke, users churned - and is about to reflex into "build more."
Last verified: not yet · Re-verify by: empty (status is not yet Proven)
With a VALUE.md, you reach a result the named recipient accepts in less total time than a builder working from a plain description of the same problem.
Breaks if: in Q3's paired test, the VALUE.md builder takes the same or more total time than the description-only builder to a result the named recipient accepts.
Check: Demonstrate - Q3's paired runs (rubric to be written when we run it).
P3 - Nothing ships until it survives the six-part gate
Status: Proposed · Recipient: the builder shipping fast on an LLM-augmented team, drowning in plausible output.
Last verified: not yet · Re-verify by: empty (status is not yet Proven)
You ship less, but everything you ship has stated its value, named what it won't break, and passed the six-part gate - including the subject test (you promise only your own behavior) and the side-effect check (you named what gets worse if you optimize hard).
Breaks if: in any rolling 90-day window of a KAV-using team, an artifact ships whose VALUE.md fails at least one of the six gate checks, and no remediation issue is opened within 14 days of the failing audit. The breakage is the combination of failed gate + no follow-up - either one alone is recoverable; both together means the standard is being treated as advice, not discipline.
Check: Inspect - quarterly audit of a random sample of shipped artifacts from teams using KAV. Each VALUE.md scored against the six-part gate; failures logged with a remediation deadline. Audit log lives at audits/p3-shipped-artifacts/YYYY-Q.md in the using team's repo.
P4 - Every miss pays you back
Status: Proposed · Recipient: the builder who just ran the check and got no - the deal didn't close, the demo broke, the result didn't land.
Last verified: not yet · Re-verify by: empty (status is not yet Proven)
Even the failure pays you - you walk away with either the value, or a direction to cut. You always leave knowing more than you started.
Breaks if: a builder uses the standard, the check fires negative, and they end the cycle with no usable diagnosis - same blind state they were in before.
Check: Field - quarterly post-miss interviews; "after this miss, did you walk away with a usable next move, or still blind?"
What we don't promise
- We don't promise your idea is worth building. We only help you tell whether you delivered the change you set out to deliver.
- We don't promise the right next move. The standard asks the questions; you still answer them.
- We don't promise your theory of change is correct. A clean VALUE.md - one that passes every gate, names a real recipient, states an observable Before / After, and wires a falsifier - can still be aimed at the wrong outcome. The standard validates that the change is statable, observable, and falsifiable. It does not validate that the change is what the recipient actually needs. Articulation is a necessary condition for landing, not a sufficient one. Being right is still the builder's problem.
What we don't break
- We won't make builders move slower as a measurable effect. If the six-part gate adds friction that round-2-style functional tests show is greater than the time it saves a builder from a wrong turn, the standard is net-negative - we re-evaluate the gate's shape, not the builder's commitment.
- We won't reduce the standard's accessibility for non-software builders to gain rigor for software ones. If non-software adoption signals drop while software signals climb, we're trading reach for depth in the wrong direction - we re-evaluate the runbook's vocabulary, not its rigor.
Status of the standard itself
Pre-1.0. The standard has been applied to 12 real projects so far. Two retrospective field studies have been published (see research/audits/2026-06-04-stranger-test-10-rounds/case-studies/) as honest interim evidence. They show that adoption is associated with reduced vague-recipient language and reduced frustration markers on greenfield work, and that 9 of 12 produced VALUE.md files in the field hold the gate's strictest rules. These are not the paired-run experiment; they are field evidence. The paired-run experiment - this standard's own Q3 falsifier and the v1.0 to v1.1 gate - has not yet been run.
We are deliberately disclosing this because the standard's own lifecycle rule says a claim Proposed for more than two cycles either gets a check wired or gets rewritten down to something checkable today. We are walking past our own rule. This is itself data: a measurement of whether articulation alone is sufficient to interrupt the build reflex. The paired-run experiment will need to measure this too.
We also do not yet know what fraction of builders in the build trap have clarity deficits versus stopping-power deficits; if the stopping-power case is the more common one, the standard's reach is narrower than the opening paragraphs imply.
Why this discipline now
Two operational facts about 2026 builder cadence change which checks against the build trap can actually run, and which cannot.
The cadence argument. A builder working with current LLMs produces roughly 50 to 300 times the per-day artifact rate they produced in 2022. Peer review on every artifact does not staff to that throughput; async review queues grow faster than they clear; LLM-on-LLM review inherits the producer's blind spots and so does not substitute for it. The classical finding that self-review misses defects peer review catches is real (Pronin, Lin & Ross 2002 on the bias blind spot; Bacchelli & Bird 2013 on what code review actually yields), and KAV does not refute it. KAV makes a different claim: at LLM-assisted cadence, the operational constraint moved. A self-administered six-part gate is the only check that runs on the long tail of artifacts where peer review is no longer affordable. A check that runs imperfectly beats a check that does not run. Peer review on the load-bearing minority of artifacts where it is still affordable remains wise; KAV is not a substitute for it there.
This argument is itself empirically untested - it predicts that a self-administered gate at LLM cadence produces better recipient-accepted outcomes than the realistic alternative (no gate at all on most artifacts, because no reviewer was available). The paired-run pilot needs to test this version of the comparison, not the classical self-vs-peer comparison the bias blind spot literature settled.
The deployment range. The standard documents the discipline as a one-page VALUE.md written at the project root before the first commit. That is one valid deployment. There is another, observed by the author across a ten-day window (research/audits/2026-06-17-author-self-observation.md): the three questions walked through in conversation - with an LLM, with a colleague, or silently - before any artifact is requested. No file. The discipline is the questions and the gate; the file is one artifact the discipline can produce. The conversational deployment matters because it is what the discipline looks like when the builder is moving at LLM cadence: not "stop everything and write a file," but "name the recipient, the change, and the verification before this next thing gets built." Both deployments are in scope. Both face the same gate.
Where this comes from
Keep a Value is a distillation, not an invention. The load-bearing borrows are Promise Theory (Burgess 2005) for the subject test, Donabedian's triad (1966) for the feelings-need-a-proxy rule, the pre-mortem (Klein 2007) for "what we don't break," and the self-explanation effect (Chi et al. 1989) for the Value Statement Test. The full lineage with primary-source citations is in lineage.md.