Why three maps and not one

A chain is a property of an object class, and this theory specifies three different kinds of object. The three chains do not share a unit, and that is checkable rather than a matter of taste: one moves downward in abstraction inside a single firm, one moves forward in time across two firms, and one moves outward from an author. Forced onto a single ladder, papers with no position on each other's chains would land at adjacent stages and imply coverage of everything in between.

The separation also keeps one number honest. The largest specification-advantage effect sizes here live on map C, and their subjects are language-model agents, not firms. On a single map they would read as evidence that specified organizations perform better. They are not evidence of that, and keeping the maps apart is what makes the difference visible.

Two papers that sit off the maps

Both are covered work. Neither is a stage on any of the three chains, for the reason given beside it.

The reading key

Owned — a paper makes a defended claim about this stage. Touched — a claim exists but is partial, or is about a neighbouring object. Ceded — deliberately not ours, by a recorded decision. Empty — nothing. Designed and unrun — pre-registered but not executed, which is a different state from empty.

Evidence grades. A executed pre-registered study with public logs · B executed but single-case or interpretive · C calibrated simulation or Monte Carlo, labelled by its own paper as not a test · D formal result, no data · nothing.

A paper counts as coverage; an article, a post or a seed does not.

Map A — Internal configuration

The object specified is one firm's own commitments. Read against: intent → strategy → operating model → process design → procedures → execution → outcome → feedback Movement along it is downward in abstraction, inside a single entity.

# Stage Verdict Papers What is actually claimed Evidence
1 Intent touched 2026ag2026m Owner intent is named as the top tier, with a distinct governor and transfer mode, and as the origin of the constraint hierarchy. There is no theory of how intent forms, or of what makes one intent better than another. It is an ontological slot, occupied and not developed. D
2 Strategy touched 2026m2026aj Business model is a tier, and cross-tier allocation has a closed-form rule under Jorgensonian user costs. Neither says anything about the content of a strategy. D
3 Operating model owned 2026i2026h2026m2026af The home ground on this map: a six-level specification cascade with acceptance contracts, a geometric bound proving exhaustive specification impossible, a cascade-compression result, and an observer-relative equivalence condition on configurations. Demonstration is one constructed domain. D + B
4 Process design owned 2026i2026m Process contracts as unit tests, executor-invariance, and rank-deficient projection between tiers. D
5 Procedures owned 2026i2026ar Where the cascade bottoms out, plus a six-level diagnostic that examines each level for what a healthy specification looks like. Worked on the same constructed case. D + B
6 Execution ceded Ceded by decision, and the reasoning travels with it, because a settled position read as a hole is a different thing from a gap. Execution is not a tier; it is the rendering of the bottom tier by the tier below, an axis orthogonal to the stack. A running instance has no stable identity and no transfer mode, so it fails the tier test.
7 Outcome designed and unrun 2026an2026am Designed and unrun — which is a different state from empty, and the difference is the credit. An information-theoretic friction-tax model and a full archival identification template exist (continuous within-firm treatment, staggered difference-in-differences, regulatory instruments, event studies, Oster bounds). The estimates have never been run. What is reported is a Monte Carlo mechanism test and a powered placebo that bounds the measure: the index does not separate 100 zero-activity from 100 matched operating filers, d = .166, p = .243. C
8 Feedback owned 2026ae2026bn2026ar2026bd The strong end. Verification as a spectral projection whose rank bounds what an organization can detect about itself; conventional audit as a degenerate rank-1 projection; an effective-sample-size treatment showing randomly oriented evaluators saturate, so diversity must be engineered rather than sampled; a six-level audit protocol; and an optimization-depth instrument executed on a pinned 40-organization filing panel. Bounded by its own pre-registered floors: 4 of 35 organizations resolve, so the decoupling hypotheses demote to descriptive reporting by rule. It measures the public record, never internal ground truth. D + A
Arrow Status
intent → strategy Unclaimed.
strategy → operating model Owned as a formal operator, with rank deficiency and a bounded information loss. No data.
operating model → process design The same: a cascade junction, formal, no data.
process design → procedures The same, plus the constructed demonstration.
procedures → execution Ceded — see the execution stage.
execution → outcome Empty. This is the arrow every practitioner reading silently assumes.
outcome → feedback Measurement of the described organization, not of outcome. The arrow as drawn is not tested.

Of the seven arrows on this map, none carries field evidence. Four are owned as formal results, which is a materially different posture from the brand-side chain map — where the corpus owns zero of its seven arrows even as theory. The difference should not be flattened in either direction: this side owns cascade arrows as mathematics and none as measurement.

Map B — Cross-firm transfer

The object specified is what crosses an ownership boundary. Read against: screen → diligence → value → structure → close → integrate → realise Movement along it is forward in time, across two entities.

# Stage Verdict Papers What is actually claimed Evidence
1 Screen owned 2026bi A pre-close negative screen on one failure mode, tested on 350 completed transactions. Conditional on no detected gap, the cascade-type failure rate is .073, with an exact upper 95% bound of .112 — low, but not zero. A
2 Diligence owned 2026ag2026ah2026bm2026ai Theory. A six-tier ontology with per-tier governors and transferability modes, a separability diagnostic, a persistence column separating sale from succession, and portfolio-capacity and recovery-salvage instruments. D
3 Value touched 2026ai2026aj A convex-kinked value geometry with a slope discontinuity at a separability threshold, and a cross-tier allocation rule. Both formal; neither estimated. D
4 Structure owned 2026bj Theory, and it is strong. An acquisition reframed as a typed bundle of fork operations over the six tiers, recovering all fourteen named deal types as signature instances of a generative rule rather than as an inductive list. D
5 Close owned 2026bi2026bj The cascade gap is defined at closing, and it is measured on real deals. A
6 Integrate touched 2026ag2026bi Failure-propagation propositions along the service hierarchy, as theory; one of them tested, at bounded strength. D + A
7 Realise empty No post-deal performance result exists anywhere in the corpus. The one field study’s outcome is the occurrence of a single failure type, not value realised.
Arrow Status
screen → diligence Unclaimed.
diligence → value Empty. Value is modelled from tier structure, but nothing links a diligence finding to a price.
value → structure Empty. The structure work is structure-side only; it does not price a bundle.
structure → close Formal: which bundles are admissible, and what a transaction leaves behind.
close → integrate The one arrow in the whole of OST with field evidence — and it is a necessary-condition claim at half strength. Necessity consistency is .517, so about half of the failures carried no gap. A bounded safe harbour and a negative screen, never a predictor of deal success.
integrate → realise Unclaimed.

Of the six arrows on this map, exactly one carries field evidence, and it is a necessary-condition claim at bounded strength.

Map C — Knowledge artifacts

The object specified is a document, a claim, a research result. Read against: author → specify → render → transmit → verify → reuse Movement along it is outward from an author.

# Stage Verdict Papers What is actually claimed Evidence
1 Author touched 2026ao2026bg Cost asymmetry across an artifact’s layers, and invention as a typed structural operation. D
2 Specify owned 2026t2026ao A machine-readable standard for scientific claims. D + B
3 Render owned 2026ap Executed: spine preservation and rendering equivalence, demonstrated on independently authored renderings. A
4 Transmit owned 2026u2026w Evaluation infrastructure, and canon as a repository. D
5 Verify owned 2026bh2026bl2026bk2026at Executed: blinded multi-model coding with inter-coder reliability, a pre-registered decomposition of extraction disagreement, a deterministic six-class compatibility check for federated vocabularies, and an internalization study whose validation failed at its own gate — which is recorded rather than quietly dropped. A
6 Reuse touched 2026aq Executed, on a non-firm subject. Specification-first prompting beats interpersonal style once value headroom exists: d = .314, p = .009 at mid capability, d = .569, p < .001 at the frontier. Ablation shows the advantage comes from teaching logrolling, it survives paraphrase, and it replicates across model families. The subject is a dyad of language-model agents. A

This is the only map where executed studies outnumber formal results. It is also the only one whose subjects are not organizations — they are documents, claims and language-model agents. Nothing here is evidence about how a firm performs.

The arrow question

Asked plainly: is there any field evidence that a specified organization performs better?

No

Not "weak evidence", not "evidence with caveats". There is no executed study anywhere in this corpus in which specification quality is an independent variable and firm performance is a dependent variable.

Three things come near it, and each falls short in a different way. The differences are the useful part, and none of the three may stand in for the claim.

  1. The design exists; the estimate does not. The specification-readiness work builds the measure and the full identification template, and reports a Monte Carlo and a powered placebo. Its own text flags the archival execution as the next deliverable.
  2. Real field data, different arrow, half strength. The one large-N field study tests whether a closing-time structural gap is necessary for one integration-failure mode, and finds a bounded safe harbour with necessity consistency .517. That is failure-avoidance on one mode, not performance — and the paper says so in its own voice.
  3. A clean causal specification advantage, on the wrong subject. The most quotable "specification pays" result in the corpus has a dyad of language-model agents as its subject, not a firm. It belongs to map C.

The warning that matters most

Two ladders, read in opposite directions

The practitioner ladder is causal. The specification cascade is a traceability relation. Only one of them has an instrument.

The practitioner ladder claims each stage produces the next, and that the last stage is money. The cascade claims each level is justified by the level above it, and its terminal object is a checkable contract rather than a result. Traceability runs upward — does this procedure serve a stated experience contract? Causation runs downward and outward — does this configuration produce results? This theory verifies the upward relation and has no executed evidence on the downward one.

Which is why the arrow finding above is an absence rather than an oversight: the downward relation was never what this work set out to measure. The two ladders are easy to read as one, so a cascade laid over a strategy-to-execution ladder will look like a performance claim. If you see it presented that way, the presentation is claiming more than the papers do.

The count

Across all three maps there is one arrow with field data behind it, and it is a necessary-condition claim about avoiding one failure mode.