Skip to content

Principles and rulings

What it is

These are the rules that hold across every tool and every LOB. Each was established once, under a ruling, when a build went wrong in a specific way. They are recorded here as product rules with the example that produced them, so that a new tool is held to the same bar without rediscovering it. The repository's standing rules file carries them for the implementing agent; this page carries them for the people who review its work .

Rules on what the engine may read

No language model call on any rules path. The harness asserts llm_construction_count == 0 per run for every rules tool. The subrogation module runs from carrier payload to thesis with zero model client constructions, and the assertion is what proves it rather than the design intent.

No rule content, threshold, pattern list, state rule, default or fallback in code; records only, read through the spec's input_map and source_path. The reserve tool once produced a complete reserve from platform constants during a database outage with nothing in the output to say so, and the WC tool defaulted the waiting period, retroactive period, maximum duration and statutory rate percentage across all states in the same file where its state-maximum work was exemplary.

A key a rule needs with no source path is a spec gap you report, never a flat default. The seed generator's validate mode reports a projection key read by a rule with no source_path as a gap; the runner walks source_path and maps nothing itself, and the "3e vehicle bridge" and prefix splits mapped in the runner are logged debt awaiting a ruling.

Defaults point the safe way. Studio authentication was correctly written and switched off because absent configuration meant off; the rule is that absent configuration means the control is on, and disabling takes an affirmative flag logged at startup. The same direction governs the bundle allowlist, where absent configuration means excluded.

Rules on what the engine may conclude

An error is never a decision. A missing resolver, section or upstream output stops the run with a named error; a fallback to fault_system when a no-fault section was missing was reverted by ruling because it turned an error into a verdict.

Absence of a catalogue is not absence of the thing catalogued. The coverage loader returned an empty exclusion list as success, so a verdict of covered was reached with zero exclusions evaluated; the screener silently dropped a chunk id that did not resolve, so a mistyped id and a claim with no recovery produced identical output. An empty catalogue is a missing governed input, and the run refuses.

Three states, never two: skipped (does not apply to this LOB), not_configured (applies but is unbuilt), error (should have run and failed). An empty dict once meant all three, so a deliberate Property liability skip rendered as a stage error and permanently blocked action synthesis. Dependency checks are positive, satisfied only when a struct says ran or skipped, never by the absence of an error key.

Nothing is estimated. Money heads are stated, computed from stated inputs, pending (rate and limit only), excluded, or indeterminate; a deadline with no date is never computed. A rental stated as $40 a day with a 30-day limit is a pending head, not $1,200, and an adapter that recomputed rate times limit was reverted twice.

Closed enums only, never parse free text. Reason codes and statuses are closed sets and a non-enum value raises. A theory exclusion reason on negligent entrustment that arrived as free text was logged as an engine finding to be made an enum, and a fifth status literal (absent, present, unknown) is refused wherever it appears .

The false-affirmative bar. A brief full of honest gaps passes; a confident brief with one fabrication does not. Three false affirmatives were found in one audit, two certain to fire on every Property claim: "no escalation required, proceed with standard handling" printed when every upstream section had errored; covered with zero exclusions evaluated; and "no indicators fired" after the detector was handed empty content. Could-not-evaluate and considered-and-rejected must never render the same.

Rules on what the trace must show

Every gate, defense, exclusion and trigger in the active corpus appears in the trace with the fields it read and its outcome, fired or not. A field that names a control claims the control ran: an empty overlay_rejected list and an inferred state_overlays_applied both asserted work never done, and "no triggers declared anywhere" reported as a clean audit was an unbuilt layer.

Every indeterminate names the fields it needed. The fleet negligence theory on the Morales worked example is indeterminate because none of its five any_of narrative fields was supplied, and the brief says which five. A refusal must be specific: the state, the missing parameter and what would resolve it.

Four fact statuses only: stated, not_stated, not_extractable, not_returned. A carrier may send only the first two; an omitted field is not_returned and resolves indeterminate. not_returned (not supplied) and not_stated (verified absence) are different facts and are never merged, mapped to each other or rendered with the same words. Defaulting an absent key to not_stated would have the harness asserting a fact about a document nobody read .

The fix protocol

A difference is a ruling before it is a fix. Every difference between expected and actual output is reported with the seed, the expected value and its source, the actual value and its trace path, and a proposed classification: engine defect, corpus defect, wrong expectation, jurisdiction gap. Then the implementer holds. It does not decide that a difference is acceptable, cosmetic or a wrong expectation.

Never edit an expectation, seed, corpus record or field guide inside the run that fails against it. Expectation changes happen only under an explicit ruling, in their own commit, with the ruling cited. The rule was written after a "32/32" pass was produced by editing expectations to fit the output, and after seeds A31 to A33 were given fitted refer expectations that had to be replaced with signed ones.

Corpus and spec changes go through the recorded path with an audit entry naming the ruling; afterwards the seed script, the database and the signed snapshot must agree, and the drift check proves it. A promotion of a theory from approved to active without a ruling is reverted, and the bundle built on it is withdrawn.

Every commit that changes behaviour re-runs the standing five commands and reports their output verbatim. See How Work Is Run.

Who decides

VS assigns work and relays rulings. The expert, the process lead in VS's Claude project, rules on every difference between expected and actual output and on every corpus, spec, seed or expectation change. Claude Code implements, measures and reports; it classifies a difference as a proposal and holds for the ruling. When a directive and the standing rules disagree, the implementer stops and reports the conflict rather than choosing. The expert's own discipline is symmetrical: a claim about what a component does carries its file and line or is marked an assumption, and a refusal is evidence of a working boundary only once the document has been read and confirmed to hold nothing to refuse.

Open items

  • Debts logged against the runner-mapping rule (3e vehicle bridge, flatten bridges, prefix splits) await rulings.
  • The negligent entrustment exclusion reason is still free text pending the enum change.
  • Trace contract step 1 replaces status envelopes printed by IntelligenceBriefGKR and ObligationEngineGKR with closed-enum determinations.