Skip to content

The governed tool evaluation kit

What it is

The evaluation kit is the set of artifacts and scripts that make a rules tool provable rather than reviewed. It was built for SubrogationScreenerGKR during the box test and applies unchanged to every rules-over-facts tool in the pipeline; each tool adds one invariant of its own. The reference for extending it is Governed_Tool_Evaluation_Method.md .

The kit has seven parts: a generated contract, a signed scenario spec with removal-axis negative controls, a per-criterion runner, static coverage of the record graph, an invariant fuzzer, corpus mutation, and the fix protocol. The extractor and SensitiveIndicatorDetectorGKR Phase A are not rules tools and use the golden-set and presentation-invariance method instead.

The generated contract

A tool's input contract is generated from the records it reads, never hand-authored. For the subrogation module that is the carrier_structured profile of external_idp_manifest, produced by generate_idp_manifest.py from the field definitions, with a --diff mode as the CI drift gate . Field paths in any seed come only from the generated schema. The seed generator has a validate mode that refuses unknown paths, theory names, channels, reason codes and states, and a projection key that a rule reads with no source_path is a spec gap the generator reports, never a flat default.

The signed scenario spec

Seeds live in a signed spec, seed_spec_v4.1.json at the time of the freeze, and every seed carries an expected outcome the expert has signed . The set grew from 32 seeds (A0 to A17 plus A3b, P0 to P12) to 53 on the carrier_structured path: A0 to A30 including A3b, with A22 Morales and A23 Martinez as payload seeds, and P0 to P21 excluding P14, which waits on the sovereign immunity section.

The negative controls are the test. Each seed has removal axes: the same scenario with one decisive fact removed, and a signed expectation of what the removal changes. The removal reason is recorded in a trace_reason field. A tool that refuses everything scores zero on the positives; a tool that fires on absence fails the removals. Seeds vary difficulty in the world, not on the page, and a seed whose correct answer two competent adjusters would dispute is not a test.

Expectations are never fitted. Seeds A31 to A33 (Michigan collision, North Dakota bodily injury, Minnesota heavy commercial) were once given refer expectations that matched the engine's output on an unpublished no-fault section; the ruling was that a referral on an unpublished section is a finding, and the seeds must carry the signed expectations after the jurisdiction payload loads.

The per-criterion runner

run_conformance.py scores ten criteria per case and produces no aggregate . The rule comes from the extraction matrix, where an aggregate F1 hid which field was weak and let a form rank best while carrying the highest false-positive rate. A criterion that fails names the seed, the criterion and the trace path, which is what the fix protocol needs. The standing result at the freeze is 53 of 53.

Static coverage of the record graph

check_corpus_coverage.py walks the record graph of corpus, spec and seeds without running a claim, so that a trigger reference that does not resolve or a record no seed reaches is found before a run rather than inferred from one . Its C3 and C6 checks were extended to the per-head no-fault trigger references and to a whole-seed walk; C7 carries two known pendings, AUTO-005 activation and P14. A theory that cannot fire is indistinguishable in the output from one that found nothing, which is why reachability is checked statically rather than inferred from a green run.

The invariant fuzzer

fuzz_subro.py --n 500 generates 500 payloads and asserts the tool's invariants on each . The standing result is zero violations. The exact invariant list is not yet documented in the sources; the rules it exists to hold are the ones in CLAUDE.md that never bend, and the fuzzer is what turns them into something a commit cannot quietly break.

Corpus mutation

Corpus mutation checks that the seeds detect a corpus change: a mutated record must fail at least one signed seed, or the seed set does not cover that record. Its script name and the mutation set are not yet documented.

The fix protocol

A difference is a ruling before it is a fix . When a run differs from a signed expectation, the implementer reports the seed, the expected value and its source, the actual value and its trace path, and a proposed classification from a closed set of four: engine defect, corpus defect, wrong expectation, jurisdiction gap. Then the implementer holds. The expert rules; the implementer never decides that a difference is acceptable, cosmetic or a wrong expectation.

Nothing is edited inside the run that fails against it: not the expectation, not a seed, not a corpus record, not the field guide. Expectation changes happen only under an explicit ruling, in their own commit, citing the ruling. Corpus and spec changes go through the recorded path with an audit entry naming the ruling, and the drift check proves the seed script, the database and the snapshot agree afterwards. Every commit that changes behaviour re-runs the standing five.

The worked example is a "32/32" pass that was reached by editing expectations and corpus records to fit the failing run. It is on the list of mistakes not to repeat. A meaning error is never cosmetic: a brief saying a payment gate "did not apply" on a claim with a stated payment is a defect, whatever the sentence looks like.

Scenario cards to seeds

The expert authors scenarios as scenario cards in YAML, starting from the payload in the engine's format, never from synthetic documents. card_to_seed.py converts a card into a seed-spec entry, the entry is validated, the expert signs the expectation, and the seed is generated, run and classified . Review is staged: author, validate, sign, generate, run, classify, ruling, fix. Testing guides use scenario cards, never seeds generated by a model, and state the CI gate.

The four-day scenario phase after the renderer track runs on this cadence: batch and review, synthetic payloads only, differences classified before any fix.

Readiness

A tool is ready when two things hold. First, a green run of the full seed set against the reloaded tenant, after a tenant re-export, so the proof runs on what will be deployed rather than on a development database. Second, a blind payload authored from the format guide alone, by someone who has not seen the seeds or the corpus, that produces a correct thesis. The second test is what shows the guide is sufficient and the engine does not depend on the authoring conventions of its own seeds.

The maturity ladder the as-built inventory reports is built, regression-tested, kit-tested, pilot-ready, each with a file or command as evidence.

Extending the kit to other tools

The seven parts apply unchanged. Each tool adds its own monotonicity invariant, an assertion that removing a fact can only move the outcome one way:

Tool Invariant
CoverageAnalyzerGKR Coverage never moves toward covered on fact removal
Authority gate Authority never lowers; a missing upstream never makes fast-track likelier
CCIEvaluatorGKR CCI never lowers

The order is resolver and obligations, coverage, liability, authority gate, SID Phase B, CCI, reserve; Auto and Property first, WC and GL after. Pipeline-level invariants across tools are added once the authority gate carries the first of them. The estimate is a day to a day and a half per rules tool.

Division of labour

The expert authors the scenario cards and signs every expectation; the internal team runs the testing and benchmarking. Rulings sit with the expert, implementation with Claude Code, and the relay between them is a file committed to the repository before it is acted on. Never describe a directive as in the repository until it has been committed.

Open items

  • Corpus mutation: script and mutation set not yet documented.
  • The subrogation module's own monotonicity invariant is not yet documented in the sources.
  • Readiness for the subrogation module is not yet reached: the tenant re-export, the readiness run and the blind-authoring test remain owed.
  • A3 (the comparative gate source path) and A31 to A33 remain owed under the protocol.