Skip to content

SensitiveIndicatorDetectorGKR

What it is

SensitiveIndicatorDetectorGKR reads a claim for fraud, severity, complexity, coverage and compliance indicators and reports, per chunk, whether the indicator fired, did not fire, or is indeterminate. Phase A is a narrative pass: a model call over the claim documents that extracts a fixed list of semantic facts with evidence spans. Phase B evaluates chunk triggers against those facts and structured claim data, with no model call .

Phase A has no cache. FNOLExtractorGKR caches on document-set hash plus model, so its second run onward is a cache hit; Phase A fires fresh every run. The claim that seven of seven stages are reproducible holds warm, not cold.

Phase A: how the non-determinism surfaced

The Directive 3 parity harness ran Morales twice against the same source database as a self-comparison step, before any cross-database comparison. The two runs diverged inside the detector. Both reported chunks_evaluated 15, signals_found 15 and phase_a_status success. Run 1: 3 citations, 0 chunks fired, 0 hard flags. Run 2: 14 citations, 1 chunk fired, 1 hard flag, Commercial Vehicle Involvement. FI-001 through FI-008 moved from not_fired to indeterminate; WSI-007 moved from indeterminate to fired.

A parity harness that compares two databases without first proving one stable cannot attribute anything it finds. Martinez did not pass either: its third run differed, but the flipped fact landed on a negative assertion, so the Martinez pass was a coincidence, not a result.

What Directive 3A established

The prompt is byte-identical across five Morales runs (one hash) and the model returns a different response every time (five distinct response hashes; ten from ten direct replays). Divergence is entirely at the inference layer. Endpoint: a LiteLLM gateway, model openai/gpt-oss-120b, temperature 0.0, max_tokens 4096, response_format json_object, no seed.

The expert then directed a rebuild of Phase A's input side without reading the tool. All of it already existed: evidence_span required for fraud and severity facts, _verify_evidence_span checking exact then whitespace-collapsed case-folded, a hallucinated span clearing the fact to indeterminate, asserted-without-span to indeterminate, invented facts discarded, confidence gated before the status branch, prompt and response hashes captured. The rebuild was withdrawn. The standing rule: a diagnostic report says what was observed, not what exists; read the source before naming the fix, and any statement about what a component does carries its file and line or is marked an assumption.

The completeness gap, confirmed in source

In _run_phase_a the validation loop iterates what the model returned. Nothing iterates required_facts to confirm every requested name came back. Both lists sit in extraction_metadata, fact_spec for what was asked and facts for what arrived, and were never compared. phase_a_status was set to success unconditionally once _run_phase_a and _verify_all_spans returned without raising, with degraded hardcoded false on that path.

So the flips are the model returning fewer keys, not asserting different things. A fact absent from the response resolves indeterminate per chunk, which is correct. What was missing was anything noticing that a third of the requested facts never arrived while the run reported success. Completeness must be measured, never assumed; it is the same shape as the empty overlay list and the empty exclusion catalogue.

Directive 3B

Part 1 measured the outcome spread nobody had: the ten replay responses run through Phase B, reported as distinct signal sets rather than response hashes. Every difference is classified by direction. Conservative drift is a signal moving toward fired or indeterminate, tolerable over-flagging. A missed indicator is a signal that fired in one outcome and returned not_fired in another, the failure that damages a demo. The two are reported separately and never averaged; a single variance figure hides the one that matters.

Part 2 added the completeness check: compare requested to returned, record facts_requested, facts_returned and facts_missing, set degraded true when any are missing, and surface a phase_a_status of degraded distinct from success and partial. Per-chunk behaviour is unchanged; the output now separates an incomplete extraction from a narrative that lacked the fact.

Explicitly not authorised in 3B: a seed, a Phase A cache, prompt changes, the strict schema, changes to _evaluate_chunk's handling of a missing fact, or resuming parity. Each would make the harness quieter without making the product more reproducible. A seed would make the failure reproducible rather than absent, with nothing in the output to say so after a model or gateway change.

What 3B returned

Ten replays through Phase B gave 3 distinct signal sets, 12 conservative-drift chunks and 0 missed indicators. The model returned not_mentioned for every fraud fact in all ten runs; what varied was its certainty about an absence, which the 0.7 confidence gate converted to indeterminate instead of not_fired. This was never a decision-path breach: the gate and the null check absorb model variance in the conservative direction by design. The earlier reading, that an SME could watch a hard flag disappear on rerun, is disproved; residual exposure is a presentation inconsistency between "not established" and "could not determine".

The completeness fix proved itself on a live run: facts_requested 26, facts_returned 24, facts_missing commercial_vehicle_involved and subrogation_demand_received, phase_a_status degraded where the same run had reported success. Key omission is intermittent: WSI-007 was stable across the ten replays and unstable live.

Seed was ruled out on measurement: seed=42 gave 10/10 distinct hashes, identical to no seed, and two runs still dropped keys. Strict json_schema eliminated key omissions entirely (zero across ten runs) without touching value variance. It was implemented and cleared, fail-closed on structured_output_unavailable with no prompt-only fallback as FNOLExtractorGKR does; the completeness check stays regardless, since the schema is a provider-side control and the check is ours. The golden set (2 cases x 2 modes x 3 runs) showed zero mode delta; the key-omission benefit is real but not demonstrable on cases that return 26 of 26 in both modes.

Two harness bugs were caught by results that were too clean. run_sid_golden.py's stub __init__ discarded config so both arms ran as strict, a comparison that could only ever pass. A single-run 12-signal "mode effect" on Martinez, first called conservative, was one sub-0.7 tail event, retracted after 12 runs; indeterminate to not_fired is a claim gaining certainty it did not have.

The Morales flip

The model returned negated (confidence 1.0) and not_mentioned (confidence 0.0) for multi_vehicle on the same Morales document. Both readings are defensible on a rear-end collision and they mean different things downstream. This is why the parity ruling put the fix in the harness: run Phase A once and feed the captured output to both the source and target runs, so any difference is attributable to the database. Removing the gate would have collapsed the diff by coincidence, and a Phase A cache would report parity the product does not have. Non-determinism stays visible in production.

The confidence gate ruling

The threshold is not the problem; the dimension is. On an asserted fact, low confidence means the model is unsure the thing is there and refusing is right. On a not_mentioned fact it is a self-report about an absence with no calibration in either direction. The gate stays on asserted facts and comes off not_mentioned, replaced by the completeness result: a fact that arrived is evaluated, a fact that did not is indeterminate. Where confidence stays in use, the threshold moves out of the Python default into the chunk record with a stated basis, because a parameter deciding whether a fraud check reads as clean or unevaluated is governed. Not POC-blocking; the not_mentioned confidence distribution is measured first, and the change is multi-run on its own, never on the run that is also its evidence.

Phase B and the carrier path

Phase B evaluates structured triggers with no model call and takes the standard rules-over-facts kit. Its Auto trigger catalogue was found empty during the classification ladder correction: no Auto signal had a structured trigger while the tool was classed shadow-ready. On the carrier-structured path there is no narrative for Phase A, and the detector records detector_not_run with a reason rather than an error; reporting error and refusing synthesis on a carrier payload was ruled a defect . Reducto changes none of this; Phase A is a separate call downstream of extraction.

The false affirmative

"No indicators fired" was printed after the detector was handed empty content. It was one of three false affirmatives in a single readiness audit, the class of defect that loses a POC outright: a brief full of honest gaps passes and a confident brief with one fabrication does not. Could-not-evaluate and considered-and-rejected must never render the same; the three-state rendering that closed it is under .

Open items

Whether WSI-007 and FI-001 through FI-008 are tagged fraud, severity, complexity, coverage or compliance is unestablished, and it decides whether WSI-007 was ever span-checked, since span discipline in _evaluate_chunk applies only to fraud and severity. The confidence-gate change awaits its measurement. The detector has never been held to the six-field vendor test with true negatives applied to the IDP.