FNOLExtractorGKR
What it is¶
FNOLExtractorGKR is the native, model-backed reader of claim documents. It returns facts with a status, a verified span and the verbatim text behind each fact, and it returns nothing else. Everything downstream is deterministic, so an extraction error stays traceable to one field. The extractor is a reference implementation and a measurement baseline; document extraction itself is bought from an IDP vendor, and the extractor's output contract is the one the vendor normalizer must reproduce .
What it returns and what it refuses¶
The extractor returns facts, not conclusions. The model is never asked for claim_type, loss_designation or priority, which are rule outputs. It never computes a total: a rental total is a stated fact or it is pending at rate and limit, never rate times days. It never reconstructs a partial policy number, never treats "N/A" as a negation, and never substitutes a discovery date for an occurrence date .
Absence is a first-class answer. Every field instruction states what to return when the fact is absent, and null, not false, is the correct return when a document is silent. The distinction matters because exposure rules read "is null or == false"; a model that returns false on silence does not misroute today but would corrupt absence-type indicators later. The golden run that first surfaced this had 17 false positives of exactly this kind, cut to 3 by a worked null-versus-false example in the core prompt.
Spans are verified on the verifier's side. The model returns verbatim text only; _verify_span locates it in the assembled composite and maps the offset to a source document and page through an offset table. The model is never asked to embed document labels in a span, and a span landing in an injected header zone is treated as unverified rather than "unknown".
Two things the extractor was asked to do and declined by ruling, recorded so they are not relitigated: fuzzy, phonetic or embedding-based entity matching (party identity drives the subro target, insured status, authority tier and who receives a demand, so a similarity judgment on that path is non-deterministic and uncited), and a composite confidence threshold to fast-track claims past intake review (that gate is AuthorityGateGKR, which sits later because fast-track needs coverage, liability, indicators and reserve, none of which the extractor holds) .
The EXTRACT-CORE record¶
The extractor loads EXTRACT-CORE as its first operation. When the record had never been seeded, every document returned intake_failure and the live Auto and GL specs were unreachable; seeding it alone unblocked the pipeline .
Core is the occurrence record: about 42 fields in seven groups (occurrence, location, insured including insured.address_state for policy-state routing, policy, parties with roles, eight exposure signals, documents and reporter), 15 critical and 27 recommended, nine decision fields with inline evidence envelopes and review thresholds. No LOB field lives in core; liability, injury, vehicle and property fields stay in the LOB specs. Core facts live only at core paths; LOB specs reference them through a core_aliases map, seeded seven per spec on Auto and GL, and _merge_core_facts fills nulls only, recording _core_alias_fills in the trace. LOB classification reads the exposure signals through a deterministic lob_exposure_rules_fnol_v1 record with three-valued evaluation; the model is advisory only, and all-null exposures route to intake_review rather than being excluded .
One record serves three consumers: the native prompt, the Reducto manifest base layer and the BYO IDP mapping template. The manifest and template are generated from field_definitions by generate_idp_manifest.py with a --diff CI gate; external_idp_manifest is never edited by hand. mappable_priority has two axes, pipeline criticality and IDP reliability: only date_of_loss, loss_state, insured.name and policy_number are mandatory, critical-but-unreliable fields are "attempt, native authoritative", and recommended fields are never mandatory .
How extraction improves without code¶
Improving extraction is a spec record change: field instruction text, enum, criticality, rule condition. The prompt is rendered from field definitions, so there is no prompt file. Code changes are reserved for mechanics (span verification, merge, the normaliser). A spec version goes active only after the golden set passes at or above the recorded floor, and a model version change is treated like a spec change requiring a regression re-run. Rollback is a previous spec version, not a code revert.
Brittleness findings¶
The first serious drop in field F1 (0.924 to 0.869) was not an extraction regression. JSON mode had silently switched off on the same model string and 26 of 33 misses were the strings "true" and "false" returned in place of booleans; the gateway then went down and zeroed a run. The rulings: the validator coerces by declared type and records coerced_fields, strings that do not coerce are validation errors, output shape is schema-enforced where supported and fails closed with structured_output_unavailable where not; the client records model_used, json_mode, gateway_backend and temperature on every trace, and the harness refuses to score a run whose profile differs from the spec's llm_profile. Tools carry no default endpoint; a floor is never confirmed on a path production will not use .
The second finding was structural. On an adjuster-narrative Auto case the LOB pass extracted parties, citations, an admission and damages with verified spans while the core pass returned an empty roster, so the screener correctly gave no_referral. The ruling was that continual prompt tweaks are not the answer: the occurrence roster became the union of core and every LOB pass with found_by provenance and a both-way merge, six cross-pass consistency checks (core_roster_empty, party_missing_in_core, vehicle and injury signal disagreement, date and insured disagreement) raise review items and trace anomalies, and cross-pass agreement is a harness metric beside F1 .
The golden set and presentation invariance¶
The golden set is 20 synthetic cases (Auto 8, Property 6, WC 5, GL 4) with 13 anchor fields, exact ISO date matching, party F1 on name and role pairs, three hallucination-suppression nulls and one multi-document case (G018) whose per-party valid_source_doc_indices is the hard attribution test. The provisional floor is 0.924 field F1; it stays provisional until the bake-off winner is pinned and three end-to-end runs land within 0.01 with every LOB pass completing. Dev-path runs cannot confirm a floor.
Part C of the robustness instruction is the measurement instrument for presentation invariance. Twenty seeds are rendered across document forms (ACORD, portal export, broker email, adjuster narrative, call transcript, police report, demand letter, multi-document package) and fact presentations (labelled, prose, table, negated, absent, contradicted). Ground truth is known by construction, and the same seed rendered every way must yield the same facts. Per-cell precision and recall are reported with invariance beside recall, and a change is accepted only if it improves its cell without lowering another. Matrix v1 (383 cells, F1 0.847) was rejected as uninstrumented: report_date was scored against renderings that could not contain it, the wall clock was returned as a default, and forms were scored on facts they cannot attest. Nothing was tuned against those numbers; v2 runs on unchanged specs .
Scalability and party identity addendum¶
An outside review raised nine points; six were accepted, two declined (above), one reclassified. The addendum is four gated packages: A moves composite_max_chars from a Python constant into the spec with a chunk_strategy block, unique counter indexes and review-item volume instrumentation; B separates party identity from role, one entity with a roles[] array and a governed entity_normalization record whose do_not_merge_signals keep "Acme Corp" apart from "Acme Corp of Texas"; C makes unresolved conflicts on date of loss, policy number and named insured blockers with no numeric threshold anywhere; D is asynchronous LOB fan-out with per-pass failure isolation. Downstream tools ask has_role(party, X) rather than reading a scalar role, because a contractor who is both certificate holder and alleged tortfeasor answers differently to coverage, screener and liability. Package A was cleared on condition that a failed unique index refuses the write (index_assertion_failed) and that held_pending_review is not introduced as a fifth persistence state .
The retired native path¶
The subrogation box test began with native extraction on the two worked claims and ended with it removed entirely. Extraction variance produced an invented employer, an invented rental total and a fifth status literal; four findings went to the extractor track and the test baseline moved to the carrier_structured path alone . The concrete failure the native path could not clear was a carrier form's four-column fault table, rendered with its value three lines from its header, which a text extractor cannot read as a key-value pair. That is the case for buying extraction: the product is the interpretation layered on top.
Open items¶
The recorded floor is 0.924 field F1 on the production gateway (2026-09-06), provisional pending the model bake-off and three stable gateway runs. Later runs on a direct model endpoint reached 0.983 field F1 and 0.957 parties F1; those are development-path figures and are not valid as a production floor. G010 (silent document returned false) and G011 (explicit negation returned null) are held as model probes. The matrix is scoped to Auto and Property with WC and GL after. The standing order after Part C is Stage 2, Stage 3, Package C, Package D of the classification and addendum work.