Document Classification
title: Document Classification description: The multi-document intake layer for FNOLExtractorGKR: how documents are classified, read under profiles, reconciled by governed precedence and attributed, and the rulings behind each stage. kind: authored status: draft source_memory: [document-classification, fnol-extract-core, idp-vendor-boundary] source_date: 2026-09-30 tags: [DocumentClassification, FNOLExtractorGKR, document_type_rules, document_profile, fact_precedence, provenance, policy_snapshot, AUTO, PROPERTY]
What it is¶
Document classification is the layer that lets the extractor read a whole claim file rather than one form. A carrier's file carries the FNOL beside a policy document, an incident report, a police report, medical records, demand letters and photographs. The extraction spec answers what is being extracted; a multi-document file adds two axes the spec does not have, what each document is and who is speaking in it. Both were flattened into the composite and lost, and that loss is load-bearing: document type is what produces the coverage analyzer's policy_source, the liability analyzer's completeness class and the compensability tool's accounts_present. Three built tools were waiting on a fact the extractor threw away .
The design position¶
There is no spec per document type. Four LOBs by ten forms would be forty specs and date_of_loss defined forty times, which is the drift the core-paths rule removed. Field definitions stay in the core and LOB specs; document profiles overlay the reading.
The sequence is: stage 0 segment and classify; stage 0b route policy documents out of the claim path into the policy_snapshot; stage 1 read with per-fact attestation; stage 2 reconcile by governed precedence rather than document order; stage 3 derive collection-level facts; stage 4 onward unchanged.
Three GKR record types carry the layer:
| Record | What it holds |
|---|---|
document_type_rules |
Deterministic structural signals, evaluated first |
document_profile |
Per type: document_class, attests, does_not_attest, reading_notes, reliability_class, speaker, own_pass |
fact_precedence |
Ranking by document type per field class, with recorded rationale |
The document catalog is governed content in the same way the indicator catalog and the rule records are. Reliability class and speaker are curation decisions on a defined scale, and adding a document form is one record.
Rulings¶
An unclassified document is read under the generic profile with provenance_uncertain and a review item, never skipped. Fail closed means the document was read and the system said it did not know what it was.
Classification is deterministic first. The model is advisory, runs only where rules return null, is batched to at most one call per file, and never overrides a rule-assigned type. IDP-supplied types are treated exactly as deterministic rule results and go through the same reconciliation.
Policy documents never feed core_facts; they feed the policy_snapshot under its own spec. The failure mode otherwise is a policy expiration date landing in occurrence.date_of_loss.
Provenance comes from the verifier-side offset table only. The model is never asked to embed document labels in a span; the extractor's step 4 ruling stands.
No model runs in reconciliation. Precedence is a governed record; equal rank or no rule raises a review item, never a silent pick. A conflict is found by a sequence with no numeric threshold anywhere: normalise first, because format differences are not conflicts and names route through the entity normalisation record rather than a second mechanism; then apply precedence; then block only what precedence cannot resolve (equal rank, no rule, or one document contradicting itself). A date threshold would silently pass the one-day discrepancy that moves a loss across a policy period boundary, and a policy number prefix match is a similarity heuristic under the same decline as fuzzy entity matching. The reason code cross_pass_disagreement separates core-versus-LOB model instability, a robustness matrix signal, from two documents disagreeing, a claim fact question for precedence.
does_not_attest is a hard suppression at merge time with a trace entry, not a fact that is merged and later outranked. Stage 2 therefore cannot be skipped ahead of stage 3: suppression must remove a fact before precedence ranks it, or a demand letter's date competes with the police report's on ranking rather than being excluded outright.
Precedence is seeded for five field classes only: date of loss, named insured and policy number, loss location, injury facts, and party identity and insurer, each with rationale recorded the way a citation is. Demand letters and complaints never rank on any factual field, and photograph sets rank nowhere. A document type absent from a ranking does not compete, which differs from ranking last and is what stops the narrative winning by default. The Part C conflict rows show which other classes actually diverge rather than seeding twenty from an armchair .
Merge order as a hidden rule¶
First-non-null by document order is an unwritten precedence ranking that decides real answers, and unwritten rules are what the program exists to eliminate. Making it a cited record with a rationale is the same move as taking thresholds out of code. The same reasoning retires form-specific reading rules patched into the shared prompt: the demand-letter date rule and the unlabelled-email insured rule were both fixed that way, applying a rule about one document form to every document. The answer to "continual prompt tweaks are not affordable" is an overlay record per form, not more prompt text.
The provenance layer¶
The extraction contract has a second axis: not only what was extracted but what attested it. Provenance was mechanically available from the offset table well before it was usable, because the missing piece was never where a span came from but what the document means. Every verified span already carries a source document and page; the profile gives that document a class, a reliability and a speaker, and precedence gives the fact a reason to win.
Incremental arrival (stage 5) is the strongest demonstration the platform has: the same tool answering the same question better because a document arrived, coverage moving from provisional to verified when the declarations land, liability from one-sided to official when the police report does.
The directive and its stages¶
The build directive is five stages, each ending in a written report and a hold, with no stage begun before the previous is signed off: 1 classification and routing; 2 profiles and per-document reading; 3 precedence reconciliation; 4 derived collection-level facts and downstream wiring; 5 incremental arrival. Each stage is independently valuable, so stopping after stage 4 leaves a coherent system. Stage 4 is where policy_source stops being hardcoded to fnol_stated in the harness, retiring the largest limitation in the cleared Auto coverage build. Package C of the extractor addendum (conflict material on date of loss, policy number and named insured) depends on stage 3 .
Stage 1 lands before any multi-document synthetic file is run. The Part C matrix already carries multi-document package as a document form and contradicted-across-documents as a fact presentation, so classification accuracy by document form and the rule-versus-advisory share become per-cell matrix metrics. The standing order is Part C first, then stage 2, stage 3, Package C, Package D.
Reducto versus native¶
The IDP vendor is expected to handle much of the segmentation from a supplied template, and the native path must be an alternative that layers cleanly on what is enabled rather than a parallel design. Vendor parity is non-negotiable: identical output contract including classification_source whether the IDP or the native path ran, and no downstream tool may branch on which. The position as of mid-September is that an IDP is the primary extraction plan and the native extractor is the fallback and the demo path; a carrier with security concerns about model use may never run it. The box test showed why segmentation matters on either path: one bundle arrived as a single scanned PDF holding a demand letter, a loss-of-use memo, a repair estimate and photographs, and bundled PDFs need page-level segmentation and classification before the composite. Unreadable scanned amounts return indeterminate, photographs are recorded as evidence available, and the other side's fault assertion is a party statement .
Open items¶
Which stages have been built and signed off is not yet documented here; the standing order places them after Part C. Precedence beyond the five seeded field classes waits on the Part C conflict rows. Page-level segmentation of bundled PDFs is a Phase 1 item. Whether classification_source is populated on the vendor path is not yet documented.