Facts, not conclusions
What it is¶
The extractor writes facts. Every later stage writes conclusions. This is the most-applied rule of the whole build and the spine of the manual. It decides what a field may contain, what a field may be called, which stage owns a derivation, and which of two disagreeing values is the defect.
A fact is something the document states, carried with its status and its span. A conclusion is anything derived from a fact plus a rule. Conclusions belong downstream, in a tool that reads records, because there they can be cited, traced and rerun. A conclusion made at the reading stage is invisible to the trace and cannot be governed.
The two competent people test¶
The test is short. Would two competent people derive it the same way from the fact plus a rule? If yes, it is a derivation and belongs downstream. If two competent adjusters could dispute it, it is not a fact at all and cannot be scored either way. A seed that resolves a relative date to a contested day is not a test.
The refusals, with the example behind each¶
Each refusal below was a live case in the build and is now a standing rule.
No computed totals. A rental stated as $150 per day for 12 days extracts as rate and duration, never as $1,800. The same rule was reverted twice more in later work: an adapter recomputed a rental total from 40 times 30 and set a medical reserve figure from a BI reserve, and both were reverted by ruling .
No resolved relative dates. "Last Tuesday" stays as text, with the report date as a separate anchor. Resolving it is a downstream derivation with an anchor the trace can show.
No comparisons. The field is prior_claim_body_part, not same_body_part. Comparisons belong in the consumer, and identity comparisons run on party_key rather than names: owner_distinct_from_driver_present is true only where an owner party and a driver party resolve to different keys, and indeterminate where either is absent, so a theory stays open pending a fact rather than being excluded on an absence.
No derived signals. A prior notice letter and a stated cause are two facts, not a third_party_at_fault_signal. Roster presence flags are derived signals with derivation records, never carrier fields.
No selection between causes. Where a document offers more than one, the extractor carries them all.
No policy_source at extraction. It is derived from the document set, so it belongs to the classification and provenance layer.
No composite confidence score. There is no confidence percentage on a recovery anywhere in the output. A related ruling keeps basis (whether an absence was verified or a vendor convention) as an audit annotation that must never adjust a confidence score, because a score adjustment is invisible in the output and compounds silently .
The naming rule¶
A field whose name contains negligence, defect, violation, fault, unfitness or indicator imports the conclusion into the input. Thirteen subro chunk fields failed that test outright when extraction moved to the vendor path, among them driver_unfitness_indicator, maintenance_violation and product_defect_indicated. The chunks change to match the catalogue, not the catalogue to match the chunks, because renaming a catalogue to existing chunk names would import the defect, and the conclusion-shaped names are also the ones no fact can populate without a conclusion .
The rule keeps recurring at the edges. The signal for a driver leaving the assigned route is route_deviation_described (a fact), not frolic_or_detour (a conclusion); frolic_or_detour was deleted from the field guide draft, and the defense that reads the fact was corrected so that a described deviation makes it indeterminate with a routine condition rather than firing on the description alone.
Why the archetype matters¶
The legacy subro projection layer is the archetype of getting this wrong. Twenty-three signals were computed inside the legacy extractors and wired into projections, which made a recovery determination at the reading stage. All twenty-three were permanently retired, and the facts behind about eight of them were restored. Restore facts, never projections.
The failure this prevents is not abstract. On the Morales worked example, a projection path returned a fault determination and a witness observation, both flagged span-verified, and both were model-constructed sentences rather than extracted text; a subrogation theory was confirmed on that evidence. A fabricated fact wearing a verification flag is the worst defect class the platform has, and it survived every review because nobody read the document behind the span.
A second consequence deserves its own line. A conclusion computed in two places will disagree. third_party_involved as an extracted boolean duplicates what the party roster already answers, and the two diverge on exactly the file where it matters. When several tools independently answer the same downstream question, the question belongs to one of them.
The retired-field decision register¶
Every legacy field gets a disposition and a reason in a decision register: restored, superseded by the roster, moved to the sensitive indicator detector, moved to the contract path, or out of scope. Retired fields are recorded with their replacements and their consumers, since several map one to many, and the reconciliation proves that the right consumer got the right field rather than only that a field exists.
The register exists because a whole field block (WC prior claims) was found missing by reading an architecture PDF; the migration from legacy tools to unified specs had never been diffed field by field. The standing fix is that a field present in a superseded spec and absent from its successor, with no recorded disposition, fails the seed. Without the register the same argument recurs in six months, which is what happened.
Two cautions on the numbers the register produces. Counting one architectural decision hundreds of times misleads: the evidence-envelope redesign inflated a 12-field gap into 372, and the redesign must be netted out before a number goes anywhere. And a conditional field must never score a correct null as a completeness miss, or every file without the rare fact is penalised.
Where the line sits on the vendor path¶
Moving extraction to a vendor sharpens the rule. A fact stated in prose is the extractor's ordinary work and needs verbatim passthrough. A conclusion drawn from prose is not the extractor's work at all. The vendor manifest answers what fields to return, and the no-derivation rule is one of the contract terms it must carry .
Contract facts are not FNOL facts either. Whether a waiver of subrogation exists lives in the contract, so it needs its own extraction path on the classification layer; what the FNOL can carry is only that a contract is referenced. A stated value and a valued value are likewise different facts and never share a field: what the insured says a vehicle is worth is a party statement with a span, and only a valuation vendor's figure can drive a settlement basis.
Open items¶
- The 13 conclusion-shaped subro chunk fields are recorded as retired; the register reconciling their replacements to consumers is referenced in the sources but its record id is not yet documented.
- The sensitive indicator signals WC-SI-013 and AUTO-SI-234 are held for their own track.