Skip to content

The intake boundary


title: The Intake Boundary description: Where an external IDP vendor stops and Adjust360 begins: the three-layer split, the canonical normalizer, the status rulings that keep an unanswered field from becoming a fact, and the findings from the first vendor schema review. kind: authored status: draft source_memory: [idp-vendor-boundary, fnol-extract-core, document-classification, subro-product-catalog] source_date: 2026-09-30 tags: [IDP, Reducto, idp_normalizer, external_idp_manifest, fact_status, not_returned, AUTO, contracts]


What it is

The intake boundary is the line between document extraction, which Adjust360 buys, and interpretation, which it sells. ElevateNow is not in the IDP business. Reducto is the vendor today and a Databricks IDP is possible; the native extractor stays as a reference implementation and a measurement baseline, not a product line. The boundary is built so that no downstream tool can tell which path read the documents .

The three-layer split

Layer 1, what to extract, comes from the database. When the harness runs with --idp reducto or --idp databricks it loads EXTRACT-CORE and the per-LOB spec, so a new field in a spec is automatically expected with no code change. Before this the harness trusted whatever the vendor file happened to contain .

Layer 2, the vendor adaptor, is legitimate vendor-specific code, because each vendor wants its own schema shape (Reducto needs JSON Schema with string types). It must read field definitions from the database rather than hardcoding them. It lives in build_auto_schema.py and is outside the harness scope .

Layer 3, the output contract, normalises vendor output into the same core_facts plus claims_by_lob shape the native path produces. The normalizer walks the spec's field list, not the vendor's output keys. This is the layer that makes vendor parity a property of the system rather than a promise: identical output contract including classification_source whether the IDP or the native path ran, and no downstream tool may branch on which. A carrier without an IDP loses throughput, not capability .

The vendor field contract itself is generated: generate_idp_manifest.py reads EXTRACT-CORE field_definitions and emits the machine-readable manifest and the BYO integration guide, with a --diff drift gate in CI. external_idp_manifest is never edited directly, and no ElevateNow component name appears in the BYO template .

Reducto adaptor conventions

The adaptor makes all types strings because Reducto scores strings reliably and mixed types it does not; booleans become the enum ["true","false",""]; an empty string means not stated because the vendor never returns null. The evidence envelope is dropped for most fields since the vendor supplies its own citations, with the thirteen subro_facts fields as the exception, each carrying status, verbatim, speaker, recorded_by, document_id, page, character offsets and confidence.

The four rulings on layers 1 and 3

An absent key is not_returned, never not_stated. not_stated means the document was read and the fact is not there, a fact a rule may act on. An absent key means nobody looked, which no rule may act on. not_extractable was also rejected for absent keys because that is a vendor's own claim about a field it could not read, and synthesising it puts words in their mouth. not_returned resolves to indeterminate everywhere, exactly as a missing fact does in the detector .

The envelope stays on every field. Bare values on the external path while the native path carries status would reintroduce the divergence layer 3 exists to close, and a bare empty string is not self-describing where an empty envelope is.

A basis of vendor_convention distinguishes a core or LOB null (the vendor's empty-string convention, an unverified absence) from a subro not_stated backed by a citation (a verified absence). The implementation made this a key that survives into the output and is testable. basis is an audit annotation and never touches confidence scoring; a rule acting on an absence should be able to see whether that absence was verified, surfaced, not scored.

Provenance records more than the spec version: vendor, vendor_model, normalizer_version, spec_id and spec_version, lob_spec_id and lob_spec_version, and field_counts across the four statuses. Two runs are only comparable if you can see they normalised against the same spec and the same vendor model.

The status vocabulary as implemented

Status Meaning on the external path
stated Vendor returned a value
not_stated, no basis Subro field; vendor searched and confirmed absent with citation; a verified absence
not_stated, basis vendor_convention Core or LOB field; vendor returned null or empty per convention; unverified
not_extractable Vendor attempted and could not resolve the field (garbled table, scan quality)
not_returned Spec expected the field; vendor JSON has no key at all

The normalizer (idp_normalizer.py: load_idp_specs, ReductoNormalizer, a DatabricksNormalizer stub raising NotImplementedError, a get_normalizer factory) ships with six negative-control tests. The fixtures were all positive cases, so a normalizer that collapsed not_stated, not_extractable and not_returned into one value would have passed every one of them. NC-5 is the guard that matters: three inputs simultaneously must produce exactly three distinct statuses .

On the Martinez smoke test the counts were 119 stated, 77 not_stated, 0 not_extractable and 16 not_returned against EXTRACT-CORE v1 and the AUTO spec v2. The 16 is the first measurement anyone has of what a vendor did not answer. field_counts.not_returned is a completeness measurement now and a procurement metric later, once a second vendor runs on the same document; it is tracked with the spec version beside it, because a rising count after a spec change means the spec grew, not that the vendor got worse.

Why absent keys are not defaulted

If any consumer reads a non-stated status as absence, a field nobody answered becomes a confirmed negative and a theory fails to fire for a reason never reported. The ruling is that every consumer branching on status is traced and its behaviour on an unrecognised value reported; default behaviour on an unknown status matters as much as explicit handling. SubrogationScreenerGKR needed a code-level check rather than an assumption. The same principle covers the harness: a fallback to a flat copy when the database is unavailable is a silent degradation into the uncanonical shape layer 3 exists to eliminate, and must fail closed with the reason instead.

Findings from the first vendor schema review

The team's Reducto schema default.auto.claim.reducto.v1 was reviewed against the engine's contract and the findings are recorded so the same shapes are not built twice :

  • a present/absent/unknown status vocabulary instead of the four canonical statuses;
  • duplicated blocks, core_facts against auto_claim, defining the same facts twice;
  • inconsistent envelopes across fields;
  • two role enums, neither matching the engine's;
  • subro fields missing: waiver, coverage_responds, employment_relationship, for_hire, weight, component fields, identity_status and the PD-1 fields.

A migration document covering all 325 nodes was required before the schema could be synced.

Three rulings govern that sync. No MANUAL_CONTRACTS in code: the contract comes from records. The signal is route_deviation_described, a fact the carrier can state, not frolic_or_detour, which is a conclusion the defense record draws; a described deviation makes the defense indeterminate with a routine condition and it fires only on deviation_personal_purpose_stated true. And no unknown enum value anywhere: a carrier states a fact or states its absence, and everything else is not_returned .

Open items

vendor_model read "unknown" on the smoke test; provenance that cannot name the model makes two runs incomparable, which is the reason the provenance ruling exists. Whether the 16 not_returned fields are a schema gap, a vendor gap or fields no AUTO document would carry is not yet established. The Databricks normalizer is a stub. Source citation from the carrier on the structured path is deferred past the POC and will be added without changing the carrier_structured payload schema.