Seed bundle and tenant install
What it is¶
A tenant is a carrier's own installation of Adjust360: a clean knowledge database seeded from the platform bundle, a separate database where processed claims land, and the four runtime components that read them. The data package that produces the knowledge database is data_packager/, a self-contained folder that a developer can zip and integrate into a deploy pipeline . Containerization is out of scope for this package; it belongs to a separate deploy package team that packages the tools, workbenches, Curation Studio and their dependencies. This page covers the data package and the product rules a tenant install must obey.
The two-database split¶
Today one Atlas cluster holds both elevatenow_gkr and fnol_claim_intel. A tenant separates them. The GKR database holds only Adjust360 knowledge: the reference collections and the platform chunks. The claims database holds what the pipeline produces when it processes a claim. The concern that forced the split was that the Curation Studio's knowledge chunks mix underwriting and claims content, which is a problem the moment the store is deployed into a carrier.
The first piece of work was therefore an audit of which GKR collections the application needs at all, since many were of unknown necessity. Directive 1 (2026-09-14) was a read-only collection-level audit of elevatenow_gkr producing GKR_COLLECTION_MANIFEST.md and adjust360_gkr_manifest.json, classifying every collection, every table_type and every chunk_type. The manifest is the generated input to the seed bundle and the build script .
The four runtime components¶
Four components ship, not two: the Claims Workbench pointing at the claims database for the adjuster persona; the Curation Studio pointing at the GKR database for the curator persona; the pipeline and tools; and a command-line utility to configure, modify and observe. Because the carrier operates the Curation Studio in this package, studio authentication and tenant scoping are a release blocker for the package rather than a later concern. See Curation Studio.
The data_packager¶
The folder has a fixed shape. README.md at root is the quick-start; docs/HANDOVER.md is the full deployment spec covering arguments, exit codes, failure modes and the two open boundary questions; scripts/ holds a360_export.py, a360_load.py, a360_verify.py and a360_bundle_lib.py; data/ holds adjust360_gkr_manifest.json, carrier_layer_note.json and recommended_indexes.json. A typical first command runs from scripts/:
python3 a360_export.py --manifest ../data/adjust360_gkr_manifest.json --source "..." --out /tmp/bundle
The source argument is a connection string supplied at run time; it never appears in code or in the bundle.
The four verdicts¶
Every collection and type in the manifest carries one of four verdicts.
| Verdict | Meaning |
|---|---|
SHIP |
A runtime tool reads it during a claim decision; it is seeded into the tenant |
STUDIO |
Only the Curation Studio reads it; seeded for the curator surface |
INIT |
Created empty with its indexes, never seeded; tenant-local state such as curation_audit, extraction_results and caches |
DROP |
Nothing reads it, proven by grep; not shipped |
INIT exists for a governance reason. Copying ElevateNow's own 424-entry curation_audit into a carrier tenant would open the carrier's governance trail with our internal history, so operational history is initialised empty rather than copied.
The closed manifest holds 69 chunk types, 1,970 knowledge_chunks and 815 reference_data records, 3,009 documents in the base package with zero remainder. Ontology collections (chunk_ontology_edges 12,332, ontology_entity_edges 1,554, ontology_entities 81) are excluded from the base bundle as studio-optional and available as a patch seed. Of the 3,009 documents, roughly 350 are read by a runtime tool during a claim decision and the remaining 2,659 are curation surface. That ratio is the honest shape of the product and belongs in this manual rather than being smoothed over.
The allowlist principle¶
The chunk_type filter is an allowlist, not a denylist, on the principle that absent configuration means excluded. A pre-seed distinct() check confirms that no type exists in the source outside the allowlist, so a type added to the source after the manifest closed cannot slip into a bundle unnoticed.
No carrier records ship¶
The export filter is carrier_id: null by construction. That excludes the 152 CARRIER_B_MM demo chunks, 32 other carrier chunks and the 13 CARRIER-AXA-AUTO operational directives. Those demo chunks are deliberately not deleted from elevatenow_gkr: deleting would write to the live source, and it would destroy the only worked example of supersession that exists. Bundles ship active or approved records only, never draft .
The verifier's five checks¶
a360_verify.py runs five checks against the loaded tenant: record counts; per-collection checksums, re-exported byte-identically from the target; cross-reference resolution across subro_spec, coverage_spec, compensability_spec and reserve_spec; indexes; and carrier absence. Check 5 is the tenant boundary and the only check that proves an absence rather than a presence. It takes a --permitted-carrier argument rather than being removed once carrier overlays are legitimately loaded, so the boundary check survives the day a carrier's own records arrive.
Versioning and checksums¶
The accepted bundle is version 2026-09-14-6ea5db45 with manifest sha256 24f12e7b. The bundle records the manifest's own sha256, so a bundle built from a different manifest is detectable. Acceptance evidence: two exports produced identical checksums; a date field round-tripped as datetime rather than str; a corrupted bundle was caught before any target connection was opened; and a non-empty target was refused.
The package was accepted after two review rounds. Round one fixed three blockers: the four FNOLExtractorGKR record types missing from the filter, knowledge_chunks shipping all 2,288 including another carrier's demo data, and operational history being copied rather than initialised empty. Round two fixed five defects: export reading from a replica secondary rather than the primary, the target URI passed as subprocess argv where ps exposes it, every failure exiting code 1, the output directory never cleaned so stale files poisoned the total checksum, and the bundle not recording the manifest's sha256.
Text indexes¶
Two text indexes, idx_text_search on knowledge_chunks and on knowledge_templates, cannot be recreated through the driver because index_information() exposes MongoDB's internal _fts and _ftsx keys. They serve content search only; no claim decision path reads them. The recreate commands ship as text_indexes.mongosh inside the bundle.
No upgrade path yet¶
The loader refuses a non-empty target, so there is no upgrade path. Changing the corpus today means a fresh database. Upgrade in place needs its own design and is named here as a gap rather than hidden.
The parity harness¶
Directive 3 established a parity harness that runs the Morales and Martinez worked examples against the source and the tenant databases with the connection string as the only variable. It diffs governance blocks, citation lists, three-state results, refusals with their reasons and per-tool record counts. Verdict-level parity is explicitly not parity: a missing record produces a clean refusal, and a clean refusal renders as a perfectly reasonable brief. The harness also captures load time, database size and per-claim runtime for the handover sizing section.
Rules a tenant install never bends¶
Three rules from the repository's standing rules apply directly to installation .
No credentials anywhere in code or config; environment variables only. The box test found a live connection string with a password printed in a delivery document, which required rotation, a purge of the repository and its history, and a pre-commit check for connection strings. Rotation is the owner's action. See How Work Is Run.
No direct database access from a tool. The API sits behind the harness, and the workbench scope agreed for the receipt view, inline brief and PDF explicitly excludes direct database access.
A tenant is written only by the bundle loader. Nothing else, no tool, no seed script run by hand, no studio path, writes to a tenant database. In the box test this was stated for the AXA tenant; it is the product rule for every tenant.
Open items¶
- Two boundary questions with the deploy team, recorded in
HANDOVER.md: whether they provision the empty database and credentials or the loader does, and whether the seed runs once at install or on every deploy. - Upgrade in place has no design.
- A tenant re-export (BT-R3) is owed before any output reaches the first carrier.