ADR-064Proposed2026-09-11
Extends ADR-033 · consumes ADR-039

Knowing the form,
not the case.

A vision model reads a scanned page and tells you what kind of document it is — a shortcut for the person triaging the Unfiled queue. It does not file it. Filing to a case is still the QR's job.

For Zareef (product / sequencing) & Matt (the triage surface). Where this sits among the ingestion phases we already outlined in ADR-033.

Two questions about an inbound scan

ADR-033's ingestion ladder already answers “which case?”. This ADR answers a second, separate question — “what kind of document is this?” — that the ladder named but never built.

A
QR / tracked-code on the page ADR-039 D7/D8
The only true auto-file. Resolves both the document type and the case number, deterministically. Works only on forms we print.
B
Printed case-number footer → OCR  ·  name + DOB match
Resolves the case when there's no QR. Still ADR-033 / ADR-039 territory.
★
Form-type recognition This ADR — new
Tells the human what the document is — even when the case is unknown. Pre-labels the queue so a director confirms at a glance instead of opening every page. A shortcut, not a filing decision.
C
Manual — the Unfiled Documents queue
Today: a director opens and reads every page. This ADR makes that queue smarter, not obsolete.
This is a triage accelerator, not auto-file. A wrong suggestion is a wrong label the human corrects in one click — never a silent misfile. Auto-file needs the case number, and only the QR delivers that.

What actually comes back from the office

Forms printed from the system, hand-annotated, wet-signed, then run through a sheet-feed scanner — which emits one PDF of many forms with no page breaks, sometimes rotated or upside-down.

So the job is two things at once: split the concatenated scan into separate forms, and identify each one. No OCR of the handwriting is needed — the form's fixed structure (masthead, title, section headings) is enough. And no change to the blank forms: no stamps, barcodes, or hidden text.

A spike on Kearney's real forms

A throwaway harness ran against the actual Kearney corpus — 78 catalogued forms, real filled documents, and a genuine scanned bundle. Findings, in order:

1
Recognition works without OCR. 94% top-1 on 34 real filled single-form documents — best on the flattened/scanned ones, which is the realistic case.
2
It survives a real scan, upside-down. A genuine Canon scan of a 23-page bundle arrived with every page rotated 180°. The harness detected and corrected orientation per page, then hit 95% instance-level — verified by hand, page by page.
3
Bundle splitting works. It recovered 20 form instances from the 23 pages, correctly merging the two 2-page forms (each with a no-masthead continuation page) into single documents.
4
The failures are confident mistakes on look-alike variants — a fingerprint-consent form read as a DNA-consent form; a KFS form read as its general twin. And they're model-independent: the expensive and cheap models each miss pages the other gets right. So one model's confidence alone is not safe to file on — which is exactly why this stays triage, not auto-file.
95%
instance-level accuracy on a real, upside-down Canon scan (human-verified)
20 / 23
form instances correctly split & identified, incl. both 2-page forms
180°
every page arrived upside-down — detected & corrected from the image, not the PDF metadata
Why not trust the PDF's own rotation flag? The scan reported “not rotated” on every page while being visibly upside-down. Real scanners bake the rotation into the pixels. So orientation is judged from the rendered image, always.

Fix it once, at onboarding — not on every document

The look-alike mistakes weren't a limit of the AI. They were an information gap: at recognition time the model only had a generic description of each form, with no explicit sense of how two near-identical forms differ.

So when we onboard a new form, we generate a one-time discriminative fingerprint: its exact title, the sections unique to it, and “here's how to tell it apart from its look-alike sibling.” Feed a form's fingerprint to the recognizer — only for the family it belongs to — and the confusion disappears.

0
confident-wrong misfiles introduced — the scoped design fixed the look-alike errors with no harmful regression across the 60-page test set
once
per form, at onboarding — vs. recognition running on every document, forever
✓ transfers
a fingerprint written by the strong model let the cheap model match it
One thing the testing taught us: inject a form's fingerprint on every page and it over-corrects — it fixes one look-alike but can push a sibling wrong. Inject it only when the page has already been recognised as that family (a cheap two-pass, firing on ~6 of 60 pages), and the fix lands with zero confident-wrong misfiles. That scoping is the design, not a detail.
Author the fingerprint once with the smart model at onboarding; run recognition forever with the cheap model. You get top-tier accuracy at budget-tier recurring cost. This is genuinely different data than we capture today — how to tell forms apart, not just what fields they have.

Priced from real token usage

Every API call reports its exact token count, so these are measured, not estimated. Two recognition strategies were compared; the cheap one won.

ApproachOpus (strong)Sonnet (cheap)Verdict
Recognition — text catalogue (1 image + cached form list)$0.027 / pg$0.010 / pgchosen
Recognition — reference-image reranker$0.137 / pg$0.047 / pgrejected — ~5× cost, and its confidence isn't trustworthy
Onboarding fingerprint~1 call per look-alike family, one timethe leverage point

At roughly 1,000 pages, cheap recognition is about $10–$27. Running two models on every document to cross-check was rejected on cost — kept only as a rare escalation for the handful of known look-alike families. The strong-onboard / cheap-recognize split is the recommended posture.

What we're not doing, and why

✕ Stamp a code on the blank forms
That's the QR path (ADR-039) and only works on forms we print. This recognizer must handle forms as they already exist.
✕ Match on the PDF's form fields (no AI)
Only works when the scan still has a fillable-field layer — most don't — and it collides across brands. Kept only as a free baseline.
✕ Reference-image reranker as primary
~5× the cost and, worse, its confidence reads high even when wrong — so it can't drive a triage display.
✕ Auto-file on one model's confidence
Confident mistakes on look-alikes are model-independent. No single model's confidence is safe to file on — hence triage, not auto-file.
✕ Two models on every document
Would catch the mistakes, but the recurring cost is unacceptable. Reserved as a rare, targeted escalation only.
✕ Full OCR of the answers to ID the form
Unnecessary — the fixed structure is enough and cheaper. Reading the filled values stays ADR-033's job.

What still needs deciding

? PHI & data residency (gates deployment)
The spike sent real PHI pages to the API on Philip's machine, by consent. Production must settle the boundary — in-region Bedrock (ca-central-1), a BAA, or redaction — before any customer use.
? Beyond Kearney
The no-regression and scoped two-pass runs are measured, but on one tenant's 60 pages. The numbers need to hold on other tenants' forms and more real scans before this is a general claim.
? Collision map
How we decide which forms are mutually confusable, so we write fingerprints only where needed — not for all 80 forms.
? Triage UX (for Matt)
How the suggested type — and a top-2 pick — surface on the Unfiled queue so the shortcut is genuinely faster than opening the document, and how a corrected suggestion is captured.