14 KiB
14 KiB
Apply Progress: OCR Ingest Integration
Current State
- Mode: Strict TDD
- Delivery: Feature-branch-chain, Units 1–4 complete; maintainer-approved
size:exceptionfor Unit 2 only - Completed tasks: 1.1, 1.2, 1.3, 1.4, 2.1, 2.2, 2.3, 2.4, 2.5
- Overall task progress: 9/30 complete
Unit 1: Migration
Implementation Summary
- Added
migrations/002_ocr_review.sqlwith durable OCR jobs, auditable page records, atomic review correction targets, and a partial unique index for non-terminal OCR identities. - Added focused schema contract coverage in
tests/catalog/migration-002.test.ts. - Preserved
migrations/001_knowledge_lifecycle.sqland all repository logic unchanged.
TDD Cycle Evidence
| Task | Test File | Layer | Safety Net | RED | GREEN | TRIANGULATE | REFACTOR |
|---|---|---|---|---|---|---|---|
| 1.2 | tests/catalog/migration-002.test.ts |
Schema contract | N/A (new files) | Valid RED: 4/4 failed with ENOENT because migration 002 did not exist. An earlier syntax-invalid run was corrected and is not counted as RED. |
4/4 passed with migration 002 present. | 4 independent schema behaviors cover jobs, pages, corrections, and pending identity uniqueness. | None needed; SQL already matches the minimal authoritative contract. |
Test Summary
- Total tests written: 4
- Total tests passing: 4
- Layers used: Schema contract (4)
- Approval tests: None — no existing production file was modified
- Pure functions created: 0
Work Unit Evidence
| Evidence | Result |
|---|---|
| Focused test command and exact result | npx --no-install tsx --test tests/catalog/migration-002.test.ts — exit 0; 4 tests, 4 passed, 0 failed. |
| Runtime harness command/scenario and exact result | N/A — the task forecast defines Unit 1 as schema-only, and the project has no PostgreSQL test harness or database-test dependency. Applying a live migration would require external database access outside this unit. |
| Rollback boundary | Delete migrations/002_ocr_review.sql and tests/catalog/migration-002.test.ts, revert task 1.2 to pending, and remove this Unit 1 progress/history entry. No existing migration or repository behavior must be reverted. |
Additional Validation
npm run check— exit 0.git diff --check— exit 0 after removing one trailing-space warning introduced in the history header.
Deviations
None — the migration follows the proposal, specifications, design, and closed OCR persistence contract.
Remaining Persistence Tasks
- 1.1 Complete repository-level RED coverage for duplicate pending ingestion, lease recovery, and review transitions in Unit 2.
- 1.2 Add migration 002 schema and pending OCR identity unique index.
- 1.3 Add catalog repository candidates and leases in Unit 2.
- 1.4 Refactor persistence code after Unit 2 reaches green.
Review Boundary
- Start: Migration 001 already supplies lifecycle states and parent catalog tables.
- End: Migration 002 and its focused schema contract test are green.
- Out of scope: Repository methods, dispatch behavior, runtime PostgreSQL migration, Unit 2, commits, pushes, PRs, deployment, and secrets.
- Rollback: Remove only the two Unit 1 implementation files and their progress metadata.
- Authored review budget: 225 additions/deletions, below the 400-line budget for this autonomous slice.
Unit 2: Repository
Implementation Summary
- Added pending OCR identity lookup, atomic queue claiming, expired-lease recovery, and guarded review transitions to
CatalogRepository. - Added a focused fake-pool repository harness that covers duplicate reuse, empty queues, remote identity preservation, and invalid transitions.
- Corrected the canonical test script so it executes both root-level and nested test suites.
TDD Cycle Evidence
| Task | Test File | Layer | Safety Net | RED | GREEN | TRIANGULATE | REFACTOR |
|---|---|---|---|---|---|---|---|
| 1.1 | tests/catalog/repository-ocr.test.ts |
Unit/fake pool | 4/4 existing tests passed | 0/6 with methods absent; then 5/6 while each completion gate was missing | 6/6 passed after task 1.3 | Found/not-found candidates, claimed/empty queue, expired lease, valid/invalid transition | Covered by 1.4 |
| 1.3 | tests/catalog/repository-ocr.test.ts |
Unit/fake pool | 4/4 existing tests passed | Tests preceded production code | 6/6 passed | Multiple inputs and state paths exercised | Shared OCR row mapping extracted only after green |
| 1.4 | tests/catalog/repository-ocr.test.ts |
Approval refactor | 6/6 before refactor | N/A — behavior-preserving approval baseline | 6/6 after refactor | Existing six cases preserved | Reusable aliased-column projection removed SQL duplication |
Work Unit Evidence
| Evidence | Result |
|---|---|
| Focused test command and exact result | NODE_ENV=test npx --no-install tsx --test tests/catalog/repository-ocr.test.ts — exit 0; 6 tests passed, 0 failed. NODE_ENV=test npx --no-install tsx --test tests/catalog/*.test.ts — exit 0; 10 tests passed, 0 failed. |
| Runtime harness command/scenario and exact result | NODE_ENV=test npx --no-install tsx --test tests/*.test.ts — exit 0; 25 root tests passed, 0 failed. Canonical npm test — exit 0; 35 total tests passed, 0 failed, proving the 25 root and 10 nested tests execute together. |
| Rollback boundary | Revert Unit 2 changes in src/modules/catalog/repository.ts, delete tests/catalog/repository-ocr.test.ts, and revert the package.json test script; Unit 1 migration remains intact. |
Validation and Boundary
- Corrective rerun for failed evidence revision
sha256:9433c0cd6628f36f4c96a27d41c6d429f4303d038ffe928f297367c274f21c2f; fresh evidence revisionsha256:a19119f84381447f54a6223e2a203acaf9d3ae08b6ae8f585e3bd52cd510909b. - The failed canonical script expanded only nested tests after nested suites appeared. The minimal correction is
NODE_ENV=test tsx --test tests/*.test.ts tests/**/*.test.ts. - Focused repository suite passed 6/6; all catalog suites passed 10/10; explicit root suites passed 25/25; canonical
npm testpassed 35/35;npm run checkandgit diff --checkexited 0. - Native Unit 2 accounting was 511 changed lines. The maintainer approved
size:exceptionexclusively for Unit 2; the corrective attempt has a separate 600-line remediation budget. All later units retain the normal 400-line budget. - No design deviations. No Unit 3 work was started.
Corrective TDD Evidence
| Correction | Safety Net / RED | GREEN | TRIANGULATE | REFACTOR |
|---|---|---|---|---|
| Canonical test discovery | Failed evidence revision proved canonical npm test executed only 10 nested tests while the explicit root command executed 25 tests. |
Updated only the package test argument list; canonical npm test passed 35/35. |
Direct root 25/25 and nested catalog 10/10 results sum to and identify the canonical 35/35 run. | None needed — one package-script argument was added. |
Unit 3: Extraction
Implementation Summary
- Added a Node 22 fixture spike proving
pdf-parsepage callbacks backed by bundled PDF.js2.0.550preserve ordered pages, including a blank page. - Added
parsePdfPageswith one-based page identity, raw native text, stable SHA-256 hashes, and fail-closed complete-page accounting. - Preserved textual ingestion as data-only, excluded
CMakeLists.txt, MDX, and shell lookalikes, and kept task 2.2 pending for Unit 4 detection/composition assertions.
TDD Cycle Evidence
| Task | Test File | Layer | Safety Net | RED | GREEN | TRIANGULATE | REFACTOR |
|---|---|---|---|---|---|---|---|
| 2.1 | tests/parsers/pdf-pages.test.ts; scripts/spike-pdfjs.ts |
Runtime spike | 20/20 lifecycle service tests passed | Focused suite failed 0/1 because parsePdfPages did not exist; the spike and parser implementation were still absent. |
Node 22.14.0 spike returned pages 1–3 exactly, including blank page 2. | Three pages prove non-empty, blank, and later-page ordering paths. | Replaced the over-budget direct dependency with the design-approved existing pdf-parse pagerender path; spike remained green. |
| 2.3 | tests/parsers/pdf-pages.test.ts |
Unit/integration fixture | 20/20 lifecycle service tests passed | Focused suite failed 0/1 on the missing export before production code. | Final focused run passed 5/5 after implementation. | Covered page hashes/blank page, native PDF composition, non-PDF rejection, data-only text parsing, and unsupported lookalikes. | Focused suite remained 5/5 after consolidating on the existing parser dependency. |
Work Unit Evidence
| Evidence | Result |
|---|---|
| Focused test command and exact result | NODE_ENV=test npx --no-install tsx --test tests/parsers/pdf-pages.test.ts — exit 0; 5 passed, 0 failed. |
| Runtime harness command/scenario and exact result | NODE_ENV=test npx --no-install tsx scripts/spike-pdfjs.ts — exit 0 on Node 22.14.0; PDF.js 2.0.550 returned three ordered pages with page 2 blank. |
| Rollback boundary | Revert src/modules/parsers/parser-registry.ts; delete scripts/spike-pdfjs.ts, tests/parsers/pdf-pages.test.ts, and tests/fixtures/ocr/native-three-pages.pdf; revert tasks 2.1/2.3 and this Unit 3 progress/history entry. Units 1–2 remain intact. |
Validation and Boundary
- Canonical
npm testpassed 40/40;npm run checkandgit diff --checkexited 0. - Native authored slice before progress/history metadata: 225 additions/deletions; no package or lockfile delta remains.
- Start: Unit 2 persistence is complete. End: per-page native extraction and parser threat routing are green.
- Out of scope: detection thresholds, OCR selection, composition, risks, Unit 4 tasks 2.4/2.5, commits, pushes, PRs, deployment, and restricted paths.
- No specification deviation. The design-authorized
pdf-parsepagerender fallback was selected only after its Node 22 fixture spike passed.
Unit 4: Detection and Composition
Implementation Summary
- Added pure per-page native detection under
pdf-detection-v1, exact OCR page selection, verified-blank classification, and fail-closed OCR quality gates. - Added deterministic candidate composition using exactly one source per page, ordered OCR lines, blank omission, canonical hashes, and unchanged risk tokens prioritized for review.
- Completed the remaining task 2.2 scenarios by combining the Unit 3 data-only/parser threat coverage with Unit 4 mixed-PDF, detection, blank, hash, and risk RED assertions.
TDD Cycle Evidence
| Task | Test File | Layer | Safety Net | RED | GREEN | TRIANGULATE | REFACTOR |
|---|---|---|---|---|---|---|---|
| 2.2 | tests/ocr/detection.test.ts; tests/parsers/pdf-pages.test.ts |
Unit/integration fixture | Unit 3 parser suite passed 5/5 before edits. | Focused Unit 4 suite failed before production with ERR_MODULE_NOT_FOUND for src/modules/ocr/detection.js. |
Focused suite passed 5/5 after detection and composition were added. | Real generated mixed PDF selected pages 2 and 3; .txt/.md stayed data-only and CMake/MDX/shell lookalikes stayed unsupported. |
None needed; routing remains separated between parser support and the pure PDF OCR planner. |
| 2.4 | tests/ocr/detection.test.ts |
Unit | N/A (new production file) | Threshold, mixed-page, blank, and quality assertions preceded production code. | Exact native, blank, and OCR quality boundaries passed in the 5/5 focused suite. | Every native and OCR threshold has a passing boundary plus a failing neighbor; visible ink with empty OCR is blocked. | Pure functions and named policy version were present at green; final focused suite remained 5/5. |
| 2.5 | tests/ocr/detection.test.ts |
Unit | N/A (new production file) | Ordered composition, blank omission, canonical hash, audit-source, and risk assertions preceded production code. | Deterministic composition and risk prioritization passed in the 5/5 focused suite. | Reordered pages, tied OCR boxes, blank middle page, repeated token, and four known ambiguous tokens exercise distinct paths. | Reused shared canonical JSON and SHA-256 utilities; final focused suite and type check remained green. |
Test Summary
- Tests: 5 written and focused-passing; 45 canonical (unit 4, integration fixture 1)
- Approval tests: None — new production files; pure functions created: 6
Work Unit Evidence
| Evidence | Result |
|---|---|
| Focused test command and exact result | NODE_ENV=test npx --no-install tsx --test tests/ocr/detection.test.ts — exit 0; 5 tests passed, 0 failed. |
| Runtime harness command/scenario and exact result | NODE_ENV=test npx --no-install tsx --test --test-name-pattern "parsed mixed PDF" tests/ocr/detection.test.ts — exit 0; 1 test passed, 0 failed. The harness generated a three-page PDF, parsed it through the Node 22 pdf-parse page callback, and selected only blank page 2 and insufficient page 3 for OCR. |
| Rollback boundary | Delete src/modules/ocr/detection.ts, src/modules/ocr/composition.ts, and tests/ocr/detection.test.ts; revert tasks 2.2/2.4/2.5 and this Unit 4 progress/history entry. Units 1–3 and their parser threat coverage remain intact. |
Validation and Boundary
- Canonical
npm testpassed 45/45;npm run check, trackedgit diff --check, and no-index whitespace checks for all three new files exited 0. - Authored Unit 4 code and tests total 322 additions and 0 deletions, below the 400-line budget with no size exception.
- Start: Unit 3 page extraction and parser threat routing are complete. End: detection, blank/quality gates, deterministic composition, hashes, and risk tokens are green.
- Out of scope: OCR service, client, dispatcher, ingest routing, review APIs, commits, pushes, PRs, deployment, and restricted paths.
- No specification or design deviation.