20 KiB
20 KiB
Apply Progress: OCR Ingest Integration
Current State
- Mode: Strict TDD
- Delivery: Feature-branch-chain, Units 1–6 complete; maintainer-approved
size:exceptionfor Unit 2 only - Completed tasks: 1.1, 1.2, 1.3, 1.4, 2.1, 2.2, 2.3, 2.4, 2.5, 3.1, 3.2, 3.3, 3.4
- Overall task progress: 13/30 complete
Unit 1: Migration
Implementation Summary
- Added
migrations/002_ocr_review.sqlwith durable OCR jobs, auditable page records, atomic review correction targets, and a partial unique index for non-terminal OCR identities. - Added focused schema contract coverage in
tests/catalog/migration-002.test.ts. - Preserved
migrations/001_knowledge_lifecycle.sqland all repository logic unchanged.
TDD Cycle Evidence
| Task | Test File | Layer | Safety Net | RED | GREEN | TRIANGULATE | REFACTOR |
|---|---|---|---|---|---|---|---|
| 1.2 | tests/catalog/migration-002.test.ts |
Schema contract | N/A (new files) | Valid RED: 4/4 failed with ENOENT because migration 002 did not exist. An earlier syntax-invalid run was corrected and is not counted as RED. |
4/4 passed with migration 002 present. | 4 independent schema behaviors cover jobs, pages, corrections, and pending identity uniqueness. | None needed; SQL already matches the minimal authoritative contract. |
Test Summary
- Total tests written: 4
- Total tests passing: 4
- Layers used: Schema contract (4)
- Approval tests: None — no existing production file was modified
- Pure functions created: 0
Work Unit Evidence
| Evidence | Result |
|---|---|
| Focused test command and exact result | npx --no-install tsx --test tests/catalog/migration-002.test.ts — exit 0; 4 tests, 4 passed, 0 failed. |
| Runtime harness command/scenario and exact result | N/A — the task forecast defines Unit 1 as schema-only, and the project has no PostgreSQL test harness or database-test dependency. Applying a live migration would require external database access outside this unit. |
| Rollback boundary | Delete migrations/002_ocr_review.sql and tests/catalog/migration-002.test.ts, revert task 1.2 to pending, and remove this Unit 1 progress/history entry. No existing migration or repository behavior must be reverted. |
Additional Validation
npm run check— exit 0.git diff --check— exit 0 after removing one trailing-space warning introduced in the history header.
Deviations
None — the migration follows the proposal, specifications, design, and closed OCR persistence contract.
Remaining Persistence Tasks
- 1.1 Complete repository-level RED coverage for duplicate pending ingestion, lease recovery, and review transitions in Unit 2.
- 1.2 Add migration 002 schema and pending OCR identity unique index.
- 1.3 Add catalog repository candidates and leases in Unit 2.
- 1.4 Refactor persistence code after Unit 2 reaches green.
Review Boundary
- Start: Migration 001 already supplies lifecycle states and parent catalog tables.
- End: Migration 002 and its focused schema contract test are green.
- Out of scope: Repository methods, dispatch behavior, runtime PostgreSQL migration, Unit 2, commits, pushes, PRs, deployment, and secrets.
- Rollback: Remove only the two Unit 1 implementation files and their progress metadata.
- Authored review budget: 225 additions/deletions, below the 400-line budget for this autonomous slice.
Unit 2: Repository
Implementation Summary
- Added pending OCR identity lookup, atomic queue claiming, expired-lease recovery, and guarded review transitions to
CatalogRepository. - Added a focused fake-pool repository harness that covers duplicate reuse, empty queues, remote identity preservation, and invalid transitions.
- Corrected the canonical test script so it executes both root-level and nested test suites.
TDD Cycle Evidence
| Task | Test File | Layer | Safety Net | RED | GREEN | TRIANGULATE | REFACTOR |
|---|---|---|---|---|---|---|---|
| 1.1 | tests/catalog/repository-ocr.test.ts |
Unit/fake pool | 4/4 existing tests passed | 0/6 with methods absent; then 5/6 while each completion gate was missing | 6/6 passed after task 1.3 | Found/not-found candidates, claimed/empty queue, expired lease, valid/invalid transition | Covered by 1.4 |
| 1.3 | tests/catalog/repository-ocr.test.ts |
Unit/fake pool | 4/4 existing tests passed | Tests preceded production code | 6/6 passed | Multiple inputs and state paths exercised | Shared OCR row mapping extracted only after green |
| 1.4 | tests/catalog/repository-ocr.test.ts |
Approval refactor | 6/6 before refactor | N/A — behavior-preserving approval baseline | 6/6 after refactor | Existing six cases preserved | Reusable aliased-column projection removed SQL duplication |
Work Unit Evidence
| Evidence | Result |
|---|---|
| Focused test command and exact result | NODE_ENV=test npx --no-install tsx --test tests/catalog/repository-ocr.test.ts — exit 0; 6 tests passed, 0 failed. NODE_ENV=test npx --no-install tsx --test tests/catalog/*.test.ts — exit 0; 10 tests passed, 0 failed. |
| Runtime harness command/scenario and exact result | NODE_ENV=test npx --no-install tsx --test tests/*.test.ts — exit 0; 25 root tests passed, 0 failed. Canonical npm test — exit 0; 35 total tests passed, 0 failed, proving the 25 root and 10 nested tests execute together. |
| Rollback boundary | Revert Unit 2 changes in src/modules/catalog/repository.ts, delete tests/catalog/repository-ocr.test.ts, and revert the package.json test script; Unit 1 migration remains intact. |
Validation and Boundary
- Corrective rerun for failed evidence revision
sha256:9433c0cd6628f36f4c96a27d41c6d429f4303d038ffe928f297367c274f21c2f; fresh evidence revisionsha256:a19119f84381447f54a6223e2a203acaf9d3ae08b6ae8f585e3bd52cd510909b. - The failed canonical script expanded only nested tests after nested suites appeared. The minimal correction is
NODE_ENV=test tsx --test tests/*.test.ts tests/**/*.test.ts. - Focused repository suite passed 6/6; all catalog suites passed 10/10; explicit root suites passed 25/25; canonical
npm testpassed 35/35;npm run checkandgit diff --checkexited 0. - Native Unit 2 accounting was 511 changed lines. The maintainer approved
size:exceptionexclusively for Unit 2; the corrective attempt has a separate 600-line remediation budget. All later units retain the normal 400-line budget. - No design deviations. No Unit 3 work was started.
Corrective TDD Evidence
| Correction | Safety Net / RED | GREEN | TRIANGULATE | REFACTOR |
|---|---|---|---|---|
| Canonical test discovery | Failed evidence revision proved canonical npm test executed only 10 nested tests while the explicit root command executed 25 tests. |
Updated only the package test argument list; canonical npm test passed 35/35. |
Direct root 25/25 and nested catalog 10/10 results sum to and identify the canonical 35/35 run. | None needed — one package-script argument was added. |
Unit 3: Extraction
Implementation Summary
- Added a Node 22 fixture spike proving
pdf-parsepage callbacks backed by bundled PDF.js2.0.550preserve ordered pages, including a blank page. - Added
parsePdfPageswith one-based page identity, raw native text, stable SHA-256 hashes, and fail-closed complete-page accounting. - Preserved textual ingestion as data-only, excluded
CMakeLists.txt, MDX, and shell lookalikes, and kept task 2.2 pending for Unit 4 detection/composition assertions.
TDD Cycle Evidence
| Task | Test File | Layer | Safety Net | RED | GREEN | TRIANGULATE | REFACTOR |
|---|---|---|---|---|---|---|---|
| 2.1 | tests/parsers/pdf-pages.test.ts; scripts/spike-pdfjs.ts |
Runtime spike | 20/20 lifecycle service tests passed | Focused suite failed 0/1 because parsePdfPages did not exist; the spike and parser implementation were still absent. |
Node 22.14.0 spike returned pages 1–3 exactly, including blank page 2. | Three pages prove non-empty, blank, and later-page ordering paths. | Replaced the over-budget direct dependency with the design-approved existing pdf-parse pagerender path; spike remained green. |
| 2.3 | tests/parsers/pdf-pages.test.ts |
Unit/integration fixture | 20/20 lifecycle service tests passed | Focused suite failed 0/1 on the missing export before production code. | Final focused run passed 5/5 after implementation. | Covered page hashes/blank page, native PDF composition, non-PDF rejection, data-only text parsing, and unsupported lookalikes. | Focused suite remained 5/5 after consolidating on the existing parser dependency. |
Work Unit Evidence
| Evidence | Result |
|---|---|
| Focused test command and exact result | NODE_ENV=test npx --no-install tsx --test tests/parsers/pdf-pages.test.ts — exit 0; 5 passed, 0 failed. |
| Runtime harness command/scenario and exact result | NODE_ENV=test npx --no-install tsx scripts/spike-pdfjs.ts — exit 0 on Node 22.14.0; PDF.js 2.0.550 returned three ordered pages with page 2 blank. |
| Rollback boundary | Revert src/modules/parsers/parser-registry.ts; delete scripts/spike-pdfjs.ts, tests/parsers/pdf-pages.test.ts, and tests/fixtures/ocr/native-three-pages.pdf; revert tasks 2.1/2.3 and this Unit 3 progress/history entry. Units 1–2 remain intact. |
Validation and Boundary
- Canonical
npm testpassed 40/40;npm run checkandgit diff --checkexited 0. - Native authored slice before progress/history metadata: 225 additions/deletions; no package or lockfile delta remains.
- Start: Unit 2 persistence is complete. End: per-page native extraction and parser threat routing are green.
- Out of scope: detection thresholds, OCR selection, composition, risks, Unit 4 tasks 2.4/2.5, commits, pushes, PRs, deployment, and restricted paths.
- No specification deviation. The design-authorized
pdf-parsepagerender fallback was selected only after its Node 22 fixture spike passed.
Unit 4: Detection and Composition
Implementation Summary
- Added pure per-page native detection under
pdf-detection-v1, exact OCR page selection, verified-blank classification, and fail-closed OCR quality gates. - Added deterministic candidate composition using exactly one source per page, ordered OCR lines, blank omission, canonical hashes, and unchanged risk tokens prioritized for review.
- Completed the remaining task 2.2 scenarios by combining the Unit 3 data-only/parser threat coverage with Unit 4 mixed-PDF, detection, blank, hash, and risk RED assertions.
TDD Cycle Evidence
| Task | Test File | Layer | Safety Net | RED | GREEN | TRIANGULATE | REFACTOR |
|---|---|---|---|---|---|---|---|
| 2.2 | tests/ocr/detection.test.ts; tests/parsers/pdf-pages.test.ts |
Unit/integration fixture | Unit 3 parser suite passed 5/5 before edits. | Focused Unit 4 suite failed before production with ERR_MODULE_NOT_FOUND for src/modules/ocr/detection.js. |
Focused suite passed 5/5 after detection and composition were added. | Real generated mixed PDF selected pages 2 and 3; .txt/.md stayed data-only and CMake/MDX/shell lookalikes stayed unsupported. |
None needed; routing remains separated between parser support and the pure PDF OCR planner. |
| 2.4 | tests/ocr/detection.test.ts |
Unit | N/A (new production file) | Threshold, mixed-page, blank, and quality assertions preceded production code. | Exact native, blank, and OCR quality boundaries passed in the 5/5 focused suite. | Every native and OCR threshold has a passing boundary plus a failing neighbor; visible ink with empty OCR is blocked. | Pure functions and named policy version were present at green; final focused suite remained 5/5. |
| 2.5 | tests/ocr/detection.test.ts |
Unit | N/A (new production file) | Ordered composition, blank omission, canonical hash, audit-source, and risk assertions preceded production code. | Deterministic composition and risk prioritization passed in the 5/5 focused suite. | Reordered pages, tied OCR boxes, blank middle page, repeated token, and four known ambiguous tokens exercise distinct paths. | Reused shared canonical JSON and SHA-256 utilities; final focused suite and type check remained green. |
Test Summary
- Tests: 5 written and focused-passing; 45 canonical (unit 4, integration fixture 1)
- Approval tests: None — new production files; pure functions created: 6
Work Unit Evidence
| Evidence | Result |
|---|---|
| Focused test command and exact result | NODE_ENV=test npx --no-install tsx --test tests/ocr/detection.test.ts — exit 0; 5 tests passed, 0 failed. |
| Runtime harness command/scenario and exact result | NODE_ENV=test npx --no-install tsx --test --test-name-pattern "parsed mixed PDF" tests/ocr/detection.test.ts — exit 0; 1 test passed, 0 failed. The harness generated a three-page PDF, parsed it through the Node 22 pdf-parse page callback, and selected only blank page 2 and insufficient page 3 for OCR. |
| Rollback boundary | Delete src/modules/ocr/detection.ts, src/modules/ocr/composition.ts, and tests/ocr/detection.test.ts; revert tasks 2.2/2.4/2.5 and this Unit 4 progress/history entry. Units 1–3 and their parser threat coverage remain intact. |
Validation and Boundary
- Canonical
npm testpassed 45/45;npm run check, trackedgit diff --check, and no-index whitespace checks for all three new files exited 0. - Authored Unit 4 code and tests total 322 additions and 0 deletions, below the 400-line budget with no size exception.
- Start: Unit 3 page extraction and parser threat routing are complete. End: detection, blank/quality gates, deterministic composition, hashes, and risk tokens are green.
- Out of scope: OCR service, client, dispatcher, ingest routing, review APIs, commits, pushes, PRs, deployment, and restricted paths.
- No specification or design deviation.
Unit 5: OCR Service
Implementation and Correction Summary
- Added pinned Unit 5 FastAPI test dependencies and retained the ignored
ocr-service/.venvenvironment. - Added RED-first HTTP tests and a SQLite-backed private API for bearer auth, idempotency, allowlisting, upload/page limits, queue pressure, status, deletion, and health.
- Preserved failed evidence
sha256:4a5c8ed8d8b302f7ff654d4f40fe0f13049bfdd29db4c76576f3847c0a9be15a: its harness returned401/400/400; inspection showed&backgrounded the preceding shell AND-list, so request variables existed only in the child shell and authenticated curls sent an emptyrequestform, reproduced as400 INVALID_REQUEST. - Corrected only the harness command boundary; the 381-line candidate implementation remained unchanged before progress persistence.
TDD Cycle Evidence
| Task | Test File | Layer | Safety Net | RED | GREEN | TRIANGULATE | REFACTOR |
|---|---|---|---|---|---|---|---|
| 3.1 | ocr-service/tests/test_api.py |
HTTP integration | N/A (new service) | Collection failed with ModuleNotFoundError: No module named 'app'; failed localhost runtime evidence then supplied correction RED. |
8/8 focused tests and corrected curl 401/202/409 passed. |
Auth variants, idempotency/conflict, allowlist, limits, pressure, integrity, and real multipart requests exercised distinct paths. | No production refactor: the defect was shell variable scope in the harness. |
| 3.2 | ocr-service/tests/test_api.py |
HTTP integration/SQLite | N/A (new service) | Same missing-service RED covered the API; failed harness blocked completion. | 8/8 focused tests and real Uvicorn harness passed. | Ready/not-ready health, queue capacity, status, deletion, and process cleanup exercised distinct paths. | Pure page hashing and centralized terminal errors retained; no correction needed. |
Work Unit Evidence
- Focused:
ocr-service/.venv/bin/python -m pytest ocr-service/tests -k auth— exit 0; 8 passed, 0 failed (one dependency deprecation warning). - Runtime: local Uvicorn plus multipart curl — unauthorized
401withUNAUTHORIZED, accepted202with queued identity, conflicting same key409withIDEMPOTENCY_CONFLICT. - Gates:
npm test45/45,npm run check, Pythoncompileall, tracked and untracked whitespace checks — all exit 0/equivalent clean. - Cleanup: Uvicorn PID 268425 terminated and absent; temporary PDF, SQLite DB, and log removed;
.venvretained. - Rollback: remove
.gitignoreUnit 5 entries andocr-service/, then revert tasks 3.1/3.2 and this Unit 5 progress/history block; Units 1–4 remain intact. - Fresh evidence revision:
sha256:0d8b2c57bbab967473eaf99a7ca900168554e64254fa80197b3626b4570b58d0; corrected final authored change count: 397 additions plus deletions, within 400.
Unit 6: OCR Rendering and Image
Implementation Summary
- Added deterministic 200 DPI PDF page rendering with a pre-render 25-megapixel guard, a PaddleOCR adapter, stable line identities, quality metrics, and the exact result schema.
- Added a deterministic fake-driven test seam and preserved ambiguous OCR tokens without correction.
- Added the private CPU image with PaddleOCR
3.4.0, PaddlePaddle CPU3.2.2, baked Latin/mobile models, one Uvicorn worker, a non-root user, private job storage, healthcheck, and documented 3 CPU/5 GiB deployment limits. - Added a Dockerfile-specific deny-by-default build context so restricted and unrelated repository paths are never sent to the Docker daemon.
TDD Cycle Evidence
| Task | Test File | Layer | Safety Net | RED | GREEN | TRIANGULATE | REFACTOR |
|---|---|---|---|---|---|---|---|
| 3.3 | ocr-service/tests/test_render.py |
Unit/integration fixture | Existing OCR API suite passed 8/8 before production edits. | Focused collection failed because app.engine did not exist. |
Focused render suite passed 5/5; complete OCR suite passed 13/13. | Selected pages 1/3, invalid page and pixel-limit paths, repeatable fake results, and Paddle payload normalization exercise distinct behavior. | Named constants and pure metric/result transforms retained; focused suite remained 5/5. |
| 3.4 | ocr-service/tests/test_render.py |
Static/container runtime | Existing OCR API suite passed 8/8 before production edits. | Docker contract was absent; deny-by-default context triangulation then failed until Dockerfile.dockerignore existed. |
Image built with baked models; offline constrained container rendered and validated PNG/result output. | Static pin/runtime assertions plus real model build and --network none execution cover configuration and runtime paths. |
Added missing OpenCV runtime libraries after the first build exposed libGL.so.1; exact corrected image passed. |
Work Unit Evidence
| Evidence | Result |
|---|---|
| Focused test command and exact result | ocr-service/.venv/bin/python -m pytest ocr-service/tests -k render -q — exit 0; 5 passed, 8 deselected, 1 dependency deprecation warning. |
| Runtime harness command/scenario and exact result | docker build --file ocr-service/Dockerfile --tag rag-ocr-service:unit6 . — corrected exact image exited 0 and baked both models. A docker run --rm --network none --cpus 3 --memory 5g ... harness loaded cached PaddleOCR 3.4.0/PaddlePaddle 3.2.2 models, rendered a 1700x2200 PNG (23,309 bytes), returned schema 1, one page, and one OCR line; host Pillow reopened it as PNG 1700x2200. |
| Rollback boundary | Remove ocr-service/{Dockerfile,Dockerfile.dockerignore,app/engine.py,app/models.py,app/render.py,tests/test_render.py} and revert Unit 6 changes in app/main.py, requirements.txt, and README.md; Units 1–5 remain intact. |
Validation and Boundary
- RED was followed by GREEN and post-refactor reruns. One expected GREEN iteration corrected a test-side non-whitespace count from 16 to the mathematically correct 15.
- The first image build failed on missing
libGL.so.1; adding the required slim-image runtime libraries produced a successful build and runtime harness. A later optional connectivity-check refactor rebuild exhausted Docker storage; that unvalidated one-line refactor was reverted, the validated Dockerfile bytes were restored, and the image tag was removed during cleanup. npm testpassed 45/45;npm run check, Pythoncompileall, tracked/untracked whitespace checks, and the full OCR pytest suite (13/13) passed.- Start: Unit 5 private queue/API is complete. End: Unit 6 render/engine/schema/container behavior is complete. Unit 7 client work was not started.
- Native Unit 6 implementation/tests/docs are 397 authored additions plus deletions, within the 400-line budget; required SDD progress and workspace history metadata are administrative evidence outside that native slice.
- No specification or design deviation. The Docker image is private by deployment contract; CPU/RAM labels document limits that the platform must enforce.