rag-service/openspec/changes/ocr-ingest-integration/apply-progress.md

226 lines
20 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Apply Progress: OCR Ingest Integration
## Current State
- **Mode:** Strict TDD
- **Delivery:** Feature-branch-chain, Units 1–6 complete; maintainer-approved `size:exception` for Unit 2 only
- **Completed tasks:** 1.1, 1.2, 1.3, 1.4, 2.1, 2.2, 2.3, 2.4, 2.5, 3.1, 3.2, 3.3, 3.4
- **Overall task progress:** 13/30 complete
## Unit 1: Migration
### Implementation Summary
- Added `migrations/002_ocr_review.sql` with durable OCR jobs, auditable page records, atomic review correction targets, and a partial unique index for non-terminal OCR identities.
- Added focused schema contract coverage in `tests/catalog/migration-002.test.ts`.
- Preserved `migrations/001_knowledge_lifecycle.sql` and all repository logic unchanged.
### TDD Cycle Evidence
| Task | Test File | Layer | Safety Net | RED | GREEN | TRIANGULATE | REFACTOR |
|---|---|---|---|---|---|---|---|
| 1.2 | `tests/catalog/migration-002.test.ts` | Schema contract | N/A (new files) | Valid RED: 4/4 failed with `ENOENT` because migration 002 did not exist. An earlier syntax-invalid run was corrected and is not counted as RED. | 4/4 passed with migration 002 present. | 4 independent schema behaviors cover jobs, pages, corrections, and pending identity uniqueness. | None needed; SQL already matches the minimal authoritative contract. |
### Test Summary
- **Total tests written:** 4
- **Total tests passing:** 4
- **Layers used:** Schema contract (4)
- **Approval tests:** None — no existing production file was modified
- **Pure functions created:** 0
### Work Unit Evidence
| Evidence | Result |
|---|---|
| Focused test command and exact result | `npx --no-install tsx --test tests/catalog/migration-002.test.ts` — exit 0; 4 tests, 4 passed, 0 failed. |
| Runtime harness command/scenario and exact result | N/A — the task forecast defines Unit 1 as schema-only, and the project has no PostgreSQL test harness or database-test dependency. Applying a live migration would require external database access outside this unit. |
| Rollback boundary | Delete `migrations/002_ocr_review.sql` and `tests/catalog/migration-002.test.ts`, revert task 1.2 to pending, and remove this Unit 1 progress/history entry. No existing migration or repository behavior must be reverted. |
### Additional Validation
- `npm run check` — exit 0.
- `git diff --check` — exit 0 after removing one trailing-space warning introduced in the history header.
### Deviations
None — the migration follows the proposal, specifications, design, and closed OCR persistence contract.
### Remaining Persistence Tasks
- [x] 1.1 Complete repository-level RED coverage for duplicate pending ingestion, lease recovery, and review transitions in Unit 2.
- [x] 1.2 Add migration 002 schema and pending OCR identity unique index.
- [x] 1.3 Add catalog repository candidates and leases in Unit 2.
- [x] 1.4 Refactor persistence code after Unit 2 reaches green.
### Review Boundary
- **Start:** Migration 001 already supplies lifecycle states and parent catalog tables.
- **End:** Migration 002 and its focused schema contract test are green.
- **Out of scope:** Repository methods, dispatch behavior, runtime PostgreSQL migration, Unit 2, commits, pushes, PRs, deployment, and secrets.
- **Rollback:** Remove only the two Unit 1 implementation files and their progress metadata.
- **Authored review budget:** 225 additions/deletions, below the 400-line budget for this autonomous slice.
## Unit 2: Repository
### Implementation Summary
- Added pending OCR identity lookup, atomic queue claiming, expired-lease recovery, and guarded review transitions to `CatalogRepository`.
- Added a focused fake-pool repository harness that covers duplicate reuse, empty queues, remote identity preservation, and invalid transitions.
- Corrected the canonical test script so it executes both root-level and nested test suites.
### TDD Cycle Evidence
| Task | Test File | Layer | Safety Net | RED | GREEN | TRIANGULATE | REFACTOR |
|---|---|---|---|---|---|---|---|
| 1.1 | `tests/catalog/repository-ocr.test.ts` | Unit/fake pool | 4/4 existing tests passed | 0/6 with methods absent; then 5/6 while each completion gate was missing | 6/6 passed after task 1.3 | Found/not-found candidates, claimed/empty queue, expired lease, valid/invalid transition | Covered by 1.4 |
| 1.3 | `tests/catalog/repository-ocr.test.ts` | Unit/fake pool | 4/4 existing tests passed | Tests preceded production code | 6/6 passed | Multiple inputs and state paths exercised | Shared OCR row mapping extracted only after green |
| 1.4 | `tests/catalog/repository-ocr.test.ts` | Approval refactor | 6/6 before refactor | N/A — behavior-preserving approval baseline | 6/6 after refactor | Existing six cases preserved | Reusable aliased-column projection removed SQL duplication |
### Work Unit Evidence
| Evidence | Result |
|---|---|
| Focused test command and exact result | `NODE_ENV=test npx --no-install tsx --test tests/catalog/repository-ocr.test.ts` — exit 0; 6 tests passed, 0 failed. `NODE_ENV=test npx --no-install tsx --test tests/catalog/*.test.ts` — exit 0; 10 tests passed, 0 failed. |
| Runtime harness command/scenario and exact result | `NODE_ENV=test npx --no-install tsx --test tests/*.test.ts` — exit 0; 25 root tests passed, 0 failed. Canonical `npm test` — exit 0; 35 total tests passed, 0 failed, proving the 25 root and 10 nested tests execute together. |
| Rollback boundary | Revert Unit 2 changes in `src/modules/catalog/repository.ts`, delete `tests/catalog/repository-ocr.test.ts`, and revert the `package.json` test script; Unit 1 migration remains intact. |
### Validation and Boundary
- Corrective rerun for failed evidence revision `sha256:9433c0cd6628f36f4c96a27d41c6d429f4303d038ffe928f297367c274f21c2f`; fresh evidence revision `sha256:a19119f84381447f54a6223e2a203acaf9d3ae08b6ae8f585e3bd52cd510909b`.
- The failed canonical script expanded only nested tests after nested suites appeared. The minimal correction is `NODE_ENV=test tsx --test tests/*.test.ts tests/**/*.test.ts`.
- Focused repository suite passed 6/6; all catalog suites passed 10/10; explicit root suites passed 25/25; canonical `npm test` passed 35/35; `npm run check` and `git diff --check` exited 0.
- Native Unit 2 accounting was 511 changed lines. The maintainer approved `size:exception` exclusively for Unit 2; the corrective attempt has a separate 600-line remediation budget. All later units retain the normal 400-line budget.
- No design deviations. No Unit 3 work was started.
### Corrective TDD Evidence
| Correction | Safety Net / RED | GREEN | TRIANGULATE | REFACTOR |
|---|---|---|---|---|
| Canonical test discovery | Failed evidence revision proved canonical `npm test` executed only 10 nested tests while the explicit root command executed 25 tests. | Updated only the package test argument list; canonical `npm test` passed 35/35. | Direct root 25/25 and nested catalog 10/10 results sum to and identify the canonical 35/35 run. | None needed — one package-script argument was added. |
## Unit 3: Extraction
### Implementation Summary
- Added a Node 22 fixture spike proving `pdf-parse` page callbacks backed by bundled PDF.js `2.0.550` preserve ordered pages, including a blank page.
- Added `parsePdfPages` with one-based page identity, raw native text, stable SHA-256 hashes, and fail-closed complete-page accounting.
- Preserved textual ingestion as data-only, excluded `CMakeLists.txt`, MDX, and shell lookalikes, and kept task 2.2 pending for Unit 4 detection/composition assertions.
### TDD Cycle Evidence
| Task | Test File | Layer | Safety Net | RED | GREEN | TRIANGULATE | REFACTOR |
|---|---|---|---|---|---|---|---|
| 2.1 | `tests/parsers/pdf-pages.test.ts`; `scripts/spike-pdfjs.ts` | Runtime spike | 20/20 lifecycle service tests passed | Focused suite failed 0/1 because `parsePdfPages` did not exist; the spike and parser implementation were still absent. | Node 22.14.0 spike returned pages 1–3 exactly, including blank page 2. | Three pages prove non-empty, blank, and later-page ordering paths. | Replaced the over-budget direct dependency with the design-approved existing `pdf-parse` pagerender path; spike remained green. |
| 2.3 | `tests/parsers/pdf-pages.test.ts` | Unit/integration fixture | 20/20 lifecycle service tests passed | Focused suite failed 0/1 on the missing export before production code. | Final focused run passed 5/5 after implementation. | Covered page hashes/blank page, native PDF composition, non-PDF rejection, data-only text parsing, and unsupported lookalikes. | Focused suite remained 5/5 after consolidating on the existing parser dependency. |
### Work Unit Evidence
| Evidence | Result |
|---|---|
| Focused test command and exact result | `NODE_ENV=test npx --no-install tsx --test tests/parsers/pdf-pages.test.ts` — exit 0; 5 passed, 0 failed. |
| Runtime harness command/scenario and exact result | `NODE_ENV=test npx --no-install tsx scripts/spike-pdfjs.ts` — exit 0 on Node 22.14.0; PDF.js 2.0.550 returned three ordered pages with page 2 blank. |
| Rollback boundary | Revert `src/modules/parsers/parser-registry.ts`; delete `scripts/spike-pdfjs.ts`, `tests/parsers/pdf-pages.test.ts`, and `tests/fixtures/ocr/native-three-pages.pdf`; revert tasks 2.1/2.3 and this Unit 3 progress/history entry. Units 1–2 remain intact. |
### Validation and Boundary
- Canonical `npm test` passed 40/40; `npm run check` and `git diff --check` exited 0.
- Native authored slice before progress/history metadata: 225 additions/deletions; no package or lockfile delta remains.
- Start: Unit 2 persistence is complete. End: per-page native extraction and parser threat routing are green.
- Out of scope: detection thresholds, OCR selection, composition, risks, Unit 4 tasks 2.4/2.5, commits, pushes, PRs, deployment, and restricted paths.
- No specification deviation. The design-authorized `pdf-parse` pagerender fallback was selected only after its Node 22 fixture spike passed.
## Unit 4: Detection and Composition
### Implementation Summary
- Added pure per-page native detection under `pdf-detection-v1`, exact OCR page selection, verified-blank classification, and fail-closed OCR quality gates.
- Added deterministic candidate composition using exactly one source per page, ordered OCR lines, blank omission, canonical hashes, and unchanged risk tokens prioritized for review.
- Completed the remaining task 2.2 scenarios by combining the Unit 3 data-only/parser threat coverage with Unit 4 mixed-PDF, detection, blank, hash, and risk RED assertions.
### TDD Cycle Evidence
| Task | Test File | Layer | Safety Net | RED | GREEN | TRIANGULATE | REFACTOR |
|---|---|---|---|---|---|---|---|
| 2.2 | `tests/ocr/detection.test.ts`; `tests/parsers/pdf-pages.test.ts` | Unit/integration fixture | Unit 3 parser suite passed 5/5 before edits. | Focused Unit 4 suite failed before production with `ERR_MODULE_NOT_FOUND` for `src/modules/ocr/detection.js`. | Focused suite passed 5/5 after detection and composition were added. | Real generated mixed PDF selected pages 2 and 3; `.txt`/`.md` stayed data-only and CMake/MDX/shell lookalikes stayed unsupported. | None needed; routing remains separated between parser support and the pure PDF OCR planner. |
| 2.4 | `tests/ocr/detection.test.ts` | Unit | N/A (new production file) | Threshold, mixed-page, blank, and quality assertions preceded production code. | Exact native, blank, and OCR quality boundaries passed in the 5/5 focused suite. | Every native and OCR threshold has a passing boundary plus a failing neighbor; visible ink with empty OCR is blocked. | Pure functions and named policy version were present at green; final focused suite remained 5/5. |
| 2.5 | `tests/ocr/detection.test.ts` | Unit | N/A (new production file) | Ordered composition, blank omission, canonical hash, audit-source, and risk assertions preceded production code. | Deterministic composition and risk prioritization passed in the 5/5 focused suite. | Reordered pages, tied OCR boxes, blank middle page, repeated token, and four known ambiguous tokens exercise distinct paths. | Reused shared canonical JSON and SHA-256 utilities; final focused suite and type check remained green. |
### Test Summary
- **Tests:** 5 written and focused-passing; 45 canonical (unit 4, integration fixture 1)
- **Approval tests:** None — new production files; **pure functions created:** 6
### Work Unit Evidence
| Evidence | Result |
|---|---|
| Focused test command and exact result | `NODE_ENV=test npx --no-install tsx --test tests/ocr/detection.test.ts` — exit 0; 5 tests passed, 0 failed. |
| Runtime harness command/scenario and exact result | `NODE_ENV=test npx --no-install tsx --test --test-name-pattern "parsed mixed PDF" tests/ocr/detection.test.ts` — exit 0; 1 test passed, 0 failed. The harness generated a three-page PDF, parsed it through the Node 22 `pdf-parse` page callback, and selected only blank page 2 and insufficient page 3 for OCR. |
| Rollback boundary | Delete `src/modules/ocr/detection.ts`, `src/modules/ocr/composition.ts`, and `tests/ocr/detection.test.ts`; revert tasks 2.2/2.4/2.5 and this Unit 4 progress/history entry. Units 1–3 and their parser threat coverage remain intact. |
### Validation and Boundary
- Canonical `npm test` passed 45/45; `npm run check`, tracked `git diff --check`, and no-index whitespace checks for all three new files exited 0.
- Authored Unit 4 code and tests total 322 additions and 0 deletions, below the 400-line budget with no size exception.
- Start: Unit 3 page extraction and parser threat routing are complete. End: detection, blank/quality gates, deterministic composition, hashes, and risk tokens are green.
- Out of scope: OCR service, client, dispatcher, ingest routing, review APIs, commits, pushes, PRs, deployment, and restricted paths.
- No specification or design deviation.
## Unit 5: OCR Service
### Implementation and Correction Summary
- Added pinned Unit 5 FastAPI test dependencies and retained the ignored `ocr-service/.venv` environment.
- Added RED-first HTTP tests and a SQLite-backed private API for bearer auth, idempotency, allowlisting, upload/page limits, queue pressure, status, deletion, and health.
- Preserved failed evidence `sha256:4a5c8ed8d8b302f7ff654d4f40fe0f13049bfdd29db4c76576f3847c0a9be15a`: its harness returned `401/400/400`; inspection showed `&` backgrounded the preceding shell AND-list, so request variables existed only in the child shell and authenticated curls sent an empty `request` form, reproduced as `400 INVALID_REQUEST`.
- Corrected only the harness command boundary; the 381-line candidate implementation remained unchanged before progress persistence.
### TDD Cycle Evidence
| Task | Test File | Layer | Safety Net | RED | GREEN | TRIANGULATE | REFACTOR |
|---|---|---|---|---|---|---|---|
| 3.1 | `ocr-service/tests/test_api.py` | HTTP integration | N/A (new service) | Collection failed with `ModuleNotFoundError: No module named 'app'`; failed localhost runtime evidence then supplied correction RED. | 8/8 focused tests and corrected curl `401/202/409` passed. | Auth variants, idempotency/conflict, allowlist, limits, pressure, integrity, and real multipart requests exercised distinct paths. | No production refactor: the defect was shell variable scope in the harness. |
| 3.2 | `ocr-service/tests/test_api.py` | HTTP integration/SQLite | N/A (new service) | Same missing-service RED covered the API; failed harness blocked completion. | 8/8 focused tests and real Uvicorn harness passed. | Ready/not-ready health, queue capacity, status, deletion, and process cleanup exercised distinct paths. | Pure page hashing and centralized terminal errors retained; no correction needed. |
### Work Unit Evidence
- Focused: `ocr-service/.venv/bin/python -m pytest ocr-service/tests -k auth` — exit 0; 8 passed, 0 failed (one dependency deprecation warning).
- Runtime: local Uvicorn plus multipart curl — unauthorized `401` with `UNAUTHORIZED`, accepted `202` with queued identity, conflicting same key `409` with `IDEMPOTENCY_CONFLICT`.
- Gates: `npm test` 45/45, `npm run check`, Python `compileall`, tracked and untracked whitespace checks — all exit 0/equivalent clean.
- Cleanup: Uvicorn PID 268425 terminated and absent; temporary PDF, SQLite DB, and log removed; `.venv` retained.
- Rollback: remove `.gitignore` Unit 5 entries and `ocr-service/`, then revert tasks 3.1/3.2 and this Unit 5 progress/history block; Units 1–4 remain intact.
- Fresh evidence revision: `sha256:0d8b2c57bbab967473eaf99a7ca900168554e64254fa80197b3626b4570b58d0`; corrected final authored change count: 397 additions plus deletions, within 400.
## Unit 6: OCR Rendering and Image
### Implementation Summary
- Added deterministic 200 DPI PDF page rendering with a pre-render 25-megapixel guard, a PaddleOCR adapter, stable line identities, quality metrics, and the exact result schema.
- Added a deterministic fake-driven test seam and preserved ambiguous OCR tokens without correction.
- Added the private CPU image with PaddleOCR `3.4.0`, PaddlePaddle CPU `3.2.2`, baked Latin/mobile models, one Uvicorn worker, a non-root user, private job storage, healthcheck, and documented 3 CPU/5 GiB deployment limits.
- Added a Dockerfile-specific deny-by-default build context so restricted and unrelated repository paths are never sent to the Docker daemon.
### TDD Cycle Evidence
| Task | Test File | Layer | Safety Net | RED | GREEN | TRIANGULATE | REFACTOR |
|---|---|---|---|---|---|---|---|
| 3.3 | `ocr-service/tests/test_render.py` | Unit/integration fixture | Existing OCR API suite passed 8/8 before production edits. | Focused collection failed because `app.engine` did not exist. | Focused render suite passed 5/5; complete OCR suite passed 13/13. | Selected pages 1/3, invalid page and pixel-limit paths, repeatable fake results, and Paddle payload normalization exercise distinct behavior. | Named constants and pure metric/result transforms retained; focused suite remained 5/5. |
| 3.4 | `ocr-service/tests/test_render.py` | Static/container runtime | Existing OCR API suite passed 8/8 before production edits. | Docker contract was absent; deny-by-default context triangulation then failed until `Dockerfile.dockerignore` existed. | Image built with baked models; offline constrained container rendered and validated PNG/result output. | Static pin/runtime assertions plus real model build and `--network none` execution cover configuration and runtime paths. | Added missing OpenCV runtime libraries after the first build exposed `libGL.so.1`; exact corrected image passed. |
### Work Unit Evidence
| Evidence | Result |
|---|---|
| Focused test command and exact result | `ocr-service/.venv/bin/python -m pytest ocr-service/tests -k render -q` — exit 0; 5 passed, 8 deselected, 1 dependency deprecation warning. |
| Runtime harness command/scenario and exact result | `docker build --file ocr-service/Dockerfile --tag rag-ocr-service:unit6 .` — corrected exact image exited 0 and baked both models. A `docker run --rm --network none --cpus 3 --memory 5g ...` harness loaded cached PaddleOCR 3.4.0/PaddlePaddle 3.2.2 models, rendered a 1700x2200 PNG (23,309 bytes), returned schema `1`, one page, and one OCR line; host Pillow reopened it as PNG 1700x2200. |
| Rollback boundary | Remove `ocr-service/{Dockerfile,Dockerfile.dockerignore,app/engine.py,app/models.py,app/render.py,tests/test_render.py}` and revert Unit 6 changes in `app/main.py`, `requirements.txt`, and `README.md`; Units 1–5 remain intact. |
### Validation and Boundary
- RED was followed by GREEN and post-refactor reruns. One expected GREEN iteration corrected a test-side non-whitespace count from 16 to the mathematically correct 15.
- The first image build failed on missing `libGL.so.1`; adding the required slim-image runtime libraries produced a successful build and runtime harness. A later optional connectivity-check refactor rebuild exhausted Docker storage; that unvalidated one-line refactor was reverted, the validated Dockerfile bytes were restored, and the image tag was removed during cleanup.
- `npm test` passed 45/45; `npm run check`, Python `compileall`, tracked/untracked whitespace checks, and the full OCR pytest suite (13/13) passed.
- Start: Unit 5 private queue/API is complete. End: Unit 6 render/engine/schema/container behavior is complete. Unit 7 client work was not started.
- Native Unit 6 implementation/tests/docs are 397 authored additions plus deletions, within the 400-line budget; required SDD progress and workspace history metadata are administrative evidence outside that native slice.
- No specification or design deviation. The Docker image is private by deployment contract; CPU/RAM labels document limits that the platform must enforce.