rag-service/openspec/changes/archive/2026-09-21-ocr-ingest-integration/exploration.md

106 lines
No EOL
20 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Exploration: OCR ingest integration (`ocr-ingest-integration`)
**Project:** rag-service — **Phase:** sdd-explore — **Date:** 2026-09-13
**Store:** hybrid (OpenSpec file + Engram topic `sdd/ocr-ingest-integration/explore`)
## Current State
Point 2 (knowledge lifecycle) is complete, enforced in production (`KNOWLEDGE_LIFECYCLE_ENFORCED=true`, 7 cataloged sources, 22,605 active points, legacy batch migrated and verified 2026-09-13). `npm test` passes 25/25. The contract in `docs/CONTRATO_CICLO_VIDA_Y_OCR.md` (section "Punto 3") is closed and authoritative: every decision below defers to it; this exploration only resolves what the contract leaves open.
Verified implementation facts (read from code, not docs):
| Area | Current behavior | OCR gap |
|---|---|---|
| Ingest | `IngestService.ingestWithLifecycle` is fully synchronous: attempt → source lock → discover → hash → reusable lookup → `createPendingVersion` → `markIndexing` → embed → upsert → count → `markReady` → `activateVersion`, all inside the HTTP request | Must become async-capable: durable original first, `202 + statusUrl` response when OCR is needed, `source_content_hash = null` until review completes |
| PDF parsing | `parser-registry.parseDocument` uses `pdf-parse` over the whole file, returns aggregated text only; no page callback in our code (library has `pagerender` option but we never use it; bundled pdf.js is v1.10.100, very old) | Page-level extraction required; contract allows replacing `pdf-parse` if unreliable |
| State machine | `SourceVersionState` already includes `review_required` and `rejected`; transitions `pending→indexing→ready→active`, `failed`, `purging→purged` implemented in `CatalogRepository`; **no transition writes into `review_required` today** | Need: `indexing→review_required` (OCR candidate complete), `review_required→indexing` (approve), `review_required→rejected`, plus `review_required`/`rejected` entries in `markPurging`-eligible states (already present), retention expiry transitions |
| Catalog repo | `createPendingVersion` requires non-null `expectedPointCount` and writes all documents with `content_hash` pre-computed; `findReusableVersion` skips when `sourceContentHash` is null (correct for OCR candidates); `hasCompleteVersionDocuments` counts `content_hash IS NULL` as incomplete | OCR flow must create a pending version **before** content exists (chunk counts unknown until review) — needs a creation mode with `expectedPointCount=0`/null-safe path, or a separate "OCR candidate" insert path |
| Reconciler | `runOnce` treats stale `pending/indexing` older than 30 min as orphans and either recovers to `ready` or fails them | Must not kill in-flight OCR versions: needs heartbeat awareness (jobs have `heartbeat_at`/`lease_expires_at` per contract) and must recover expired OCR leases by re-dispatching with the same remote idempotency key |
| Uploads | `multer.memoryStorage()` in `src/app.ts`, no size limit; upload writes to `/tmp` with timestamp, deleted in `finally` — no durable original | Contract requires disk-based multer with limits and durable original under `/data/ingestions/<versionId>/` before processing |
| Env/config | `env.ts` has PostgreSQL + lifecycle tokens only | Needs `OCR_SERVICE_URL`, `OCR_INTERNAL_TOKEN`, artifact volume path, limits (50 MiB, 100 pages, timeouts), retention intervals |
| Migrations | `migrations/001_knowledge_lifecycle.sql` applied via `rag_schema_migrations` with checksums + advisory lock; migration runner sorts `^\d+_.*\.sql$` | `002_ocr_review.sql` fits the existing runner unchanged (tables `rag_ocr_jobs`, `rag_document_pages`, `rag_review_corrections` per contract) |
| OpenAPI | `IngestResponse` is 201-only under lifecycle; `securitySchemes.bearerAuth` already defined and used by lifecycle endpoints | Needs `202` variant with `phase`/`statusUrl`, `/ingestions/:versionId` status route, review/approve/reject routes |
| Playground | `public/playground/app.js` supports activate/expectedActiveVersionId; no review UI | Review view (page images, native vs OCR text, corrections) required by acceptance criteria |
| Docker | Node runtime only; `CMD ["node", "dist/server.js"]` | Needs `/data/ingestions` volume; ocr-service gets its own Dockerfile (Python CPU, PaddleOCR 3.4.0 + PaddlePaddle 3.2.2 pinned, baked models) |
| Artifacts | No filesystem artifact manager exists; `rag_version_documents.artifact_manifest_path` and `artifact_state`/`retention_due_at` columns exist but are never written (`retention_deleted` only set on purge) | Whole artifact subsystem (write/verify 0600, manifest, retention sweeper) is new |
## Affected Areas
- `src/modules/ingest/service.ts` — async OCR branch, durable originals, version creation before content
- `src/modules/parsers/parser-registry.ts` — per-page PDF extraction API
- `src/modules/catalog/repository.ts` — new state transitions, OCR-safe version creation, page/job/correction persistence
- `src/modules/catalog/reconciler.ts` — lease-aware orphan handling, OCR job re-dispatch
- `src/app.ts` — disk multer, 202 async responses, `/ingestions/*` + review endpoints
- `src/api/openapi.ts`, `public/playground/*` — new contracts and review UI
- `src/config/env.ts`, `migrations/002_ocr_review.sql`, `Dockerfile`, `docs/API_RAG.md`, `docs/INGESTA.md`
- New: `src/modules/ocr/{client,detection,artifacts,review}.ts`, `ocr-service/` (Python)
## Open Decisions Resolved
| # | Decision | Resolution | Rationale |
|---|---|---|---|
| 1 | Page-level PDF extraction | **Replace `pdf-parse` with `pdfjs-dist` (legacy build) used directly**, wrapped in a new `parsePdfPages(buffer)` that returns `Array<{pageNumber, text}>`. Keep `pdf-parse` only if a spike proves its `pagerender` callback reliable on Node 22 — the contract permits either, but pdf.js v1.10.100 (2018) is too old to trust for the metrics contract. Fallback order: pdfjs-dist direct → `pdf-parse` with custom pagerender. Must be settled by a fixture-backed spike in Block 2 before anything downstream builds on it | Contract mandates "callback per pagina probado con fixture" and allows substitution. pdfjs-dist is the same engine, maintained, permissively licensed (Apache-2.0), and gives page objects with text items we need for metrics |
| 2 | Job/page state machine & progress | **Reuse the contract schema exactly**: `rag_ocr_jobs` (one per version+document, `UNIQUE(version_id, document_id)`), `rag_document_pages` per page, dispatcher uses `FOR UPDATE SKIP LOCKED`, heartbeats. Progress = `completed_pages/totalPages` surfaced via `GET /ingestions/:versionId` from a single aggregate query over jobs+pages. No new states beyond contract | Schema is already specified in the contract (Punto 3 "Persistencia OCR"); inventing alternatives violates the closed decision |
| 3 | Multi-document semantics | **Version-level candidate gate**: a version reaches `review_required` only when ALL documents' pages are native-sufficient or OCR-complete (contract quality gate is per-document AND per-version). One failed document fails the version (`OCR_QUALITY_BLOCKED` or the document's `error_code`). Mixed sources (md + pdf) in one folder ingest: only PDFs go through OCR; non-PDF docs keep native content_hash immediately, PDF docs get `content_hash` after review | The contract's gate ("el documento solo llega a review_required cuando…") is per-document; activation is version-level, so the join is forced |
| 4 | Reconciler behavior | Extend `recoverOrphanedVersions` to distinguish: (a) versions in `indexing` with live OCR jobs (`lease_expires_at > now`) → leave alone; (b) expired leases → re-enqueue with same `remote_idempotency_key`, reset `next_attempt_at`; (c) stale native versions → existing behavior. The dispatcher itself (new module) owns lease renewal; reconciler only recovers dead leases | Contract: "Un reconciliador periodico reenvia con la misma clave remota y recupera remote_job_id sin duplicar OCR" |
| 5 | Durable originals/artifacts | **Write original + manifest under `/data/ingestions/<versionId>/documents/<documentArtifactId>/` BEFORE `createPendingVersion`** (contract sequence: "se guarda el original durable; se crea la version…"). `documentArtifactId` = UUIDv5(document_id). If ingest dies between artifact write and version creation, a startup/periodic sweep deletes orphan version dirs (no catalog row = not referenceable). `artifact_state` column transitions: `none → present` at write, `present → retention_deleting → retention_deleted` by sweeper | Ordering is mandated by the contract's async section; sweep covers the one crash window |
| 6 | Purge/retention | **Single daily sweeper** per contract: acquires the same version CAS/lock used by review and purge (`withVersionExclusiveLock`), sets `artifact_state='retention_deleting'`, deletes idempotently, marks `retention_deleted`. State TTLs: review_required 30d → `rejected`/`REVIEW_EXPIRED`; failed/rejected 7d → `purging→purged`; review images 7d post-decision; active keeps original+reviewed while active +30d. An active version is never TTL-deleted | All values are fixed by the contract's "Artefactos y retencion" section |
| 7 | Worker topology | **In-process dispatcher inside the RAG Node service** (single instance, single EasyPanel app): a `setInterval`-driven poller that claims jobs with `FOR UPDATE SKIP LOCKED`, heartbeats, polls the OCR service, never blocks the ingest HTTP path beyond the initial enqueue. No separate Node worker deploy. The OCR service itself is one worker + one replica per contract | Adding a second Node worker would double deploy surface for zero benefit at current scale (≤3 queued jobs); contract fixes OCR service at 1 worker/1 replica, and the RAG API is the only writer |
| 8 | OCR service/image boundary | **Exactly as contracted**: `ocr-service/` in-repo, FastAPI or plain HTTP Python app (implementation detail for design phase), PaddleOCR 3.4.0 + PaddlePaddle 3.2.2 CPU pinned, models baked at build, `/v1/jobs` multipart + Idempotency-Key, `/health/live` + `/health/ready`, internal network only, `OCR_INTERNAL_TOKEN` bearer, SQLite queue on ephemeral `/data/jobs` with 24h TTL, 50 MiB limit, config allowlist. RAG's `src/modules/ocr/client.ts` validates `documentSha256`, page sets, and result schema strictly | Closed decision in contract ("Despliegue OCR", "Contrato interno del servicio OCR") |
| 9 | Durable uploads | Replace `multer.memoryStorage()` with disk storage into a temp staging dir, enforce 50 MiB + 100 pages + PDF sanity (reject encrypted/corrupt before full render), and move the original into the version artifact dir once identity is resolved | Contract: "Antes de habilitar OCR, sustituir multer.memoryStorage() por upload a disco con limite y limpieza garantizada" |
| 10 | Representative fixtures | **Three committed fixtures** under `tests/fixtures/ocr/`: (1) `native.pdf` — text-rich, all pages pass detection (no OCR called); (2) `scanned.pdf` — image-only pages (2–3 pages, small) requiring OCR on all non-blank pages; (3) `mixed.pdf` — some native pages + some scanned pages, containing tokens shaped like the benchmark's risk identifiers (`CBG04a`-like). Plus unit fixtures: single-page blank image, oversized-rejection corpus. The real FacturaTech PDF stays out of the repo (client data); its 34-entry acceptance run happens in production per the contract's gate. Fixtures must be tiny (<200 KiB total) and generated with a script committed for reproducibility, so tests stay offline | `npm test` has no network; OCR service tests run inside its own Python test suite (pytest) — Node tests fake the OCR client. Contract acceptance criteria name the exact behaviors the fixtures must exercise |
## Approaches
1. **Contract-as-specified, in-process dispatcher (recommended)** — implement Punto 3 exactly as the contract defines: async ingest branch with 202, new tables 002, OCR client with retries/polling, artifacts manager, review endpoints, in-process job dispatcher.
- Pros: zero contract amendments; all closed decisions honored; single deployable RAG change + one new OCR service; reconciler extension is small; rollback = feature-flag off.
- Cons: dispatcher lifetime couples to RAG process (acceptable: single-instance deploy already assumed); ingest service grows substantially.
- Effort: High (largest change in RAG history), but decomposable into 7 blocks.
2. **Separate Node worker process** — dispatcher runs as its own EasyPanel service consuming the same PostgreSQL queue.
- Pros: isolates OCR polling latency; RAG API stays minimal.
- Cons: violates the current single-service ops model; duplicate pool/lock code; two deploys to coordinate for what is currently ≤3 concurrent jobs; no requirement justifies it.
- Effort: High+.
3. **Minimal sync OCR first, review later** — call OCR inline during ingest with a long timeout, add review workflow in a later change.
- Pros: fewer moving parts initially.
- Cons: directly violates the closed decision ("responde sin esperar los minutos de OCR", 202 async, review before activation); 245s benchmark runtime would hold HTTP connections; rejected.
**Recommendation: Approach 1.** It is the only one consistent with the closed contract, and the in-process dispatcher keeps deployment identical in shape to today's (RAG + private OCR service).
## Seven Work Blocks
Dependencies flow strictly downward; each block maps to a chained PR slice.
| # | Block | Contents | Depends on | Key risk |
|---|---|---|---|---|
| 1 | Persistence & state machine | `migrations/002_ocr_review.sql`; repo methods: create OCRCandidateVersion (null-safe point counts), `markReviewRequired`, `approve→indexing`, `reject`, lease claim/renew methods; state-machine unit tests over fake pool | — | Migration 002 is additive (new tables + `extraction_method` values already exist); zero impact on live 001 data. Rollback: drop 002 tables (safe — no production data references them until OCR ships) |
| 2 | Page-level PDF extraction | pdfjs-dist spike + `parsePdfPages`; detection metrics (N/A/W/R thresholds); candidate composition (`--- Page N ---` separators); canonical JSON; detection unit tests with fixtures | 1 (types only) | Library choice must be settled by spike evidence early; if pdfjs-dist legacy build fails on Node 22, fall back to `pdf-parse` custom pagerender |
| 3 | PaddleOCR service | `ocr-service/` FastAPI app: jobs (SQLite queue, idempotency, TTL), engine wrapper (PaddleOCR 3.4.0 pinned), render pipeline (200 DPI, 25 MP cap), auth, limits, health; pytest suite with a stub engine | none (parallel-safe; contract fully specifies it) | Image build size/time (Paddle ~2 GB); models must bake at build; readiness false until models load |
| 4 | Orchestration & artifacts | `ocr/client.ts` (retries, backoff 2–15s, strict validation); `ocr/artifacts.ts` (0600 writes, manifest, artifactId); in-process dispatcher (SKIP LOCKED, heartbeats); async ingest branch with 202 + `GET /ingestions/:versionId`; durable disk multer upload; reconciler lease-awareness | 1, 2, 3 | Biggest integration surface: touches ingest, app, reconciler; must not regress native sync path (regression tests required) |
| 5 | Review, corrections, indexing | Risk-token detection (regex + priority rules); `GET /ingestions/:versionId/review`; approve with `candidateSha256` + per-line `expectedLineSha256` corrections (all-or-nothing 409); duplicate-reusable check; post-approval chunk/embed/index/activate via existing point-2 path; playground review UI | 4 | Correction collision semantics (409 without partial application) must be exact; concurrency with activation changes |
| 6 | Security, retention, deployment | Retention sweeper (TTLs, artifact_state CAS); `/health` OCR section; OpenAPI full update (`202`, `/ingestions/*`, review routes, bearer); Dockerfile volume + OCR compose config; docs (API_RAG, INGESTA, deploy); no-secret review | 4, 5 | Deployment ordering: OCR service first, RAG with OCR disabled, then enable — irreversible risk is low if ordering respected |
| 7 | End-to-end validation | Fixture-driven e2e (native/scanned/mixed), idempotent-resend test, fail-closed test (OCR down), full static suite (`npm run check`, `npm run build`, Node tests + pytest), production validation plan for the real FacturaTech PDF | 1–6 | Production acceptance (34 entries) requires human review step; cannot be automated |
Boundary adjustments made (with evidence): the contract's own task order lists "añadir migracion OCR y contratos TypeScript" first and extraction second, but **Block 3 (OCR service) has no code dependency on Blocks 1–2** — the contract fully specifies its HTTP interface. Starting it in parallel is the only way to fit the overall work into review-budget-sized slices without artificial sequencing. This is an execution-order adjustment only; all contract sequencing gates (OCR deployed before RAG enables it) remain.
## Dependencies & Risk Map
- **Hard dependency chain:** 1 → 2 → 4 → 5 → 6 → 7; 3 is parallel, joins at 4.
- **Irreversible operations:** only `migrations/002_ocr_review.sql` (additive DDL — new tables, no ALTER on existing ones; safe rollback = `DROP TABLE rag_ocr_jobs, rag_document_pages, rag_review_corrections` + `DELETE FROM rag_schema_migrations WHERE name='002_ocr_review.sql'`). Qdrant unchanged until approval-driven indexing, which uses existing versioned point IDs. Production artifact volume starts empty; retention never deletes active versions.
- **Rollback boundary:** feature-flag `OCR_INGEST_ENABLED` (new env). Off = today's native-only sync behavior byte-for-byte (regression suite proves it). The 202 async path and `/ingestions/*` routes exist but return `503 OCR_DISABLED`-style errors when off — or routes can be conditionally mounted; design phase decides. Data written while enabled (candidates, jobs) is invisible to retrieval by construction (never `active` without review).
- **Deploy ordering (contract-mandated):** OCR service private deploy + readiness check → RAG deploy with OCR disabled → verify native regressions → enable flag → run real PDF validation.
- **Review-budget pressure (400 lines, auto-chain):** **High across the whole change.** Realistic estimate: Block 1 ~450–550 lines (SQL+repo+tests), Block 2 ~350–450, Block 3 ~600–800 (Python, reviewed separately from Node PRs), Block 4 ~700–900, Block 5 ~600–750, Block 6 ~350–450, Block 7 ~250–350. Every block except possibly Block 7 individually risks exceeding 400 changed lines; chained/stacked PRs are mandatory, and Blocks 4 and 5 may each need two slices (e.g., 4a client+artifacts, 4b dispatcher+ingest async).
## Risks
- **pdfjs-dist integration uncertainty on Node 22** (legacy vs modern build, worker config) — mitigate with a day-one spike in Block 2, contract already allows the `pdf-parse` fallback.
- **OCR image build fragility**: PaddlePaddle 3.2.2 CPU wheels are large and platform-sensitive; baked models add build time. Mitigate: pin exact versions, build once, keep ocr-service deploy independent (contract requirement).
- **Ingest regression on the native path**: the async branch rewrites the hottest service. Mitigate: strict TDD per config (strict_tdd: true), the 25 existing tests must stay green, plus new native-regression tests before any OCR wiring.
- **State-machine drift from the contract**: the contract's diagram is normative; any deviation needs contract amendment FIRST. Mitigation: spec phase encodes each transition as Given/When/Then.
- **Playground review UI scope creep**: image rendering, side-by-side diffs, corrections editor could balloon. Mitigation: MVP review UI = native/OCR text + corrections JSON input + image links; defer polish (the playground doc says it's an internal tool).
- **VPS2 resource contention**: OCR service capped at 3 CPU/5 GiB of the 6 vCPU/11 GiB host; a 25-page/245s benchmark job can saturate one core for minutes. Mitigate: contract's queue depth 3, single worker, and 60s/page timeout are already sized for this.
## Ready for Proposal
**Yes.** Every open decision above is resolved within the closed contract; no human input is required to draft the proposal. The proposal must state: contract-conformant implementation (Approach 1), seven chained blocks as scoped, Block-3 parallelism, feature-flag rollback, and the deploy ordering. Contract amendments are NOT needed — the contract's Punto 3 already anticipates all resolved details; the only judgment calls (worker topology, fixture strategy, library spike order) are implementation-level and recorded here.