# Exploration: OCR ingest integration (`ocr-ingest-integration`) **Project:** rag-service — **Phase:** sdd-explore — **Date:** 2026-09-13 **Store:** hybrid (OpenSpec file + Engram topic `sdd/ocr-ingest-integration/explore`) ## Current State Point 2 (knowledge lifecycle) is complete, enforced in production (`KNOWLEDGE_LIFECYCLE_ENFORCED=true`, 7 cataloged sources, 22,605 active points, legacy batch migrated and verified 2026-09-13). `npm test` passes 25/25. The contract in `docs/CONTRATO_CICLO_VIDA_Y_OCR.md` (section "Punto 3") is closed and authoritative: every decision below defers to it; this exploration only resolves what the contract leaves open. Verified implementation facts (read from code, not docs): | Area | Current behavior | OCR gap | |---|---|---| | Ingest | `IngestService.ingestWithLifecycle` is fully synchronous: attempt → source lock → discover → hash → reusable lookup → `createPendingVersion` → `markIndexing` → embed → upsert → count → `markReady` → `activateVersion`, all inside the HTTP request | Must become async-capable: durable original first, `202 + statusUrl` response when OCR is needed, `source_content_hash = null` until review completes | | PDF parsing | `parser-registry.parseDocument` uses `pdf-parse` over the whole file, returns aggregated text only; no page callback in our code (library has `pagerender` option but we never use it; bundled pdf.js is v1.10.100, very old) | Page-level extraction required; contract allows replacing `pdf-parse` if unreliable | | State machine | `SourceVersionState` already includes `review_required` and `rejected`; transitions `pending→indexing→ready→active`, `failed`, `purging→purged` implemented in `CatalogRepository`; **no transition writes into `review_required` today** | Need: `indexing→review_required` (OCR candidate complete), `review_required→indexing` (approve), `review_required→rejected`, plus `review_required`/`rejected` entries in `markPurging`-eligible states (already present), retention expiry transitions | | Catalog repo | `createPendingVersion` requires non-null `expectedPointCount` and writes all documents with `content_hash` pre-computed; `findReusableVersion` skips when `sourceContentHash` is null (correct for OCR candidates); `hasCompleteVersionDocuments` counts `content_hash IS NULL` as incomplete | OCR flow must create a pending version **before** content exists (chunk counts unknown until review) — needs a creation mode with `expectedPointCount=0`/null-safe path, or a separate "OCR candidate" insert path | | Reconciler | `runOnce` treats stale `pending/indexing` older than 30 min as orphans and either recovers to `ready` or fails them | Must not kill in-flight OCR versions: needs heartbeat awareness (jobs have `heartbeat_at`/`lease_expires_at` per contract) and must recover expired OCR leases by re-dispatching with the same remote idempotency key | | Uploads | `multer.memoryStorage()` in `src/app.ts`, no size limit; upload writes to `/tmp` with timestamp, deleted in `finally` — no durable original | Contract requires disk-based multer with limits and durable original under `/data/ingestions//` before processing | | Env/config | `env.ts` has PostgreSQL + lifecycle tokens only | Needs `OCR_SERVICE_URL`, `OCR_INTERNAL_TOKEN`, artifact volume path, limits (50 MiB, 100 pages, timeouts), retention intervals | | Migrations | `migrations/001_knowledge_lifecycle.sql` applied via `rag_schema_migrations` with checksums + advisory lock; migration runner sorts `^\d+_.*\.sql$` | `002_ocr_review.sql` fits the existing runner unchanged (tables `rag_ocr_jobs`, `rag_document_pages`, `rag_review_corrections` per contract) | | OpenAPI | `IngestResponse` is 201-only under lifecycle; `securitySchemes.bearerAuth` already defined and used by lifecycle endpoints | Needs `202` variant with `phase`/`statusUrl`, `/ingestions/:versionId` status route, review/approve/reject routes | | Playground | `public/playground/app.js` supports activate/expectedActiveVersionId; no review UI | Review view (page images, native vs OCR text, corrections) required by acceptance criteria | | Docker | Node runtime only; `CMD ["node", "dist/server.js"]` | Needs `/data/ingestions` volume; ocr-service gets its own Dockerfile (Python CPU, PaddleOCR 3.4.0 + PaddlePaddle 3.2.2 pinned, baked models) | | Artifacts | No filesystem artifact manager exists; `rag_version_documents.artifact_manifest_path` and `artifact_state`/`retention_due_at` columns exist but are never written (`retention_deleted` only set on purge) | Whole artifact subsystem (write/verify 0600, manifest, retention sweeper) is new | ## Affected Areas - `src/modules/ingest/service.ts` — async OCR branch, durable originals, version creation before content - `src/modules/parsers/parser-registry.ts` — per-page PDF extraction API - `src/modules/catalog/repository.ts` — new state transitions, OCR-safe version creation, page/job/correction persistence - `src/modules/catalog/reconciler.ts` — lease-aware orphan handling, OCR job re-dispatch - `src/app.ts` — disk multer, 202 async responses, `/ingestions/*` + review endpoints - `src/api/openapi.ts`, `public/playground/*` — new contracts and review UI - `src/config/env.ts`, `migrations/002_ocr_review.sql`, `Dockerfile`, `docs/API_RAG.md`, `docs/INGESTA.md` - New: `src/modules/ocr/{client,detection,artifacts,review}.ts`, `ocr-service/` (Python) ## Open Decisions Resolved | # | Decision | Resolution | Rationale | |---|---|---|---| | 1 | Page-level PDF extraction | **Replace `pdf-parse` with `pdfjs-dist` (legacy build) used directly**, wrapped in a new `parsePdfPages(buffer)` that returns `Array<{pageNumber, text}>`. Keep `pdf-parse` only if a spike proves its `pagerender` callback reliable on Node 22 — the contract permits either, but pdf.js v1.10.100 (2018) is too old to trust for the metrics contract. Fallback order: pdfjs-dist direct → `pdf-parse` with custom pagerender. Must be settled by a fixture-backed spike in Block 2 before anything downstream builds on it | Contract mandates "callback per pagina probado con fixture" and allows substitution. pdfjs-dist is the same engine, maintained, permissively licensed (Apache-2.0), and gives page objects with text items we need for metrics | | 2 | Job/page state machine & progress | **Reuse the contract schema exactly**: `rag_ocr_jobs` (one per version+document, `UNIQUE(version_id, document_id)`), `rag_document_pages` per page, dispatcher uses `FOR UPDATE SKIP LOCKED`, heartbeats. Progress = `completed_pages/totalPages` surfaced via `GET /ingestions/:versionId` from a single aggregate query over jobs+pages. No new states beyond contract | Schema is already specified in the contract (Punto 3 "Persistencia OCR"); inventing alternatives violates the closed decision | | 3 | Multi-document semantics | **Version-level candidate gate**: a version reaches `review_required` only when ALL documents' pages are native-sufficient or OCR-complete (contract quality gate is per-document AND per-version). One failed document fails the version (`OCR_QUALITY_BLOCKED` or the document's `error_code`). Mixed sources (md + pdf) in one folder ingest: only PDFs go through OCR; non-PDF docs keep native content_hash immediately, PDF docs get `content_hash` after review | The contract's gate ("el documento solo llega a review_required cuando…") is per-document; activation is version-level, so the join is forced | | 4 | Reconciler behavior | Extend `recoverOrphanedVersions` to distinguish: (a) versions in `indexing` with live OCR jobs (`lease_expires_at > now`) → leave alone; (b) expired leases → re-enqueue with same `remote_idempotency_key`, reset `next_attempt_at`; (c) stale native versions → existing behavior. The dispatcher itself (new module) owns lease renewal; reconciler only recovers dead leases | Contract: "Un reconciliador periodico reenvia con la misma clave remota y recupera remote_job_id sin duplicar OCR" | | 5 | Durable originals/artifacts | **Write original + manifest under `/data/ingestions//documents//` BEFORE `createPendingVersion`** (contract sequence: "se guarda el original durable; se crea la version…"). `documentArtifactId` = UUIDv5(document_id). If ingest dies between artifact write and version creation, a startup/periodic sweep deletes orphan version dirs (no catalog row = not referenceable). `artifact_state` column transitions: `none → present` at write, `present → retention_deleting → retention_deleted` by sweeper | Ordering is mandated by the contract's async section; sweep covers the one crash window | | 6 | Purge/retention | **Single daily sweeper** per contract: acquires the same version CAS/lock used by review and purge (`withVersionExclusiveLock`), sets `artifact_state='retention_deleting'`, deletes idempotently, marks `retention_deleted`. State TTLs: review_required 30d → `rejected`/`REVIEW_EXPIRED`; failed/rejected 7d → `purging→purged`; review images 7d post-decision; active keeps original+reviewed while active +30d. An active version is never TTL-deleted | All values are fixed by the contract's "Artefactos y retencion" section | | 7 | Worker topology | **In-process dispatcher inside the RAG Node service** (single instance, single EasyPanel app): a `setInterval`-driven poller that claims jobs with `FOR UPDATE SKIP LOCKED`, heartbeats, polls the OCR service, never blocks the ingest HTTP path beyond the initial enqueue. No separate Node worker deploy. The OCR service itself is one worker + one replica per contract | Adding a second Node worker would double deploy surface for zero benefit at current scale (≤3 queued jobs); contract fixes OCR service at 1 worker/1 replica, and the RAG API is the only writer | | 8 | OCR service/image boundary | **Exactly as contracted**: `ocr-service/` in-repo, FastAPI or plain HTTP Python app (implementation detail for design phase), PaddleOCR 3.4.0 + PaddlePaddle 3.2.2 CPU pinned, models baked at build, `/v1/jobs` multipart + Idempotency-Key, `/health/live` + `/health/ready`, internal network only, `OCR_INTERNAL_TOKEN` bearer, SQLite queue on ephemeral `/data/jobs` with 24h TTL, 50 MiB limit, config allowlist. RAG's `src/modules/ocr/client.ts` validates `documentSha256`, page sets, and result schema strictly | Closed decision in contract ("Despliegue OCR", "Contrato interno del servicio OCR") | | 9 | Durable uploads | Replace `multer.memoryStorage()` with disk storage into a temp staging dir, enforce 50 MiB + 100 pages + PDF sanity (reject encrypted/corrupt before full render), and move the original into the version artifact dir once identity is resolved | Contract: "Antes de habilitar OCR, sustituir multer.memoryStorage() por upload a disco con limite y limpieza garantizada" | | 10 | Representative fixtures | **Three committed fixtures** under `tests/fixtures/ocr/`: (1) `native.pdf` — text-rich, all pages pass detection (no OCR called); (2) `scanned.pdf` — image-only pages (2–3 pages, small) requiring OCR on all non-blank pages; (3) `mixed.pdf` — some native pages + some scanned pages, containing tokens shaped like the benchmark's risk identifiers (`CBG04a`-like). Plus unit fixtures: single-page blank image, oversized-rejection corpus. The real FacturaTech PDF stays out of the repo (client data); its 34-entry acceptance run happens in production per the contract's gate. Fixtures must be tiny (<200 KiB total) and generated with a script committed for reproducibility, so tests stay offline | `npm test` has no network; OCR service tests run inside its own Python test suite (pytest) — Node tests fake the OCR client. Contract acceptance criteria name the exact behaviors the fixtures must exercise | ## Approaches 1. **Contract-as-specified, in-process dispatcher (recommended)** — implement Punto 3 exactly as the contract defines: async ingest branch with 202, new tables 002, OCR client with retries/polling, artifacts manager, review endpoints, in-process job dispatcher. - Pros: zero contract amendments; all closed decisions honored; single deployable RAG change + one new OCR service; reconciler extension is small; rollback = feature-flag off. - Cons: dispatcher lifetime couples to RAG process (acceptable: single-instance deploy already assumed); ingest service grows substantially. - Effort: High (largest change in RAG history), but decomposable into 7 blocks. 2. **Separate Node worker process** — dispatcher runs as its own EasyPanel service consuming the same PostgreSQL queue. - Pros: isolates OCR polling latency; RAG API stays minimal. - Cons: violates the current single-service ops model; duplicate pool/lock code; two deploys to coordinate for what is currently ≤3 concurrent jobs; no requirement justifies it. - Effort: High+. 3. **Minimal sync OCR first, review later** — call OCR inline during ingest with a long timeout, add review workflow in a later change. - Pros: fewer moving parts initially. - Cons: directly violates the closed decision ("responde sin esperar los minutos de OCR", 202 async, review before activation); 245s benchmark runtime would hold HTTP connections; rejected. **Recommendation: Approach 1.** It is the only one consistent with the closed contract, and the in-process dispatcher keeps deployment identical in shape to today's (RAG + private OCR service). ## Seven Work Blocks Dependencies flow strictly downward; each block maps to a chained PR slice. | # | Block | Contents | Depends on | Key risk | |---|---|---|---|---| | 1 | Persistence & state machine | `migrations/002_ocr_review.sql`; repo methods: create OCRCandidateVersion (null-safe point counts), `markReviewRequired`, `approve→indexing`, `reject`, lease claim/renew methods; state-machine unit tests over fake pool | — | Migration 002 is additive (new tables + `extraction_method` values already exist); zero impact on live 001 data. Rollback: drop 002 tables (safe — no production data references them until OCR ships) | | 2 | Page-level PDF extraction | pdfjs-dist spike + `parsePdfPages`; detection metrics (N/A/W/R thresholds); candidate composition (`--- Page N ---` separators); canonical JSON; detection unit tests with fixtures | 1 (types only) | Library choice must be settled by spike evidence early; if pdfjs-dist legacy build fails on Node 22, fall back to `pdf-parse` custom pagerender | | 3 | PaddleOCR service | `ocr-service/` FastAPI app: jobs (SQLite queue, idempotency, TTL), engine wrapper (PaddleOCR 3.4.0 pinned), render pipeline (200 DPI, 25 MP cap), auth, limits, health; pytest suite with a stub engine | none (parallel-safe; contract fully specifies it) | Image build size/time (Paddle ~2 GB); models must bake at build; readiness false until models load | | 4 | Orchestration & artifacts | `ocr/client.ts` (retries, backoff 2–15s, strict validation); `ocr/artifacts.ts` (0600 writes, manifest, artifactId); in-process dispatcher (SKIP LOCKED, heartbeats); async ingest branch with 202 + `GET /ingestions/:versionId`; durable disk multer upload; reconciler lease-awareness | 1, 2, 3 | Biggest integration surface: touches ingest, app, reconciler; must not regress native sync path (regression tests required) | | 5 | Review, corrections, indexing | Risk-token detection (regex + priority rules); `GET /ingestions/:versionId/review`; approve with `candidateSha256` + per-line `expectedLineSha256` corrections (all-or-nothing 409); duplicate-reusable check; post-approval chunk/embed/index/activate via existing point-2 path; playground review UI | 4 | Correction collision semantics (409 without partial application) must be exact; concurrency with activation changes | | 6 | Security, retention, deployment | Retention sweeper (TTLs, artifact_state CAS); `/health` OCR section; OpenAPI full update (`202`, `/ingestions/*`, review routes, bearer); Dockerfile volume + OCR compose config; docs (API_RAG, INGESTA, deploy); no-secret review | 4, 5 | Deployment ordering: OCR service first, RAG with OCR disabled, then enable — irreversible risk is low if ordering respected | | 7 | End-to-end validation | Fixture-driven e2e (native/scanned/mixed), idempotent-resend test, fail-closed test (OCR down), full static suite (`npm run check`, `npm run build`, Node tests + pytest), production validation plan for the real FacturaTech PDF | 1–6 | Production acceptance (34 entries) requires human review step; cannot be automated | Boundary adjustments made (with evidence): the contract's own task order lists "añadir migracion OCR y contratos TypeScript" first and extraction second, but **Block 3 (OCR service) has no code dependency on Blocks 1–2** — the contract fully specifies its HTTP interface. Starting it in parallel is the only way to fit the overall work into review-budget-sized slices without artificial sequencing. This is an execution-order adjustment only; all contract sequencing gates (OCR deployed before RAG enables it) remain. ## Dependencies & Risk Map - **Hard dependency chain:** 1 → 2 → 4 → 5 → 6 → 7; 3 is parallel, joins at 4. - **Irreversible operations:** only `migrations/002_ocr_review.sql` (additive DDL — new tables, no ALTER on existing ones; safe rollback = `DROP TABLE rag_ocr_jobs, rag_document_pages, rag_review_corrections` + `DELETE FROM rag_schema_migrations WHERE name='002_ocr_review.sql'`). Qdrant unchanged until approval-driven indexing, which uses existing versioned point IDs. Production artifact volume starts empty; retention never deletes active versions. - **Rollback boundary:** feature-flag `OCR_INGEST_ENABLED` (new env). Off = today's native-only sync behavior byte-for-byte (regression suite proves it). The 202 async path and `/ingestions/*` routes exist but return `503 OCR_DISABLED`-style errors when off — or routes can be conditionally mounted; design phase decides. Data written while enabled (candidates, jobs) is invisible to retrieval by construction (never `active` without review). - **Deploy ordering (contract-mandated):** OCR service private deploy + readiness check → RAG deploy with OCR disabled → verify native regressions → enable flag → run real PDF validation. - **Review-budget pressure (400 lines, auto-chain):** **High across the whole change.** Realistic estimate: Block 1 ~450–550 lines (SQL+repo+tests), Block 2 ~350–450, Block 3 ~600–800 (Python, reviewed separately from Node PRs), Block 4 ~700–900, Block 5 ~600–750, Block 6 ~350–450, Block 7 ~250–350. Every block except possibly Block 7 individually risks exceeding 400 changed lines; chained/stacked PRs are mandatory, and Blocks 4 and 5 may each need two slices (e.g., 4a client+artifacts, 4b dispatcher+ingest async). ## Risks - **pdfjs-dist integration uncertainty on Node 22** (legacy vs modern build, worker config) — mitigate with a day-one spike in Block 2, contract already allows the `pdf-parse` fallback. - **OCR image build fragility**: PaddlePaddle 3.2.2 CPU wheels are large and platform-sensitive; baked models add build time. Mitigate: pin exact versions, build once, keep ocr-service deploy independent (contract requirement). - **Ingest regression on the native path**: the async branch rewrites the hottest service. Mitigate: strict TDD per config (strict_tdd: true), the 25 existing tests must stay green, plus new native-regression tests before any OCR wiring. - **State-machine drift from the contract**: the contract's diagram is normative; any deviation needs contract amendment FIRST. Mitigation: spec phase encodes each transition as Given/When/Then. - **Playground review UI scope creep**: image rendering, side-by-side diffs, corrections editor could balloon. Mitigation: MVP review UI = native/OCR text + corrections JSON input + image links; defer polish (the playground doc says it's an internal tool). - **VPS2 resource contention**: OCR service capped at 3 CPU/5 GiB of the 6 vCPU/11 GiB host; a 25-page/245s benchmark job can saturate one core for minutes. Mitigate: contract's queue depth 3, single worker, and 60s/page timeout are already sized for this. ## Ready for Proposal **Yes.** Every open decision above is resolved within the closed contract; no human input is required to draft the proposal. The proposal must state: contract-conformant implementation (Approach 1), seven chained blocks as scoped, Block-3 parallelism, feature-flag rollback, and the deploy ordering. Contract amendments are NOT needed — the contract's Punto 3 already anticipates all resolved details; the only judgment calls (worker topology, fixture strategy, library spike order) are implementation-level and recorded here.