rag-service/openspec/changes/archive/2026-09-21-ocr-ingest-integration/exploration.md

20 KiB
Raw Permalink Blame History

Exploration: OCR ingest integration (ocr-ingest-integration)

Project: rag-service — Phase: sdd-explore — Date: 2026-09-13 Store: hybrid (OpenSpec file + Engram topic sdd/ocr-ingest-integration/explore)

Current State

Point 2 (knowledge lifecycle) is complete, enforced in production (KNOWLEDGE_LIFECYCLE_ENFORCED=true, 7 cataloged sources, 22,605 active points, legacy batch migrated and verified 2026-09-13). npm test passes 25/25. The contract in docs/CONTRATO_CICLO_VIDA_Y_OCR.md (section "Punto 3") is closed and authoritative: every decision below defers to it; this exploration only resolves what the contract leaves open.

Verified implementation facts (read from code, not docs):

Area Current behavior OCR gap
Ingest IngestService.ingestWithLifecycle is fully synchronous: attempt → source lock → discover → hash → reusable lookup → createPendingVersion → markIndexing → embed → upsert → count → markReady → activateVersion, all inside the HTTP request Must become async-capable: durable original first, 202 + statusUrl response when OCR is needed, source_content_hash = null until review completes
PDF parsing parser-registry.parseDocument uses pdf-parse over the whole file, returns aggregated text only; no page callback in our code (library has pagerender option but we never use it; bundled pdf.js is v1.10.100, very old) Page-level extraction required; contract allows replacing pdf-parse if unreliable
State machine SourceVersionState already includes review_required and rejected; transitions pending→indexing→ready→active, failed, purging→purged implemented in CatalogRepository; no transition writes into review_required today Need: indexing→review_required (OCR candidate complete), review_required→indexing (approve), review_required→rejected, plus review_required/rejected entries in markPurging-eligible states (already present), retention expiry transitions
Catalog repo createPendingVersion requires non-null expectedPointCount and writes all documents with content_hash pre-computed; findReusableVersion skips when sourceContentHash is null (correct for OCR candidates); hasCompleteVersionDocuments counts content_hash IS NULL as incomplete OCR flow must create a pending version before content exists (chunk counts unknown until review) — needs a creation mode with expectedPointCount=0/null-safe path, or a separate "OCR candidate" insert path
Reconciler runOnce treats stale pending/indexing older than 30 min as orphans and either recovers to ready or fails them Must not kill in-flight OCR versions: needs heartbeat awareness (jobs have heartbeat_at/lease_expires_at per contract) and must recover expired OCR leases by re-dispatching with the same remote idempotency key
Uploads multer.memoryStorage() in src/app.ts, no size limit; upload writes to /tmp with timestamp, deleted in finally — no durable original Contract requires disk-based multer with limits and durable original under /data/ingestions/<versionId>/ before processing
Env/config env.ts has PostgreSQL + lifecycle tokens only Needs OCR_SERVICE_URL, OCR_INTERNAL_TOKEN, artifact volume path, limits (50 MiB, 100 pages, timeouts), retention intervals
Migrations migrations/001_knowledge_lifecycle.sql applied via rag_schema_migrations with checksums + advisory lock; migration runner sorts ^\d+_.*\.sql$ 002_ocr_review.sql fits the existing runner unchanged (tables rag_ocr_jobs, rag_document_pages, rag_review_corrections per contract)
OpenAPI IngestResponse is 201-only under lifecycle; securitySchemes.bearerAuth already defined and used by lifecycle endpoints Needs 202 variant with phase/statusUrl, /ingestions/:versionId status route, review/approve/reject routes
Playground public/playground/app.js supports activate/expectedActiveVersionId; no review UI Review view (page images, native vs OCR text, corrections) required by acceptance criteria
Docker Node runtime only; CMD ["node", "dist/server.js"] Needs /data/ingestions volume; ocr-service gets its own Dockerfile (Python CPU, PaddleOCR 3.4.0 + PaddlePaddle 3.2.2 pinned, baked models)
Artifacts No filesystem artifact manager exists; rag_version_documents.artifact_manifest_path and artifact_state/retention_due_at columns exist but are never written (retention_deleted only set on purge) Whole artifact subsystem (write/verify 0600, manifest, retention sweeper) is new

Affected Areas

  • src/modules/ingest/service.ts — async OCR branch, durable originals, version creation before content
  • src/modules/parsers/parser-registry.ts — per-page PDF extraction API
  • src/modules/catalog/repository.ts — new state transitions, OCR-safe version creation, page/job/correction persistence
  • src/modules/catalog/reconciler.ts — lease-aware orphan handling, OCR job re-dispatch
  • src/app.ts — disk multer, 202 async responses, /ingestions/* + review endpoints
  • src/api/openapi.ts, public/playground/* — new contracts and review UI
  • src/config/env.ts, migrations/002_ocr_review.sql, Dockerfile, docs/API_RAG.md, docs/INGESTA.md
  • New: src/modules/ocr/{client,detection,artifacts,review}.ts, ocr-service/ (Python)

Open Decisions Resolved

# Decision Resolution Rationale
1 Page-level PDF extraction Replace pdf-parse with pdfjs-dist (legacy build) used directly, wrapped in a new parsePdfPages(buffer) that returns Array<{pageNumber, text}>. Keep pdf-parse only if a spike proves its pagerender callback reliable on Node 22 — the contract permits either, but pdf.js v1.10.100 (2018) is too old to trust for the metrics contract. Fallback order: pdfjs-dist direct → pdf-parse with custom pagerender. Must be settled by a fixture-backed spike in Block 2 before anything downstream builds on it Contract mandates "callback per pagina probado con fixture" and allows substitution. pdfjs-dist is the same engine, maintained, permissively licensed (Apache-2.0), and gives page objects with text items we need for metrics
2 Job/page state machine & progress Reuse the contract schema exactly: rag_ocr_jobs (one per version+document, UNIQUE(version_id, document_id)), rag_document_pages per page, dispatcher uses FOR UPDATE SKIP LOCKED, heartbeats. Progress = completed_pages/totalPages surfaced via GET /ingestions/:versionId from a single aggregate query over jobs+pages. No new states beyond contract Schema is already specified in the contract (Punto 3 "Persistencia OCR"); inventing alternatives violates the closed decision
3 Multi-document semantics Version-level candidate gate: a version reaches review_required only when ALL documents' pages are native-sufficient or OCR-complete (contract quality gate is per-document AND per-version). One failed document fails the version (OCR_QUALITY_BLOCKED or the document's error_code). Mixed sources (md + pdf) in one folder ingest: only PDFs go through OCR; non-PDF docs keep native content_hash immediately, PDF docs get content_hash after review The contract's gate ("el documento solo llega a review_required cuando…") is per-document; activation is version-level, so the join is forced
4 Reconciler behavior Extend recoverOrphanedVersions to distinguish: (a) versions in indexing with live OCR jobs (lease_expires_at > now) → leave alone; (b) expired leases → re-enqueue with same remote_idempotency_key, reset next_attempt_at; (c) stale native versions → existing behavior. The dispatcher itself (new module) owns lease renewal; reconciler only recovers dead leases Contract: "Un reconciliador periodico reenvia con la misma clave remota y recupera remote_job_id sin duplicar OCR"
5 Durable originals/artifacts Write original + manifest under /data/ingestions/<versionId>/documents/<documentArtifactId>/ BEFORE createPendingVersion (contract sequence: "se guarda el original durable; se crea la version…"). documentArtifactId = UUIDv5(document_id). If ingest dies between artifact write and version creation, a startup/periodic sweep deletes orphan version dirs (no catalog row = not referenceable). artifact_state column transitions: none → present at write, present → retention_deleting → retention_deleted by sweeper Ordering is mandated by the contract's async section; sweep covers the one crash window
6 Purge/retention Single daily sweeper per contract: acquires the same version CAS/lock used by review and purge (withVersionExclusiveLock), sets artifact_state='retention_deleting', deletes idempotently, marks retention_deleted. State TTLs: review_required 30d → rejected/REVIEW_EXPIRED; failed/rejected 7d → purging→purged; review images 7d post-decision; active keeps original+reviewed while active +30d. An active version is never TTL-deleted All values are fixed by the contract's "Artefactos y retencion" section
7 Worker topology In-process dispatcher inside the RAG Node service (single instance, single EasyPanel app): a setInterval-driven poller that claims jobs with FOR UPDATE SKIP LOCKED, heartbeats, polls the OCR service, never blocks the ingest HTTP path beyond the initial enqueue. No separate Node worker deploy. The OCR service itself is one worker + one replica per contract Adding a second Node worker would double deploy surface for zero benefit at current scale (≤3 queued jobs); contract fixes OCR service at 1 worker/1 replica, and the RAG API is the only writer
8 OCR service/image boundary Exactly as contracted: ocr-service/ in-repo, FastAPI or plain HTTP Python app (implementation detail for design phase), PaddleOCR 3.4.0 + PaddlePaddle 3.2.2 CPU pinned, models baked at build, /v1/jobs multipart + Idempotency-Key, /health/live + /health/ready, internal network only, OCR_INTERNAL_TOKEN bearer, SQLite queue on ephemeral /data/jobs with 24h TTL, 50 MiB limit, config allowlist. RAG's src/modules/ocr/client.ts validates documentSha256, page sets, and result schema strictly Closed decision in contract ("Despliegue OCR", "Contrato interno del servicio OCR")
9 Durable uploads Replace multer.memoryStorage() with disk storage into a temp staging dir, enforce 50 MiB + 100 pages + PDF sanity (reject encrypted/corrupt before full render), and move the original into the version artifact dir once identity is resolved Contract: "Antes de habilitar OCR, sustituir multer.memoryStorage() por upload a disco con limite y limpieza garantizada"
10 Representative fixtures Three committed fixtures under tests/fixtures/ocr/: (1) native.pdf — text-rich, all pages pass detection (no OCR called); (2) scanned.pdf — image-only pages (2–3 pages, small) requiring OCR on all non-blank pages; (3) mixed.pdf — some native pages + some scanned pages, containing tokens shaped like the benchmark's risk identifiers (CBG04a-like). Plus unit fixtures: single-page blank image, oversized-rejection corpus. The real FacturaTech PDF stays out of the repo (client data); its 34-entry acceptance run happens in production per the contract's gate. Fixtures must be tiny (<200 KiB total) and generated with a script committed for reproducibility, so tests stay offline npm test has no network; OCR service tests run inside its own Python test suite (pytest) — Node tests fake the OCR client. Contract acceptance criteria name the exact behaviors the fixtures must exercise

Approaches

  1. Contract-as-specified, in-process dispatcher (recommended) — implement Punto 3 exactly as the contract defines: async ingest branch with 202, new tables 002, OCR client with retries/polling, artifacts manager, review endpoints, in-process job dispatcher.

    • Pros: zero contract amendments; all closed decisions honored; single deployable RAG change + one new OCR service; reconciler extension is small; rollback = feature-flag off.
    • Cons: dispatcher lifetime couples to RAG process (acceptable: single-instance deploy already assumed); ingest service grows substantially.
    • Effort: High (largest change in RAG history), but decomposable into 7 blocks.
  2. Separate Node worker process — dispatcher runs as its own EasyPanel service consuming the same PostgreSQL queue.

    • Pros: isolates OCR polling latency; RAG API stays minimal.
    • Cons: violates the current single-service ops model; duplicate pool/lock code; two deploys to coordinate for what is currently ≤3 concurrent jobs; no requirement justifies it.
    • Effort: High+.
  3. Minimal sync OCR first, review later — call OCR inline during ingest with a long timeout, add review workflow in a later change.

    • Pros: fewer moving parts initially.
    • Cons: directly violates the closed decision ("responde sin esperar los minutos de OCR", 202 async, review before activation); 245s benchmark runtime would hold HTTP connections; rejected.

Recommendation: Approach 1. It is the only one consistent with the closed contract, and the in-process dispatcher keeps deployment identical in shape to today's (RAG + private OCR service).

Seven Work Blocks

Dependencies flow strictly downward; each block maps to a chained PR slice.

# Block Contents Depends on Key risk
1 Persistence & state machine migrations/002_ocr_review.sql; repo methods: create OCRCandidateVersion (null-safe point counts), markReviewRequired, approve→indexing, reject, lease claim/renew methods; state-machine unit tests over fake pool — Migration 002 is additive (new tables + extraction_method values already exist); zero impact on live 001 data. Rollback: drop 002 tables (safe — no production data references them until OCR ships)
2 Page-level PDF extraction pdfjs-dist spike + parsePdfPages; detection metrics (N/A/W/R thresholds); candidate composition (--- Page N --- separators); canonical JSON; detection unit tests with fixtures 1 (types only) Library choice must be settled by spike evidence early; if pdfjs-dist legacy build fails on Node 22, fall back to pdf-parse custom pagerender
3 PaddleOCR service ocr-service/ FastAPI app: jobs (SQLite queue, idempotency, TTL), engine wrapper (PaddleOCR 3.4.0 pinned), render pipeline (200 DPI, 25 MP cap), auth, limits, health; pytest suite with a stub engine none (parallel-safe; contract fully specifies it) Image build size/time (Paddle ~2 GB); models must bake at build; readiness false until models load
4 Orchestration & artifacts ocr/client.ts (retries, backoff 2–15s, strict validation); ocr/artifacts.ts (0600 writes, manifest, artifactId); in-process dispatcher (SKIP LOCKED, heartbeats); async ingest branch with 202 + GET /ingestions/:versionId; durable disk multer upload; reconciler lease-awareness 1, 2, 3 Biggest integration surface: touches ingest, app, reconciler; must not regress native sync path (regression tests required)
5 Review, corrections, indexing Risk-token detection (regex + priority rules); GET /ingestions/:versionId/review; approve with candidateSha256 + per-line expectedLineSha256 corrections (all-or-nothing 409); duplicate-reusable check; post-approval chunk/embed/index/activate via existing point-2 path; playground review UI 4 Correction collision semantics (409 without partial application) must be exact; concurrency with activation changes
6 Security, retention, deployment Retention sweeper (TTLs, artifact_state CAS); /health OCR section; OpenAPI full update (202, /ingestions/*, review routes, bearer); Dockerfile volume + OCR compose config; docs (API_RAG, INGESTA, deploy); no-secret review 4, 5 Deployment ordering: OCR service first, RAG with OCR disabled, then enable — irreversible risk is low if ordering respected
7 End-to-end validation Fixture-driven e2e (native/scanned/mixed), idempotent-resend test, fail-closed test (OCR down), full static suite (npm run check, npm run build, Node tests + pytest), production validation plan for the real FacturaTech PDF 1–6 Production acceptance (34 entries) requires human review step; cannot be automated

Boundary adjustments made (with evidence): the contract's own task order lists "añadir migracion OCR y contratos TypeScript" first and extraction second, but Block 3 (OCR service) has no code dependency on Blocks 1–2 — the contract fully specifies its HTTP interface. Starting it in parallel is the only way to fit the overall work into review-budget-sized slices without artificial sequencing. This is an execution-order adjustment only; all contract sequencing gates (OCR deployed before RAG enables it) remain.

Dependencies & Risk Map

  • Hard dependency chain: 1 → 2 → 4 → 5 → 6 → 7; 3 is parallel, joins at 4.
  • Irreversible operations: only migrations/002_ocr_review.sql (additive DDL — new tables, no ALTER on existing ones; safe rollback = DROP TABLE rag_ocr_jobs, rag_document_pages, rag_review_corrections + DELETE FROM rag_schema_migrations WHERE name='002_ocr_review.sql'). Qdrant unchanged until approval-driven indexing, which uses existing versioned point IDs. Production artifact volume starts empty; retention never deletes active versions.
  • Rollback boundary: feature-flag OCR_INGEST_ENABLED (new env). Off = today's native-only sync behavior byte-for-byte (regression suite proves it). The 202 async path and /ingestions/* routes exist but return 503 OCR_DISABLED-style errors when off — or routes can be conditionally mounted; design phase decides. Data written while enabled (candidates, jobs) is invisible to retrieval by construction (never active without review).
  • Deploy ordering (contract-mandated): OCR service private deploy + readiness check → RAG deploy with OCR disabled → verify native regressions → enable flag → run real PDF validation.
  • Review-budget pressure (400 lines, auto-chain): High across the whole change. Realistic estimate: Block 1 ~450–550 lines (SQL+repo+tests), Block 2 ~350–450, Block 3 ~600–800 (Python, reviewed separately from Node PRs), Block 4 ~700–900, Block 5 ~600–750, Block 6 ~350–450, Block 7 ~250–350. Every block except possibly Block 7 individually risks exceeding 400 changed lines; chained/stacked PRs are mandatory, and Blocks 4 and 5 may each need two slices (e.g., 4a client+artifacts, 4b dispatcher+ingest async).

Risks

  • pdfjs-dist integration uncertainty on Node 22 (legacy vs modern build, worker config) — mitigate with a day-one spike in Block 2, contract already allows the pdf-parse fallback.
  • OCR image build fragility: PaddlePaddle 3.2.2 CPU wheels are large and platform-sensitive; baked models add build time. Mitigate: pin exact versions, build once, keep ocr-service deploy independent (contract requirement).
  • Ingest regression on the native path: the async branch rewrites the hottest service. Mitigate: strict TDD per config (strict_tdd: true), the 25 existing tests must stay green, plus new native-regression tests before any OCR wiring.
  • State-machine drift from the contract: the contract's diagram is normative; any deviation needs contract amendment FIRST. Mitigation: spec phase encodes each transition as Given/When/Then.
  • Playground review UI scope creep: image rendering, side-by-side diffs, corrections editor could balloon. Mitigation: MVP review UI = native/OCR text + corrections JSON input + image links; defer polish (the playground doc says it's an internal tool).
  • VPS2 resource contention: OCR service capped at 3 CPU/5 GiB of the 6 vCPU/11 GiB host; a 25-page/245s benchmark job can saturate one core for minutes. Mitigate: contract's queue depth 3, single worker, and 60s/page timeout are already sized for this.

Ready for Proposal

Yes. Every open decision above is resolved within the closed contract; no human input is required to draft the proposal. The proposal must state: contract-conformant implementation (Approach 1), seven chained blocks as scoped, Block-3 parallelism, feature-flag rollback, and the deploy ordering. Contract amendments are NOT needed — the contract's Punto 3 already anticipates all resolved details; the only judgment calls (worker topology, fixture strategy, library spike order) are implementation-level and recorded here.