# Proposal: OCR Ingest Integration (`ocr-ingest-integration`) **Project:** rag-service — **Phase:** sdd-propose — 2026-09-13 **Store:** hybrid (OpenSpec + Engram `sdd/ocr-ingest-integration/proposal`) **Contract:** `docs/CONTRATO_CICLO_VIDA_Y_OCR.md` Punto 3 (closed) ## Intent Auditable OCR ingestion for scanned PDFs: durable originals, async processing behind a review gate, approval before OCR content enters retrieval. Native ingestion and Punto 2 must not regress. ## Scope **In:** async OCR branch (`202 + statusUrl`, `GET /ingestions/:versionId`, version-level review gate); migration `002_ocr_review.sql` with transitions `indexing→review_required→indexing|rejected`; page-level PDF extraction (pdfjs-dist spike; `pdf-parse` fallback), detection; private `ocr-service/` (Python, PaddleOCR pinned, idempotent `/v1/jobs`); dispatcher, OCR client, artifacts, lease-aware reconciler; review endpoints (`candidateSha256`, corrections, 409), sweeper, review UI; `OCR_INGEST_ENABLED` flag, limits, deploy order, docs, fixtures. **Out:** client PDFs in repo; UI polish; extra Node workers; contract amendments. FacturaTech acceptance: production. ## Capabilities > `openspec/specs/` is empty — all new. - `ocr-ingest-orchestration`: async ingestion, durable originals, job/page state machine, dispatcher, reconciliation - `ocr-processing`: page extraction, detection, composition, service contract - `ocr-review-workflow`: review gate, corrections, approval, retention Modified: none. ## Approach Contract-as-specified (exploration Approach 1): in-process Node dispatcher plus private PaddleOCR service. Blocks: 1) persistence/state, 2) page extraction, 3) OCR service (parallel; joins at 4), 4) orchestration/artifacts, 5) review/indexing, 6) security/retention/deploy, 7) e2e validation. Durable originals precede version creation; unapproved candidates stay out of retrieval. TDD. ## Affected Areas | Area | Impact | |------|--------| | `src/modules/ingest/service.ts`, `parsers/parser-registry.ts` | Modified: async OCR branch; per-page `parsePdfPages` | | `src/modules/catalog/{repository,reconciler}.ts` | Modified: new transitions; lease recovery | | `src/app.ts`, `src/api/openapi.ts`, `public/playground/*` | Modified: disk multer; 202; `/ingestions/*`; review UI | | `src/modules/ocr/`, `ocr-service/`, `migrations/002_ocr_review.sql`, `src/config/env.ts`, `Dockerfile`, docs | New | ## Risks | Risk | Likelihood | Mitigation | |------|------------|------------| | Native ingest regression | Medium | TDD; native-regression tests first | | pdfjs-dist on Node 22 | Medium | Day-one spike; contract fallback | | OCR image fragility | Medium | Pinned versions; baked models | | Review budget overrun | High | Auto-chain; split blocks 4–5 | ## Rollback Plan `OCR_INGEST_ENABLED=false` restores native-only behavior. Migration 002 rollback: drop the new tables plus migration row (no production data references them). Qdrant untouched until approval; candidates invisible to retrieval. ## Dependencies Contract Punto 3; Punto 2 lifecycle (in production). Deploy: OCR service → RAG (flag off) → native verify → enable → acceptance. ## Success Criteria - [ ] Native PDFs ingest sync, zero behavior change; tests green - [ ] Scanned PDF: `202` → OCR → `review_required` → approve → active - [ ] Candidate never retrievable pre-approval (fail-closed) - [ ] Expired leases redispatched idempotently; corrections 409 atomic - [ ] Sweeper honors TTLs; active versions never deleted - [ ] `npm run check`, `npm run build`, `npm test`, pytest offline - [ ] FacturaTech: 34 entries human-approved