docs(sdd): add OCR integration plan
This commit is contained in:
parent
270f693587
commit
11139bd14e
6 changed files with 309 additions and 0 deletions
0
openspec/changes/archive/.gitkeep
Normal file
0
openspec/changes/archive/.gitkeep
Normal file
106
openspec/changes/ocr-ingest-integration/exploration.md
Normal file
106
openspec/changes/ocr-ingest-integration/exploration.md
Normal file
|
|
@ -0,0 +1,106 @@
|
|||
# Exploration: OCR ingest integration (`ocr-ingest-integration`)
|
||||
|
||||
**Project:** rag-service — **Phase:** sdd-explore — **Date:** 2026-09-13
|
||||
**Store:** hybrid (OpenSpec file + Engram topic `sdd/ocr-ingest-integration/explore`)
|
||||
|
||||
## Current State
|
||||
|
||||
Point 2 (knowledge lifecycle) is complete, enforced in production (`KNOWLEDGE_LIFECYCLE_ENFORCED=true`, 7 cataloged sources, 22,605 active points, legacy batch migrated and verified 2026-09-13). `npm test` passes 25/25. The contract in `docs/CONTRATO_CICLO_VIDA_Y_OCR.md` (section "Punto 3") is closed and authoritative: every decision below defers to it; this exploration only resolves what the contract leaves open.
|
||||
|
||||
Verified implementation facts (read from code, not docs):
|
||||
|
||||
| Area | Current behavior | OCR gap |
|
||||
|---|---|---|
|
||||
| Ingest | `IngestService.ingestWithLifecycle` is fully synchronous: attempt → source lock → discover → hash → reusable lookup → `createPendingVersion` → `markIndexing` → embed → upsert → count → `markReady` → `activateVersion`, all inside the HTTP request | Must become async-capable: durable original first, `202 + statusUrl` response when OCR is needed, `source_content_hash = null` until review completes |
|
||||
| PDF parsing | `parser-registry.parseDocument` uses `pdf-parse` over the whole file, returns aggregated text only; no page callback in our code (library has `pagerender` option but we never use it; bundled pdf.js is v1.10.100, very old) | Page-level extraction required; contract allows replacing `pdf-parse` if unreliable |
|
||||
| State machine | `SourceVersionState` already includes `review_required` and `rejected`; transitions `pending→indexing→ready→active`, `failed`, `purging→purged` implemented in `CatalogRepository`; **no transition writes into `review_required` today** | Need: `indexing→review_required` (OCR candidate complete), `review_required→indexing` (approve), `review_required→rejected`, plus `review_required`/`rejected` entries in `markPurging`-eligible states (already present), retention expiry transitions |
|
||||
| Catalog repo | `createPendingVersion` requires non-null `expectedPointCount` and writes all documents with `content_hash` pre-computed; `findReusableVersion` skips when `sourceContentHash` is null (correct for OCR candidates); `hasCompleteVersionDocuments` counts `content_hash IS NULL` as incomplete | OCR flow must create a pending version **before** content exists (chunk counts unknown until review) — needs a creation mode with `expectedPointCount=0`/null-safe path, or a separate "OCR candidate" insert path |
|
||||
| Reconciler | `runOnce` treats stale `pending/indexing` older than 30 min as orphans and either recovers to `ready` or fails them | Must not kill in-flight OCR versions: needs heartbeat awareness (jobs have `heartbeat_at`/`lease_expires_at` per contract) and must recover expired OCR leases by re-dispatching with the same remote idempotency key |
|
||||
| Uploads | `multer.memoryStorage()` in `src/app.ts`, no size limit; upload writes to `/tmp` with timestamp, deleted in `finally` — no durable original | Contract requires disk-based multer with limits and durable original under `/data/ingestions/<versionId>/` before processing |
|
||||
| Env/config | `env.ts` has PostgreSQL + lifecycle tokens only | Needs `OCR_SERVICE_URL`, `OCR_INTERNAL_TOKEN`, artifact volume path, limits (50 MiB, 100 pages, timeouts), retention intervals |
|
||||
| Migrations | `migrations/001_knowledge_lifecycle.sql` applied via `rag_schema_migrations` with checksums + advisory lock; migration runner sorts `^\d+_.*\.sql$` | `002_ocr_review.sql` fits the existing runner unchanged (tables `rag_ocr_jobs`, `rag_document_pages`, `rag_review_corrections` per contract) |
|
||||
| OpenAPI | `IngestResponse` is 201-only under lifecycle; `securitySchemes.bearerAuth` already defined and used by lifecycle endpoints | Needs `202` variant with `phase`/`statusUrl`, `/ingestions/:versionId` status route, review/approve/reject routes |
|
||||
| Playground | `public/playground/app.js` supports activate/expectedActiveVersionId; no review UI | Review view (page images, native vs OCR text, corrections) required by acceptance criteria |
|
||||
| Docker | Node runtime only; `CMD ["node", "dist/server.js"]` | Needs `/data/ingestions` volume; ocr-service gets its own Dockerfile (Python CPU, PaddleOCR 3.4.0 + PaddlePaddle 3.2.2 pinned, baked models) |
|
||||
| Artifacts | No filesystem artifact manager exists; `rag_version_documents.artifact_manifest_path` and `artifact_state`/`retention_due_at` columns exist but are never written (`retention_deleted` only set on purge) | Whole artifact subsystem (write/verify 0600, manifest, retention sweeper) is new |
|
||||
|
||||
## Affected Areas
|
||||
|
||||
- `src/modules/ingest/service.ts` — async OCR branch, durable originals, version creation before content
|
||||
- `src/modules/parsers/parser-registry.ts` — per-page PDF extraction API
|
||||
- `src/modules/catalog/repository.ts` — new state transitions, OCR-safe version creation, page/job/correction persistence
|
||||
- `src/modules/catalog/reconciler.ts` — lease-aware orphan handling, OCR job re-dispatch
|
||||
- `src/app.ts` — disk multer, 202 async responses, `/ingestions/*` + review endpoints
|
||||
- `src/api/openapi.ts`, `public/playground/*` — new contracts and review UI
|
||||
- `src/config/env.ts`, `migrations/002_ocr_review.sql`, `Dockerfile`, `docs/API_RAG.md`, `docs/INGESTA.md`
|
||||
- New: `src/modules/ocr/{client,detection,artifacts,review}.ts`, `ocr-service/` (Python)
|
||||
|
||||
## Open Decisions Resolved
|
||||
|
||||
| # | Decision | Resolution | Rationale |
|
||||
|---|---|---|---|
|
||||
| 1 | Page-level PDF extraction | **Replace `pdf-parse` with `pdfjs-dist` (legacy build) used directly**, wrapped in a new `parsePdfPages(buffer)` that returns `Array<{pageNumber, text}>`. Keep `pdf-parse` only if a spike proves its `pagerender` callback reliable on Node 22 — the contract permits either, but pdf.js v1.10.100 (2018) is too old to trust for the metrics contract. Fallback order: pdfjs-dist direct → `pdf-parse` with custom pagerender. Must be settled by a fixture-backed spike in Block 2 before anything downstream builds on it | Contract mandates "callback per pagina probado con fixture" and allows substitution. pdfjs-dist is the same engine, maintained, permissively licensed (Apache-2.0), and gives page objects with text items we need for metrics |
|
||||
| 2 | Job/page state machine & progress | **Reuse the contract schema exactly**: `rag_ocr_jobs` (one per version+document, `UNIQUE(version_id, document_id)`), `rag_document_pages` per page, dispatcher uses `FOR UPDATE SKIP LOCKED`, heartbeats. Progress = `completed_pages/totalPages` surfaced via `GET /ingestions/:versionId` from a single aggregate query over jobs+pages. No new states beyond contract | Schema is already specified in the contract (Punto 3 "Persistencia OCR"); inventing alternatives violates the closed decision |
|
||||
| 3 | Multi-document semantics | **Version-level candidate gate**: a version reaches `review_required` only when ALL documents' pages are native-sufficient or OCR-complete (contract quality gate is per-document AND per-version). One failed document fails the version (`OCR_QUALITY_BLOCKED` or the document's `error_code`). Mixed sources (md + pdf) in one folder ingest: only PDFs go through OCR; non-PDF docs keep native content_hash immediately, PDF docs get `content_hash` after review | The contract's gate ("el documento solo llega a review_required cuando…") is per-document; activation is version-level, so the join is forced |
|
||||
| 4 | Reconciler behavior | Extend `recoverOrphanedVersions` to distinguish: (a) versions in `indexing` with live OCR jobs (`lease_expires_at > now`) → leave alone; (b) expired leases → re-enqueue with same `remote_idempotency_key`, reset `next_attempt_at`; (c) stale native versions → existing behavior. The dispatcher itself (new module) owns lease renewal; reconciler only recovers dead leases | Contract: "Un reconciliador periodico reenvia con la misma clave remota y recupera remote_job_id sin duplicar OCR" |
|
||||
| 5 | Durable originals/artifacts | **Write original + manifest under `/data/ingestions/<versionId>/documents/<documentArtifactId>/` BEFORE `createPendingVersion`** (contract sequence: "se guarda el original durable; se crea la version…"). `documentArtifactId` = UUIDv5(document_id). If ingest dies between artifact write and version creation, a startup/periodic sweep deletes orphan version dirs (no catalog row = not referenceable). `artifact_state` column transitions: `none → present` at write, `present → retention_deleting → retention_deleted` by sweeper | Ordering is mandated by the contract's async section; sweep covers the one crash window |
|
||||
| 6 | Purge/retention | **Single daily sweeper** per contract: acquires the same version CAS/lock used by review and purge (`withVersionExclusiveLock`), sets `artifact_state='retention_deleting'`, deletes idempotently, marks `retention_deleted`. State TTLs: review_required 30d → `rejected`/`REVIEW_EXPIRED`; failed/rejected 7d → `purging→purged`; review images 7d post-decision; active keeps original+reviewed while active +30d. An active version is never TTL-deleted | All values are fixed by the contract's "Artefactos y retencion" section |
|
||||
| 7 | Worker topology | **In-process dispatcher inside the RAG Node service** (single instance, single EasyPanel app): a `setInterval`-driven poller that claims jobs with `FOR UPDATE SKIP LOCKED`, heartbeats, polls the OCR service, never blocks the ingest HTTP path beyond the initial enqueue. No separate Node worker deploy. The OCR service itself is one worker + one replica per contract | Adding a second Node worker would double deploy surface for zero benefit at current scale (≤3 queued jobs); contract fixes OCR service at 1 worker/1 replica, and the RAG API is the only writer |
|
||||
| 8 | OCR service/image boundary | **Exactly as contracted**: `ocr-service/` in-repo, FastAPI or plain HTTP Python app (implementation detail for design phase), PaddleOCR 3.4.0 + PaddlePaddle 3.2.2 CPU pinned, models baked at build, `/v1/jobs` multipart + Idempotency-Key, `/health/live` + `/health/ready`, internal network only, `OCR_INTERNAL_TOKEN` bearer, SQLite queue on ephemeral `/data/jobs` with 24h TTL, 50 MiB limit, config allowlist. RAG's `src/modules/ocr/client.ts` validates `documentSha256`, page sets, and result schema strictly | Closed decision in contract ("Despliegue OCR", "Contrato interno del servicio OCR") |
|
||||
| 9 | Durable uploads | Replace `multer.memoryStorage()` with disk storage into a temp staging dir, enforce 50 MiB + 100 pages + PDF sanity (reject encrypted/corrupt before full render), and move the original into the version artifact dir once identity is resolved | Contract: "Antes de habilitar OCR, sustituir multer.memoryStorage() por upload a disco con limite y limpieza garantizada" |
|
||||
| 10 | Representative fixtures | **Three committed fixtures** under `tests/fixtures/ocr/`: (1) `native.pdf` — text-rich, all pages pass detection (no OCR called); (2) `scanned.pdf` — image-only pages (2–3 pages, small) requiring OCR on all non-blank pages; (3) `mixed.pdf` — some native pages + some scanned pages, containing tokens shaped like the benchmark's risk identifiers (`CBG04a`-like). Plus unit fixtures: single-page blank image, oversized-rejection corpus. The real FacturaTech PDF stays out of the repo (client data); its 34-entry acceptance run happens in production per the contract's gate. Fixtures must be tiny (<200 KiB total) and generated with a script committed for reproducibility, so tests stay offline | `npm test` has no network; OCR service tests run inside its own Python test suite (pytest) — Node tests fake the OCR client. Contract acceptance criteria name the exact behaviors the fixtures must exercise |
|
||||
|
||||
## Approaches
|
||||
|
||||
1. **Contract-as-specified, in-process dispatcher (recommended)** — implement Punto 3 exactly as the contract defines: async ingest branch with 202, new tables 002, OCR client with retries/polling, artifacts manager, review endpoints, in-process job dispatcher.
|
||||
- Pros: zero contract amendments; all closed decisions honored; single deployable RAG change + one new OCR service; reconciler extension is small; rollback = feature-flag off.
|
||||
- Cons: dispatcher lifetime couples to RAG process (acceptable: single-instance deploy already assumed); ingest service grows substantially.
|
||||
- Effort: High (largest change in RAG history), but decomposable into 7 blocks.
|
||||
|
||||
2. **Separate Node worker process** — dispatcher runs as its own EasyPanel service consuming the same PostgreSQL queue.
|
||||
- Pros: isolates OCR polling latency; RAG API stays minimal.
|
||||
- Cons: violates the current single-service ops model; duplicate pool/lock code; two deploys to coordinate for what is currently ≤3 concurrent jobs; no requirement justifies it.
|
||||
- Effort: High+.
|
||||
|
||||
3. **Minimal sync OCR first, review later** — call OCR inline during ingest with a long timeout, add review workflow in a later change.
|
||||
- Pros: fewer moving parts initially.
|
||||
- Cons: directly violates the closed decision ("responde sin esperar los minutos de OCR", 202 async, review before activation); 245s benchmark runtime would hold HTTP connections; rejected.
|
||||
|
||||
**Recommendation: Approach 1.** It is the only one consistent with the closed contract, and the in-process dispatcher keeps deployment identical in shape to today's (RAG + private OCR service).
|
||||
|
||||
## Seven Work Blocks
|
||||
|
||||
Dependencies flow strictly downward; each block maps to a chained PR slice.
|
||||
|
||||
| # | Block | Contents | Depends on | Key risk |
|
||||
|---|---|---|---|---|
|
||||
| 1 | Persistence & state machine | `migrations/002_ocr_review.sql`; repo methods: create OCRCandidateVersion (null-safe point counts), `markReviewRequired`, `approve→indexing`, `reject`, lease claim/renew methods; state-machine unit tests over fake pool | — | Migration 002 is additive (new tables + `extraction_method` values already exist); zero impact on live 001 data. Rollback: drop 002 tables (safe — no production data references them until OCR ships) |
|
||||
| 2 | Page-level PDF extraction | pdfjs-dist spike + `parsePdfPages`; detection metrics (N/A/W/R thresholds); candidate composition (`--- Page N ---` separators); canonical JSON; detection unit tests with fixtures | 1 (types only) | Library choice must be settled by spike evidence early; if pdfjs-dist legacy build fails on Node 22, fall back to `pdf-parse` custom pagerender |
|
||||
| 3 | PaddleOCR service | `ocr-service/` FastAPI app: jobs (SQLite queue, idempotency, TTL), engine wrapper (PaddleOCR 3.4.0 pinned), render pipeline (200 DPI, 25 MP cap), auth, limits, health; pytest suite with a stub engine | none (parallel-safe; contract fully specifies it) | Image build size/time (Paddle ~2 GB); models must bake at build; readiness false until models load |
|
||||
| 4 | Orchestration & artifacts | `ocr/client.ts` (retries, backoff 2–15s, strict validation); `ocr/artifacts.ts` (0600 writes, manifest, artifactId); in-process dispatcher (SKIP LOCKED, heartbeats); async ingest branch with 202 + `GET /ingestions/:versionId`; durable disk multer upload; reconciler lease-awareness | 1, 2, 3 | Biggest integration surface: touches ingest, app, reconciler; must not regress native sync path (regression tests required) |
|
||||
| 5 | Review, corrections, indexing | Risk-token detection (regex + priority rules); `GET /ingestions/:versionId/review`; approve with `candidateSha256` + per-line `expectedLineSha256` corrections (all-or-nothing 409); duplicate-reusable check; post-approval chunk/embed/index/activate via existing point-2 path; playground review UI | 4 | Correction collision semantics (409 without partial application) must be exact; concurrency with activation changes |
|
||||
| 6 | Security, retention, deployment | Retention sweeper (TTLs, artifact_state CAS); `/health` OCR section; OpenAPI full update (`202`, `/ingestions/*`, review routes, bearer); Dockerfile volume + OCR compose config; docs (API_RAG, INGESTA, deploy); no-secret review | 4, 5 | Deployment ordering: OCR service first, RAG with OCR disabled, then enable — irreversible risk is low if ordering respected |
|
||||
| 7 | End-to-end validation | Fixture-driven e2e (native/scanned/mixed), idempotent-resend test, fail-closed test (OCR down), full static suite (`npm run check`, `npm run build`, Node tests + pytest), production validation plan for the real FacturaTech PDF | 1–6 | Production acceptance (34 entries) requires human review step; cannot be automated |
|
||||
|
||||
Boundary adjustments made (with evidence): the contract's own task order lists "añadir migracion OCR y contratos TypeScript" first and extraction second, but **Block 3 (OCR service) has no code dependency on Blocks 1–2** — the contract fully specifies its HTTP interface. Starting it in parallel is the only way to fit the overall work into review-budget-sized slices without artificial sequencing. This is an execution-order adjustment only; all contract sequencing gates (OCR deployed before RAG enables it) remain.
|
||||
|
||||
## Dependencies & Risk Map
|
||||
|
||||
- **Hard dependency chain:** 1 → 2 → 4 → 5 → 6 → 7; 3 is parallel, joins at 4.
|
||||
- **Irreversible operations:** only `migrations/002_ocr_review.sql` (additive DDL — new tables, no ALTER on existing ones; safe rollback = `DROP TABLE rag_ocr_jobs, rag_document_pages, rag_review_corrections` + `DELETE FROM rag_schema_migrations WHERE name='002_ocr_review.sql'`). Qdrant unchanged until approval-driven indexing, which uses existing versioned point IDs. Production artifact volume starts empty; retention never deletes active versions.
|
||||
- **Rollback boundary:** feature-flag `OCR_INGEST_ENABLED` (new env). Off = today's native-only sync behavior byte-for-byte (regression suite proves it). The 202 async path and `/ingestions/*` routes exist but return `503 OCR_DISABLED`-style errors when off — or routes can be conditionally mounted; design phase decides. Data written while enabled (candidates, jobs) is invisible to retrieval by construction (never `active` without review).
|
||||
- **Deploy ordering (contract-mandated):** OCR service private deploy + readiness check → RAG deploy with OCR disabled → verify native regressions → enable flag → run real PDF validation.
|
||||
- **Review-budget pressure (400 lines, auto-chain):** **High across the whole change.** Realistic estimate: Block 1 ~450–550 lines (SQL+repo+tests), Block 2 ~350–450, Block 3 ~600–800 (Python, reviewed separately from Node PRs), Block 4 ~700–900, Block 5 ~600–750, Block 6 ~350–450, Block 7 ~250–350. Every block except possibly Block 7 individually risks exceeding 400 changed lines; chained/stacked PRs are mandatory, and Blocks 4 and 5 may each need two slices (e.g., 4a client+artifacts, 4b dispatcher+ingest async).
|
||||
|
||||
## Risks
|
||||
|
||||
- **pdfjs-dist integration uncertainty on Node 22** (legacy vs modern build, worker config) — mitigate with a day-one spike in Block 2, contract already allows the `pdf-parse` fallback.
|
||||
- **OCR image build fragility**: PaddlePaddle 3.2.2 CPU wheels are large and platform-sensitive; baked models add build time. Mitigate: pin exact versions, build once, keep ocr-service deploy independent (contract requirement).
|
||||
- **Ingest regression on the native path**: the async branch rewrites the hottest service. Mitigate: strict TDD per config (strict_tdd: true), the 25 existing tests must stay green, plus new native-regression tests before any OCR wiring.
|
||||
- **State-machine drift from the contract**: the contract's diagram is normative; any deviation needs contract amendment FIRST. Mitigation: spec phase encodes each transition as Given/When/Then.
|
||||
- **Playground review UI scope creep**: image rendering, side-by-side diffs, corrections editor could balloon. Mitigation: MVP review UI = native/OCR text + corrections JSON input + image links; defer polish (the playground doc says it's an internal tool).
|
||||
- **VPS2 resource contention**: OCR service capped at 3 CPU/5 GiB of the 6 vCPU/11 GiB host; a 25-page/245s benchmark job can saturate one core for minutes. Mitigate: contract's queue depth 3, single worker, and 60s/page timeout are already sized for this.
|
||||
|
||||
## Ready for Proposal
|
||||
|
||||
**Yes.** Every open decision above is resolved within the closed contract; no human input is required to draft the proposal. The proposal must state: contract-conformant implementation (Approach 1), seven chained blocks as scoped, Block-3 parallelism, feature-flag rollback, and the deploy ordering. Contract amendments are NOT needed — the contract's Punto 3 already anticipates all resolved details; the only judgment calls (worker topology, fixture strategy, library spike order) are implementation-level and recorded here.
|
||||
65
openspec/changes/ocr-ingest-integration/proposal.md
Normal file
65
openspec/changes/ocr-ingest-integration/proposal.md
Normal file
|
|
@ -0,0 +1,65 @@
|
|||
# Proposal: OCR Ingest Integration (`ocr-ingest-integration`)
|
||||
|
||||
**Project:** rag-service — **Phase:** sdd-propose — 2026-09-13
|
||||
**Store:** hybrid (OpenSpec + Engram `sdd/ocr-ingest-integration/proposal`)
|
||||
**Contract:** `docs/CONTRATO_CICLO_VIDA_Y_OCR.md` Punto 3 (closed)
|
||||
|
||||
## Intent
|
||||
|
||||
Auditable OCR ingestion for scanned PDFs: durable originals, async processing behind a review gate, approval before OCR content enters retrieval. Native ingestion and Punto 2 must not regress.
|
||||
|
||||
## Scope
|
||||
|
||||
**In:** async OCR branch (`202 + statusUrl`, `GET /ingestions/:versionId`, version-level review gate); migration `002_ocr_review.sql` with transitions `indexing→review_required→indexing|rejected`; page-level PDF extraction (pdfjs-dist spike; `pdf-parse` fallback), detection; private `ocr-service/` (Python, PaddleOCR pinned, idempotent `/v1/jobs`); dispatcher, OCR client, artifacts, lease-aware reconciler; review endpoints (`candidateSha256`, corrections, 409), sweeper, review UI; `OCR_INGEST_ENABLED` flag, limits, deploy order, docs, fixtures.
|
||||
|
||||
**Out:** client PDFs in repo; UI polish; extra Node workers; contract amendments. FacturaTech acceptance: production.
|
||||
|
||||
## Capabilities
|
||||
|
||||
> `openspec/specs/` is empty — all new.
|
||||
|
||||
- `ocr-ingest-orchestration`: async ingestion, durable originals, job/page state machine, dispatcher, reconciliation
|
||||
- `ocr-processing`: page extraction, detection, composition, service contract
|
||||
- `ocr-review-workflow`: review gate, corrections, approval, retention
|
||||
|
||||
Modified: none.
|
||||
|
||||
## Approach
|
||||
|
||||
Contract-as-specified (exploration Approach 1): in-process Node dispatcher plus private PaddleOCR service. Blocks: 1) persistence/state, 2) page extraction, 3) OCR service (parallel; joins at 4), 4) orchestration/artifacts, 5) review/indexing, 6) security/retention/deploy, 7) e2e validation. Durable originals precede version creation; unapproved candidates stay out of retrieval. TDD.
|
||||
|
||||
## Affected Areas
|
||||
|
||||
| Area | Impact |
|
||||
|------|--------|
|
||||
| `src/modules/ingest/service.ts`, `parsers/parser-registry.ts` | Modified: async OCR branch; per-page `parsePdfPages` |
|
||||
| `src/modules/catalog/{repository,reconciler}.ts` | Modified: new transitions; lease recovery |
|
||||
| `src/app.ts`, `src/api/openapi.ts`, `public/playground/*` | Modified: disk multer; 202; `/ingestions/*`; review UI |
|
||||
| `src/modules/ocr/`, `ocr-service/`, `migrations/002_ocr_review.sql`, `src/config/env.ts`, `Dockerfile`, docs | New |
|
||||
|
||||
## Risks
|
||||
|
||||
| Risk | Likelihood | Mitigation |
|
||||
|------|------------|------------|
|
||||
| Native ingest regression | Medium | TDD; native-regression tests first |
|
||||
| pdfjs-dist on Node 22 | Medium | Day-one spike; contract fallback |
|
||||
| OCR image fragility | Medium | Pinned versions; baked models |
|
||||
| Review budget overrun | High | Auto-chain; split blocks 4–5 |
|
||||
|
||||
## Rollback Plan
|
||||
|
||||
`OCR_INGEST_ENABLED=false` restores native-only behavior. Migration 002 rollback: drop the new tables plus migration row (no production data references them). Qdrant untouched until approval; candidates invisible to retrieval.
|
||||
|
||||
## Dependencies
|
||||
|
||||
Contract Punto 3; Punto 2 lifecycle (in production). Deploy: OCR service → RAG (flag off) → native verify → enable → acceptance.
|
||||
|
||||
## Success Criteria
|
||||
|
||||
- [ ] Native PDFs ingest sync, zero behavior change; tests green
|
||||
- [ ] Scanned PDF: `202` → OCR → `review_required` → approve → active
|
||||
- [ ] Candidate never retrievable pre-approval (fail-closed)
|
||||
- [ ] Expired leases redispatched idempotently; corrections 409 atomic
|
||||
- [ ] Sweeper honors TTLs; active versions never deleted
|
||||
- [ ] `npm run check`, `npm run build`, `npm test`, pytest offline
|
||||
- [ ] FacturaTech: 34 entries human-approved
|
||||
79
openspec/changes/ocr-ingest-integration/tasks.md
Normal file
79
openspec/changes/ocr-ingest-integration/tasks.md
Normal file
|
|
@ -0,0 +1,79 @@
|
|||
# Tasks: OCR Ingest Integration
|
||||
|
||||
## Forecast
|
||||
|
||||
Estimate: 3,400–4,300 lines.
|
||||
|
||||
Decision needed before apply: No
|
||||
Chained PRs recommended: Yes
|
||||
Chain strategy: feature-branch-chain
|
||||
400-line budget risk: High
|
||||
|
||||
Tracker feature/ocr-ingest-integration is draft/no-merge and sole main target. PR1 targets tracker; children target predecessors. Coverage: 17 requirements/34 scenarios (O10/P12/V12).
|
||||
|
||||
### Units: command; harness; rollback
|
||||
|
||||
1 Migration: tsx --test tests/catalog/migration-002.test.ts; N/A schema-only; drop OCR schema
|
||||
2 Repository: tsx --test tests/catalog/repository-ocr.test.ts; fake pool; revert candidate/leases
|
||||
3 Extraction: tsx --test tests/parsers/pdf-pages.test.ts; PDF spike; revert parser/fixtures
|
||||
4 Detection: tsx --test tests/ocr/detection.test.ts; mixed PDF; revert transforms
|
||||
5 Service: pytest ocr-service/tests -k auth; curl 401/409; remove ocr-service
|
||||
6 Render: pytest ocr-service/tests -k render; Docker PNG; revert wrapper/image
|
||||
7 Client: tsx --test tests/ocr/client.test.ts; retry stub; revert client/artifacts
|
||||
8 Dispatcher: tsx --test tests/ocr/dispatcher.test.ts; curl 202/status; revert routing
|
||||
9 Review: tsx --test tests/ocr/review.test.ts; curl review/409; revert review/routes
|
||||
10 Approval: tsx --test tests/ocr/approval.test.ts; playground approval; revert indexing/UI
|
||||
11 Retention: tsx --test tests/ocr/retention.test.ts; fake-fs sweep; revert retention
|
||||
12 Config: npm run check; N/A static-only; revert config/docs
|
||||
13 E2E: npm test; production OCR flow; disable OCR flag
|
||||
|
||||
## 1 Persistence (Units 1–2; O2/O3)
|
||||
|
||||
- [x] 1.1 RED: duplicate pending, lease recovery, review transitions
|
||||
- [x] 1.2 GREEN: `migrations/002_ocr_review.sql` schema and unique index
|
||||
- [x] 1.3 GREEN: `src/modules/catalog/repository.ts` candidates and leases
|
||||
- [x] 1.4 REFACTOR after green
|
||||
|
||||
## 2 Extraction (Units 3–4; P1–P4)
|
||||
|
||||
- [x] 2.1 Spike `scripts/spike-pdfjs.ts` for Node 22 parser
|
||||
- [x] 2.2 RED: PDF-only routing; text files never execute; unsupported build/script formats; pages, blanks, hashes, risks
|
||||
- [x] 2.3 GREEN: `tests/fixtures/ocr/` and `src/modules/parsers/parser-registry.ts`
|
||||
- [x] 2.4 GREEN: `src/modules/ocr/detection.ts` thresholds
|
||||
- [x] 2.5 GREEN: `src/modules/ocr/composition.ts` text, hashes, risks
|
||||
|
||||
## 3 OCR Service (Units 5–6; P5/P6)
|
||||
|
||||
- [ ] 3.1 RED HTTP: bearer 401; idempotency/conflict 409; allowlist; limits 413/422; pressure 429; no-retry; integrity mismatch
|
||||
- [ ] 3.2 GREEN: `ocr-service/` API, queue, allowlist, limits, health
|
||||
- [ ] 3.3 GREEN: PaddleOCR 3.4.0, 200 DPI, schema, stub
|
||||
- [ ] 3.4 GREEN: `ocr-service/Dockerfile` pinned wheels/models/resources
|
||||
|
||||
## 4 Orchestration (Units 7–8; O1–O4)
|
||||
|
||||
- [ ] 4.1 RED HTTP: flag-off 201; OCR 202/status; OCR-down fail-closed; catalog-down 503; retries/integrity
|
||||
- [ ] 4.2 GREEN: `src/modules/ocr/client.ts` retries and integrity
|
||||
- [ ] 4.3 GREEN: `src/modules/ocr/artifacts.ts` originals, manifest, sweep
|
||||
- [ ] 4.4 GREEN: `src/modules/ocr/dispatcher.ts` leases; `src/modules/catalog/reconciler.ts`
|
||||
- [ ] 4.5 GREEN: `src/modules/ingest/service.ts` branch; `src/app.ts` upload/status
|
||||
|
||||
## 5 Review (Units 9–10; V1–V4)
|
||||
|
||||
- [ ] 5.1 RED: auth/premature access; atomic corrections/stale 409; new/reusable activation race; reject/no embeddings
|
||||
- [ ] 5.2 GREEN: `src/modules/ocr/review.ts` view, corrections, conflicts
|
||||
- [ ] 5.3 GREEN: `src/modules/ocr/indexing.ts` CAS; new/reusable activate_requested=true activates, false stays ready
|
||||
- [ ] 5.4 GREEN: rejection and `public/playground/` review UI
|
||||
|
||||
## 6 Retention/Deploy (Units 11–12; O5/V5)
|
||||
|
||||
- [ ] 6.1 RED: TTL/CAS/resume/active protection; flag-off native-only and candidate invisibility
|
||||
- [ ] 6.2 GREEN: `src/modules/ocr/retention.ts`
|
||||
- [ ] 6.3 GREEN: `src/api/openapi.ts` OCR/status/review contracts
|
||||
- [ ] 6.4 GREEN: `src/config/env.ts`, `Dockerfile`; follow `docs/CONTRATO_CICLO_VIDA_Y_OCR.md` (read-only)
|
||||
|
||||
## 7 E2E (Unit 13; V6)
|
||||
|
||||
- [ ] 7.1 E2E: native 201; scanned 202→review→approve→active; mixed
|
||||
- [ ] 7.2 E2E: resend, OCR-down fail-closed, catalog-down 503
|
||||
- [ ] 7.3 Gate: check, build, test, offline pytest
|
||||
- [ ] 7.4 Production: flag-off verify→enable→FacturaTech 34 entries; CBG04a/FAT07/DSAU08/NSAV06
|
||||
59
openspec/config.yaml
Normal file
59
openspec/config.yaml
Normal file
|
|
@ -0,0 +1,59 @@
|
|||
schema: spec-driven
|
||||
|
||||
context: |
|
||||
Project: rag-service (RAG knowledge service for enterprise AI tooling)
|
||||
Tech stack: Node.js 22 (ESM), TypeScript 5.8 strict, Express 4, tsx runner
|
||||
Architecture: modular service — src/modules/{ingest,process,parsers,retrieve,answer,catalog,embeddings,vectorstore,logs}, src/api (OpenAPI), src/shared
|
||||
Data stores: PostgreSQL (pg) + Qdrant (@qdrant/js-client-rest) + OpenAI embeddings
|
||||
Testing: node:test + node:assert/strict executed via `npm test` (NODE_ENV=test tsx --test tests/**/*.test.ts)
|
||||
Quality: `npm run check` (tsc --noEmit, strict), `npm run build` (tsc emit to dist/)
|
||||
Deploy: Docker + EasyPanel; migrations are manual scripts (migrate:lifecycle, migrate:legacy:lifecycle)
|
||||
No linter, formatter, or coverage tooling is configured.
|
||||
|
||||
strict_tdd: true
|
||||
|
||||
testing:
|
||||
projects:
|
||||
- path: ./
|
||||
stack: Node.js 22 + TypeScript 5.8 (ESM)
|
||||
test_command: npm test
|
||||
test_framework: node:test (node --test via tsx)
|
||||
layers:
|
||||
unit: true
|
||||
integration: true
|
||||
e2e: false
|
||||
coverage:
|
||||
available: false
|
||||
command: ""
|
||||
quality:
|
||||
linter: false
|
||||
type_checker: true
|
||||
type_checker_command: npm run check
|
||||
formatter: false
|
||||
|
||||
rules:
|
||||
proposal:
|
||||
- Include rollback plan for risky changes (migrations and Qdrant/PostgreSQL data operations)
|
||||
- Reference existing module boundaries under src/modules/
|
||||
specs:
|
||||
- Use Given/When/Then for scenarios
|
||||
- Use RFC 2119 keywords (MUST, SHALL, SHOULD, MAY)
|
||||
design:
|
||||
- Document architecture decisions with rationale
|
||||
- Keep module boundaries; parsers/ingest extensions must follow the parser-registry pattern
|
||||
tasks:
|
||||
- Group by phase, use hierarchical numbering
|
||||
- Keep tasks completable in one session
|
||||
- Respect review_budget_lines: 400 (auto-chain delivery when exceeded)
|
||||
apply:
|
||||
guidelines:
|
||||
- Follow existing code patterns (ESM imports with .js extensions, strict types)
|
||||
- Never read or expose .env* files, llaves, or backups/ contents
|
||||
tdd: true
|
||||
test_command: npm test
|
||||
verify:
|
||||
test_command: npm test
|
||||
build_command: npm run build
|
||||
coverage_threshold: 0
|
||||
archive:
|
||||
- Warn before merging destructive deltas
|
||||
0
openspec/specs/.gitkeep
Normal file
0
openspec/specs/.gitkeep
Normal file
Loading…
Add table
Reference in a new issue