diff --git a/openspec/changes/archive/.gitkeep b/openspec/changes/archive/.gitkeep new file mode 100644 index 0000000..e69de29 diff --git a/openspec/changes/ocr-ingest-integration/exploration.md b/openspec/changes/ocr-ingest-integration/exploration.md new file mode 100644 index 0000000..79283a5 --- /dev/null +++ b/openspec/changes/ocr-ingest-integration/exploration.md @@ -0,0 +1,106 @@ +# Exploration: OCR ingest integration (`ocr-ingest-integration`) + +**Project:** rag-service — **Phase:** sdd-explore — **Date:** 2026-09-13 +**Store:** hybrid (OpenSpec file + Engram topic `sdd/ocr-ingest-integration/explore`) + +## Current State + +Point 2 (knowledge lifecycle) is complete, enforced in production (`KNOWLEDGE_LIFECYCLE_ENFORCED=true`, 7 cataloged sources, 22,605 active points, legacy batch migrated and verified 2026-09-13). `npm test` passes 25/25. The contract in `docs/CONTRATO_CICLO_VIDA_Y_OCR.md` (section "Punto 3") is closed and authoritative: every decision below defers to it; this exploration only resolves what the contract leaves open. + +Verified implementation facts (read from code, not docs): + +| Area | Current behavior | OCR gap | +|---|---|---| +| Ingest | `IngestService.ingestWithLifecycle` is fully synchronous: attempt → source lock → discover → hash → reusable lookup → `createPendingVersion` → `markIndexing` → embed → upsert → count → `markReady` → `activateVersion`, all inside the HTTP request | Must become async-capable: durable original first, `202 + statusUrl` response when OCR is needed, `source_content_hash = null` until review completes | +| PDF parsing | `parser-registry.parseDocument` uses `pdf-parse` over the whole file, returns aggregated text only; no page callback in our code (library has `pagerender` option but we never use it; bundled pdf.js is v1.10.100, very old) | Page-level extraction required; contract allows replacing `pdf-parse` if unreliable | +| State machine | `SourceVersionState` already includes `review_required` and `rejected`; transitions `pending→indexing→ready→active`, `failed`, `purging→purged` implemented in `CatalogRepository`; **no transition writes into `review_required` today** | Need: `indexing→review_required` (OCR candidate complete), `review_required→indexing` (approve), `review_required→rejected`, plus `review_required`/`rejected` entries in `markPurging`-eligible states (already present), retention expiry transitions | +| Catalog repo | `createPendingVersion` requires non-null `expectedPointCount` and writes all documents with `content_hash` pre-computed; `findReusableVersion` skips when `sourceContentHash` is null (correct for OCR candidates); `hasCompleteVersionDocuments` counts `content_hash IS NULL` as incomplete | OCR flow must create a pending version **before** content exists (chunk counts unknown until review) — needs a creation mode with `expectedPointCount=0`/null-safe path, or a separate "OCR candidate" insert path | +| Reconciler | `runOnce` treats stale `pending/indexing` older than 30 min as orphans and either recovers to `ready` or fails them | Must not kill in-flight OCR versions: needs heartbeat awareness (jobs have `heartbeat_at`/`lease_expires_at` per contract) and must recover expired OCR leases by re-dispatching with the same remote idempotency key | +| Uploads | `multer.memoryStorage()` in `src/app.ts`, no size limit; upload writes to `/tmp` with timestamp, deleted in `finally` — no durable original | Contract requires disk-based multer with limits and durable original under `/data/ingestions//` before processing | +| Env/config | `env.ts` has PostgreSQL + lifecycle tokens only | Needs `OCR_SERVICE_URL`, `OCR_INTERNAL_TOKEN`, artifact volume path, limits (50 MiB, 100 pages, timeouts), retention intervals | +| Migrations | `migrations/001_knowledge_lifecycle.sql` applied via `rag_schema_migrations` with checksums + advisory lock; migration runner sorts `^\d+_.*\.sql$` | `002_ocr_review.sql` fits the existing runner unchanged (tables `rag_ocr_jobs`, `rag_document_pages`, `rag_review_corrections` per contract) | +| OpenAPI | `IngestResponse` is 201-only under lifecycle; `securitySchemes.bearerAuth` already defined and used by lifecycle endpoints | Needs `202` variant with `phase`/`statusUrl`, `/ingestions/:versionId` status route, review/approve/reject routes | +| Playground | `public/playground/app.js` supports activate/expectedActiveVersionId; no review UI | Review view (page images, native vs OCR text, corrections) required by acceptance criteria | +| Docker | Node runtime only; `CMD ["node", "dist/server.js"]` | Needs `/data/ingestions` volume; ocr-service gets its own Dockerfile (Python CPU, PaddleOCR 3.4.0 + PaddlePaddle 3.2.2 pinned, baked models) | +| Artifacts | No filesystem artifact manager exists; `rag_version_documents.artifact_manifest_path` and `artifact_state`/`retention_due_at` columns exist but are never written (`retention_deleted` only set on purge) | Whole artifact subsystem (write/verify 0600, manifest, retention sweeper) is new | + +## Affected Areas + +- `src/modules/ingest/service.ts` — async OCR branch, durable originals, version creation before content +- `src/modules/parsers/parser-registry.ts` — per-page PDF extraction API +- `src/modules/catalog/repository.ts` — new state transitions, OCR-safe version creation, page/job/correction persistence +- `src/modules/catalog/reconciler.ts` — lease-aware orphan handling, OCR job re-dispatch +- `src/app.ts` — disk multer, 202 async responses, `/ingestions/*` + review endpoints +- `src/api/openapi.ts`, `public/playground/*` — new contracts and review UI +- `src/config/env.ts`, `migrations/002_ocr_review.sql`, `Dockerfile`, `docs/API_RAG.md`, `docs/INGESTA.md` +- New: `src/modules/ocr/{client,detection,artifacts,review}.ts`, `ocr-service/` (Python) + +## Open Decisions Resolved + +| # | Decision | Resolution | Rationale | +|---|---|---|---| +| 1 | Page-level PDF extraction | **Replace `pdf-parse` with `pdfjs-dist` (legacy build) used directly**, wrapped in a new `parsePdfPages(buffer)` that returns `Array<{pageNumber, text}>`. Keep `pdf-parse` only if a spike proves its `pagerender` callback reliable on Node 22 — the contract permits either, but pdf.js v1.10.100 (2018) is too old to trust for the metrics contract. Fallback order: pdfjs-dist direct → `pdf-parse` with custom pagerender. Must be settled by a fixture-backed spike in Block 2 before anything downstream builds on it | Contract mandates "callback per pagina probado con fixture" and allows substitution. pdfjs-dist is the same engine, maintained, permissively licensed (Apache-2.0), and gives page objects with text items we need for metrics | +| 2 | Job/page state machine & progress | **Reuse the contract schema exactly**: `rag_ocr_jobs` (one per version+document, `UNIQUE(version_id, document_id)`), `rag_document_pages` per page, dispatcher uses `FOR UPDATE SKIP LOCKED`, heartbeats. Progress = `completed_pages/totalPages` surfaced via `GET /ingestions/:versionId` from a single aggregate query over jobs+pages. No new states beyond contract | Schema is already specified in the contract (Punto 3 "Persistencia OCR"); inventing alternatives violates the closed decision | +| 3 | Multi-document semantics | **Version-level candidate gate**: a version reaches `review_required` only when ALL documents' pages are native-sufficient or OCR-complete (contract quality gate is per-document AND per-version). One failed document fails the version (`OCR_QUALITY_BLOCKED` or the document's `error_code`). Mixed sources (md + pdf) in one folder ingest: only PDFs go through OCR; non-PDF docs keep native content_hash immediately, PDF docs get `content_hash` after review | The contract's gate ("el documento solo llega a review_required cuando…") is per-document; activation is version-level, so the join is forced | +| 4 | Reconciler behavior | Extend `recoverOrphanedVersions` to distinguish: (a) versions in `indexing` with live OCR jobs (`lease_expires_at > now`) → leave alone; (b) expired leases → re-enqueue with same `remote_idempotency_key`, reset `next_attempt_at`; (c) stale native versions → existing behavior. The dispatcher itself (new module) owns lease renewal; reconciler only recovers dead leases | Contract: "Un reconciliador periodico reenvia con la misma clave remota y recupera remote_job_id sin duplicar OCR" | +| 5 | Durable originals/artifacts | **Write original + manifest under `/data/ingestions//documents//` BEFORE `createPendingVersion`** (contract sequence: "se guarda el original durable; se crea la version…"). `documentArtifactId` = UUIDv5(document_id). If ingest dies between artifact write and version creation, a startup/periodic sweep deletes orphan version dirs (no catalog row = not referenceable). `artifact_state` column transitions: `none → present` at write, `present → retention_deleting → retention_deleted` by sweeper | Ordering is mandated by the contract's async section; sweep covers the one crash window | +| 6 | Purge/retention | **Single daily sweeper** per contract: acquires the same version CAS/lock used by review and purge (`withVersionExclusiveLock`), sets `artifact_state='retention_deleting'`, deletes idempotently, marks `retention_deleted`. State TTLs: review_required 30d → `rejected`/`REVIEW_EXPIRED`; failed/rejected 7d → `purging→purged`; review images 7d post-decision; active keeps original+reviewed while active +30d. An active version is never TTL-deleted | All values are fixed by the contract's "Artefactos y retencion" section | +| 7 | Worker topology | **In-process dispatcher inside the RAG Node service** (single instance, single EasyPanel app): a `setInterval`-driven poller that claims jobs with `FOR UPDATE SKIP LOCKED`, heartbeats, polls the OCR service, never blocks the ingest HTTP path beyond the initial enqueue. No separate Node worker deploy. The OCR service itself is one worker + one replica per contract | Adding a second Node worker would double deploy surface for zero benefit at current scale (≤3 queued jobs); contract fixes OCR service at 1 worker/1 replica, and the RAG API is the only writer | +| 8 | OCR service/image boundary | **Exactly as contracted**: `ocr-service/` in-repo, FastAPI or plain HTTP Python app (implementation detail for design phase), PaddleOCR 3.4.0 + PaddlePaddle 3.2.2 CPU pinned, models baked at build, `/v1/jobs` multipart + Idempotency-Key, `/health/live` + `/health/ready`, internal network only, `OCR_INTERNAL_TOKEN` bearer, SQLite queue on ephemeral `/data/jobs` with 24h TTL, 50 MiB limit, config allowlist. RAG's `src/modules/ocr/client.ts` validates `documentSha256`, page sets, and result schema strictly | Closed decision in contract ("Despliegue OCR", "Contrato interno del servicio OCR") | +| 9 | Durable uploads | Replace `multer.memoryStorage()` with disk storage into a temp staging dir, enforce 50 MiB + 100 pages + PDF sanity (reject encrypted/corrupt before full render), and move the original into the version artifact dir once identity is resolved | Contract: "Antes de habilitar OCR, sustituir multer.memoryStorage() por upload a disco con limite y limpieza garantizada" | +| 10 | Representative fixtures | **Three committed fixtures** under `tests/fixtures/ocr/`: (1) `native.pdf` — text-rich, all pages pass detection (no OCR called); (2) `scanned.pdf` — image-only pages (2–3 pages, small) requiring OCR on all non-blank pages; (3) `mixed.pdf` — some native pages + some scanned pages, containing tokens shaped like the benchmark's risk identifiers (`CBG04a`-like). Plus unit fixtures: single-page blank image, oversized-rejection corpus. The real FacturaTech PDF stays out of the repo (client data); its 34-entry acceptance run happens in production per the contract's gate. Fixtures must be tiny (<200 KiB total) and generated with a script committed for reproducibility, so tests stay offline | `npm test` has no network; OCR service tests run inside its own Python test suite (pytest) — Node tests fake the OCR client. Contract acceptance criteria name the exact behaviors the fixtures must exercise | + +## Approaches + +1. **Contract-as-specified, in-process dispatcher (recommended)** — implement Punto 3 exactly as the contract defines: async ingest branch with 202, new tables 002, OCR client with retries/polling, artifacts manager, review endpoints, in-process job dispatcher. + - Pros: zero contract amendments; all closed decisions honored; single deployable RAG change + one new OCR service; reconciler extension is small; rollback = feature-flag off. + - Cons: dispatcher lifetime couples to RAG process (acceptable: single-instance deploy already assumed); ingest service grows substantially. + - Effort: High (largest change in RAG history), but decomposable into 7 blocks. + +2. **Separate Node worker process** — dispatcher runs as its own EasyPanel service consuming the same PostgreSQL queue. + - Pros: isolates OCR polling latency; RAG API stays minimal. + - Cons: violates the current single-service ops model; duplicate pool/lock code; two deploys to coordinate for what is currently ≤3 concurrent jobs; no requirement justifies it. + - Effort: High+. + +3. **Minimal sync OCR first, review later** — call OCR inline during ingest with a long timeout, add review workflow in a later change. + - Pros: fewer moving parts initially. + - Cons: directly violates the closed decision ("responde sin esperar los minutos de OCR", 202 async, review before activation); 245s benchmark runtime would hold HTTP connections; rejected. + +**Recommendation: Approach 1.** It is the only one consistent with the closed contract, and the in-process dispatcher keeps deployment identical in shape to today's (RAG + private OCR service). + +## Seven Work Blocks + +Dependencies flow strictly downward; each block maps to a chained PR slice. + +| # | Block | Contents | Depends on | Key risk | +|---|---|---|---|---| +| 1 | Persistence & state machine | `migrations/002_ocr_review.sql`; repo methods: create OCRCandidateVersion (null-safe point counts), `markReviewRequired`, `approve→indexing`, `reject`, lease claim/renew methods; state-machine unit tests over fake pool | — | Migration 002 is additive (new tables + `extraction_method` values already exist); zero impact on live 001 data. Rollback: drop 002 tables (safe — no production data references them until OCR ships) | +| 2 | Page-level PDF extraction | pdfjs-dist spike + `parsePdfPages`; detection metrics (N/A/W/R thresholds); candidate composition (`--- Page N ---` separators); canonical JSON; detection unit tests with fixtures | 1 (types only) | Library choice must be settled by spike evidence early; if pdfjs-dist legacy build fails on Node 22, fall back to `pdf-parse` custom pagerender | +| 3 | PaddleOCR service | `ocr-service/` FastAPI app: jobs (SQLite queue, idempotency, TTL), engine wrapper (PaddleOCR 3.4.0 pinned), render pipeline (200 DPI, 25 MP cap), auth, limits, health; pytest suite with a stub engine | none (parallel-safe; contract fully specifies it) | Image build size/time (Paddle ~2 GB); models must bake at build; readiness false until models load | +| 4 | Orchestration & artifacts | `ocr/client.ts` (retries, backoff 2–15s, strict validation); `ocr/artifacts.ts` (0600 writes, manifest, artifactId); in-process dispatcher (SKIP LOCKED, heartbeats); async ingest branch with 202 + `GET /ingestions/:versionId`; durable disk multer upload; reconciler lease-awareness | 1, 2, 3 | Biggest integration surface: touches ingest, app, reconciler; must not regress native sync path (regression tests required) | +| 5 | Review, corrections, indexing | Risk-token detection (regex + priority rules); `GET /ingestions/:versionId/review`; approve with `candidateSha256` + per-line `expectedLineSha256` corrections (all-or-nothing 409); duplicate-reusable check; post-approval chunk/embed/index/activate via existing point-2 path; playground review UI | 4 | Correction collision semantics (409 without partial application) must be exact; concurrency with activation changes | +| 6 | Security, retention, deployment | Retention sweeper (TTLs, artifact_state CAS); `/health` OCR section; OpenAPI full update (`202`, `/ingestions/*`, review routes, bearer); Dockerfile volume + OCR compose config; docs (API_RAG, INGESTA, deploy); no-secret review | 4, 5 | Deployment ordering: OCR service first, RAG with OCR disabled, then enable — irreversible risk is low if ordering respected | +| 7 | End-to-end validation | Fixture-driven e2e (native/scanned/mixed), idempotent-resend test, fail-closed test (OCR down), full static suite (`npm run check`, `npm run build`, Node tests + pytest), production validation plan for the real FacturaTech PDF | 1–6 | Production acceptance (34 entries) requires human review step; cannot be automated | + +Boundary adjustments made (with evidence): the contract's own task order lists "añadir migracion OCR y contratos TypeScript" first and extraction second, but **Block 3 (OCR service) has no code dependency on Blocks 1–2** — the contract fully specifies its HTTP interface. Starting it in parallel is the only way to fit the overall work into review-budget-sized slices without artificial sequencing. This is an execution-order adjustment only; all contract sequencing gates (OCR deployed before RAG enables it) remain. + +## Dependencies & Risk Map + +- **Hard dependency chain:** 1 → 2 → 4 → 5 → 6 → 7; 3 is parallel, joins at 4. +- **Irreversible operations:** only `migrations/002_ocr_review.sql` (additive DDL — new tables, no ALTER on existing ones; safe rollback = `DROP TABLE rag_ocr_jobs, rag_document_pages, rag_review_corrections` + `DELETE FROM rag_schema_migrations WHERE name='002_ocr_review.sql'`). Qdrant unchanged until approval-driven indexing, which uses existing versioned point IDs. Production artifact volume starts empty; retention never deletes active versions. +- **Rollback boundary:** feature-flag `OCR_INGEST_ENABLED` (new env). Off = today's native-only sync behavior byte-for-byte (regression suite proves it). The 202 async path and `/ingestions/*` routes exist but return `503 OCR_DISABLED`-style errors when off — or routes can be conditionally mounted; design phase decides. Data written while enabled (candidates, jobs) is invisible to retrieval by construction (never `active` without review). +- **Deploy ordering (contract-mandated):** OCR service private deploy + readiness check → RAG deploy with OCR disabled → verify native regressions → enable flag → run real PDF validation. +- **Review-budget pressure (400 lines, auto-chain):** **High across the whole change.** Realistic estimate: Block 1 ~450–550 lines (SQL+repo+tests), Block 2 ~350–450, Block 3 ~600–800 (Python, reviewed separately from Node PRs), Block 4 ~700–900, Block 5 ~600–750, Block 6 ~350–450, Block 7 ~250–350. Every block except possibly Block 7 individually risks exceeding 400 changed lines; chained/stacked PRs are mandatory, and Blocks 4 and 5 may each need two slices (e.g., 4a client+artifacts, 4b dispatcher+ingest async). + +## Risks + +- **pdfjs-dist integration uncertainty on Node 22** (legacy vs modern build, worker config) — mitigate with a day-one spike in Block 2, contract already allows the `pdf-parse` fallback. +- **OCR image build fragility**: PaddlePaddle 3.2.2 CPU wheels are large and platform-sensitive; baked models add build time. Mitigate: pin exact versions, build once, keep ocr-service deploy independent (contract requirement). +- **Ingest regression on the native path**: the async branch rewrites the hottest service. Mitigate: strict TDD per config (strict_tdd: true), the 25 existing tests must stay green, plus new native-regression tests before any OCR wiring. +- **State-machine drift from the contract**: the contract's diagram is normative; any deviation needs contract amendment FIRST. Mitigation: spec phase encodes each transition as Given/When/Then. +- **Playground review UI scope creep**: image rendering, side-by-side diffs, corrections editor could balloon. Mitigation: MVP review UI = native/OCR text + corrections JSON input + image links; defer polish (the playground doc says it's an internal tool). +- **VPS2 resource contention**: OCR service capped at 3 CPU/5 GiB of the 6 vCPU/11 GiB host; a 25-page/245s benchmark job can saturate one core for minutes. Mitigate: contract's queue depth 3, single worker, and 60s/page timeout are already sized for this. + +## Ready for Proposal + +**Yes.** Every open decision above is resolved within the closed contract; no human input is required to draft the proposal. The proposal must state: contract-conformant implementation (Approach 1), seven chained blocks as scoped, Block-3 parallelism, feature-flag rollback, and the deploy ordering. Contract amendments are NOT needed — the contract's Punto 3 already anticipates all resolved details; the only judgment calls (worker topology, fixture strategy, library spike order) are implementation-level and recorded here. \ No newline at end of file diff --git a/openspec/changes/ocr-ingest-integration/proposal.md b/openspec/changes/ocr-ingest-integration/proposal.md new file mode 100644 index 0000000..eb48832 --- /dev/null +++ b/openspec/changes/ocr-ingest-integration/proposal.md @@ -0,0 +1,65 @@ +# Proposal: OCR Ingest Integration (`ocr-ingest-integration`) + +**Project:** rag-service — **Phase:** sdd-propose — 2026-09-13 +**Store:** hybrid (OpenSpec + Engram `sdd/ocr-ingest-integration/proposal`) +**Contract:** `docs/CONTRATO_CICLO_VIDA_Y_OCR.md` Punto 3 (closed) + +## Intent + +Auditable OCR ingestion for scanned PDFs: durable originals, async processing behind a review gate, approval before OCR content enters retrieval. Native ingestion and Punto 2 must not regress. + +## Scope + +**In:** async OCR branch (`202 + statusUrl`, `GET /ingestions/:versionId`, version-level review gate); migration `002_ocr_review.sql` with transitions `indexing→review_required→indexing|rejected`; page-level PDF extraction (pdfjs-dist spike; `pdf-parse` fallback), detection; private `ocr-service/` (Python, PaddleOCR pinned, idempotent `/v1/jobs`); dispatcher, OCR client, artifacts, lease-aware reconciler; review endpoints (`candidateSha256`, corrections, 409), sweeper, review UI; `OCR_INGEST_ENABLED` flag, limits, deploy order, docs, fixtures. + +**Out:** client PDFs in repo; UI polish; extra Node workers; contract amendments. FacturaTech acceptance: production. + +## Capabilities + +> `openspec/specs/` is empty — all new. + +- `ocr-ingest-orchestration`: async ingestion, durable originals, job/page state machine, dispatcher, reconciliation +- `ocr-processing`: page extraction, detection, composition, service contract +- `ocr-review-workflow`: review gate, corrections, approval, retention + +Modified: none. + +## Approach + +Contract-as-specified (exploration Approach 1): in-process Node dispatcher plus private PaddleOCR service. Blocks: 1) persistence/state, 2) page extraction, 3) OCR service (parallel; joins at 4), 4) orchestration/artifacts, 5) review/indexing, 6) security/retention/deploy, 7) e2e validation. Durable originals precede version creation; unapproved candidates stay out of retrieval. TDD. + +## Affected Areas + +| Area | Impact | +|------|--------| +| `src/modules/ingest/service.ts`, `parsers/parser-registry.ts` | Modified: async OCR branch; per-page `parsePdfPages` | +| `src/modules/catalog/{repository,reconciler}.ts` | Modified: new transitions; lease recovery | +| `src/app.ts`, `src/api/openapi.ts`, `public/playground/*` | Modified: disk multer; 202; `/ingestions/*`; review UI | +| `src/modules/ocr/`, `ocr-service/`, `migrations/002_ocr_review.sql`, `src/config/env.ts`, `Dockerfile`, docs | New | + +## Risks + +| Risk | Likelihood | Mitigation | +|------|------------|------------| +| Native ingest regression | Medium | TDD; native-regression tests first | +| pdfjs-dist on Node 22 | Medium | Day-one spike; contract fallback | +| OCR image fragility | Medium | Pinned versions; baked models | +| Review budget overrun | High | Auto-chain; split blocks 4–5 | + +## Rollback Plan + +`OCR_INGEST_ENABLED=false` restores native-only behavior. Migration 002 rollback: drop the new tables plus migration row (no production data references them). Qdrant untouched until approval; candidates invisible to retrieval. + +## Dependencies + +Contract Punto 3; Punto 2 lifecycle (in production). Deploy: OCR service → RAG (flag off) → native verify → enable → acceptance. + +## Success Criteria + +- [ ] Native PDFs ingest sync, zero behavior change; tests green +- [ ] Scanned PDF: `202` → OCR → `review_required` → approve → active +- [ ] Candidate never retrievable pre-approval (fail-closed) +- [ ] Expired leases redispatched idempotently; corrections 409 atomic +- [ ] Sweeper honors TTLs; active versions never deleted +- [ ] `npm run check`, `npm run build`, `npm test`, pytest offline +- [ ] FacturaTech: 34 entries human-approved \ No newline at end of file diff --git a/openspec/changes/ocr-ingest-integration/tasks.md b/openspec/changes/ocr-ingest-integration/tasks.md new file mode 100644 index 0000000..004c93e --- /dev/null +++ b/openspec/changes/ocr-ingest-integration/tasks.md @@ -0,0 +1,79 @@ +# Tasks: OCR Ingest Integration + +## Forecast + +Estimate: 3,400–4,300 lines. + +Decision needed before apply: No +Chained PRs recommended: Yes +Chain strategy: feature-branch-chain +400-line budget risk: High + +Tracker feature/ocr-ingest-integration is draft/no-merge and sole main target. PR1 targets tracker; children target predecessors. Coverage: 17 requirements/34 scenarios (O10/P12/V12). + +### Units: command; harness; rollback + +1 Migration: tsx --test tests/catalog/migration-002.test.ts; N/A schema-only; drop OCR schema +2 Repository: tsx --test tests/catalog/repository-ocr.test.ts; fake pool; revert candidate/leases +3 Extraction: tsx --test tests/parsers/pdf-pages.test.ts; PDF spike; revert parser/fixtures +4 Detection: tsx --test tests/ocr/detection.test.ts; mixed PDF; revert transforms +5 Service: pytest ocr-service/tests -k auth; curl 401/409; remove ocr-service +6 Render: pytest ocr-service/tests -k render; Docker PNG; revert wrapper/image +7 Client: tsx --test tests/ocr/client.test.ts; retry stub; revert client/artifacts +8 Dispatcher: tsx --test tests/ocr/dispatcher.test.ts; curl 202/status; revert routing +9 Review: tsx --test tests/ocr/review.test.ts; curl review/409; revert review/routes +10 Approval: tsx --test tests/ocr/approval.test.ts; playground approval; revert indexing/UI +11 Retention: tsx --test tests/ocr/retention.test.ts; fake-fs sweep; revert retention +12 Config: npm run check; N/A static-only; revert config/docs +13 E2E: npm test; production OCR flow; disable OCR flag + +## 1 Persistence (Units 1–2; O2/O3) + +- [x] 1.1 RED: duplicate pending, lease recovery, review transitions +- [x] 1.2 GREEN: `migrations/002_ocr_review.sql` schema and unique index +- [x] 1.3 GREEN: `src/modules/catalog/repository.ts` candidates and leases +- [x] 1.4 REFACTOR after green + +## 2 Extraction (Units 3–4; P1–P4) + +- [x] 2.1 Spike `scripts/spike-pdfjs.ts` for Node 22 parser +- [x] 2.2 RED: PDF-only routing; text files never execute; unsupported build/script formats; pages, blanks, hashes, risks +- [x] 2.3 GREEN: `tests/fixtures/ocr/` and `src/modules/parsers/parser-registry.ts` +- [x] 2.4 GREEN: `src/modules/ocr/detection.ts` thresholds +- [x] 2.5 GREEN: `src/modules/ocr/composition.ts` text, hashes, risks + +## 3 OCR Service (Units 5–6; P5/P6) + +- [ ] 3.1 RED HTTP: bearer 401; idempotency/conflict 409; allowlist; limits 413/422; pressure 429; no-retry; integrity mismatch +- [ ] 3.2 GREEN: `ocr-service/` API, queue, allowlist, limits, health +- [ ] 3.3 GREEN: PaddleOCR 3.4.0, 200 DPI, schema, stub +- [ ] 3.4 GREEN: `ocr-service/Dockerfile` pinned wheels/models/resources + +## 4 Orchestration (Units 7–8; O1–O4) + +- [ ] 4.1 RED HTTP: flag-off 201; OCR 202/status; OCR-down fail-closed; catalog-down 503; retries/integrity +- [ ] 4.2 GREEN: `src/modules/ocr/client.ts` retries and integrity +- [ ] 4.3 GREEN: `src/modules/ocr/artifacts.ts` originals, manifest, sweep +- [ ] 4.4 GREEN: `src/modules/ocr/dispatcher.ts` leases; `src/modules/catalog/reconciler.ts` +- [ ] 4.5 GREEN: `src/modules/ingest/service.ts` branch; `src/app.ts` upload/status + +## 5 Review (Units 9–10; V1–V4) + +- [ ] 5.1 RED: auth/premature access; atomic corrections/stale 409; new/reusable activation race; reject/no embeddings +- [ ] 5.2 GREEN: `src/modules/ocr/review.ts` view, corrections, conflicts +- [ ] 5.3 GREEN: `src/modules/ocr/indexing.ts` CAS; new/reusable activate_requested=true activates, false stays ready +- [ ] 5.4 GREEN: rejection and `public/playground/` review UI + +## 6 Retention/Deploy (Units 11–12; O5/V5) + +- [ ] 6.1 RED: TTL/CAS/resume/active protection; flag-off native-only and candidate invisibility +- [ ] 6.2 GREEN: `src/modules/ocr/retention.ts` +- [ ] 6.3 GREEN: `src/api/openapi.ts` OCR/status/review contracts +- [ ] 6.4 GREEN: `src/config/env.ts`, `Dockerfile`; follow `docs/CONTRATO_CICLO_VIDA_Y_OCR.md` (read-only) + +## 7 E2E (Unit 13; V6) + +- [ ] 7.1 E2E: native 201; scanned 202→review→approve→active; mixed +- [ ] 7.2 E2E: resend, OCR-down fail-closed, catalog-down 503 +- [ ] 7.3 Gate: check, build, test, offline pytest +- [ ] 7.4 Production: flag-off verify→enable→FacturaTech 34 entries; CBG04a/FAT07/DSAU08/NSAV06 diff --git a/openspec/config.yaml b/openspec/config.yaml new file mode 100644 index 0000000..3be56d8 --- /dev/null +++ b/openspec/config.yaml @@ -0,0 +1,59 @@ +schema: spec-driven + +context: | + Project: rag-service (RAG knowledge service for enterprise AI tooling) + Tech stack: Node.js 22 (ESM), TypeScript 5.8 strict, Express 4, tsx runner + Architecture: modular service — src/modules/{ingest,process,parsers,retrieve,answer,catalog,embeddings,vectorstore,logs}, src/api (OpenAPI), src/shared + Data stores: PostgreSQL (pg) + Qdrant (@qdrant/js-client-rest) + OpenAI embeddings + Testing: node:test + node:assert/strict executed via `npm test` (NODE_ENV=test tsx --test tests/**/*.test.ts) + Quality: `npm run check` (tsc --noEmit, strict), `npm run build` (tsc emit to dist/) + Deploy: Docker + EasyPanel; migrations are manual scripts (migrate:lifecycle, migrate:legacy:lifecycle) + No linter, formatter, or coverage tooling is configured. + +strict_tdd: true + +testing: + projects: + - path: ./ + stack: Node.js 22 + TypeScript 5.8 (ESM) + test_command: npm test + test_framework: node:test (node --test via tsx) + layers: + unit: true + integration: true + e2e: false + coverage: + available: false + command: "" + quality: + linter: false + type_checker: true + type_checker_command: npm run check + formatter: false + +rules: + proposal: + - Include rollback plan for risky changes (migrations and Qdrant/PostgreSQL data operations) + - Reference existing module boundaries under src/modules/ + specs: + - Use Given/When/Then for scenarios + - Use RFC 2119 keywords (MUST, SHALL, SHOULD, MAY) + design: + - Document architecture decisions with rationale + - Keep module boundaries; parsers/ingest extensions must follow the parser-registry pattern + tasks: + - Group by phase, use hierarchical numbering + - Keep tasks completable in one session + - Respect review_budget_lines: 400 (auto-chain delivery when exceeded) + apply: + guidelines: + - Follow existing code patterns (ESM imports with .js extensions, strict types) + - Never read or expose .env* files, llaves, or backups/ contents + tdd: true + test_command: npm test + verify: + test_command: npm test + build_command: npm run build + coverage_threshold: 0 + archive: + - Warn before merging destructive deltas \ No newline at end of file diff --git a/openspec/specs/.gitkeep b/openspec/specs/.gitkeep new file mode 100644 index 0000000..e69de29