rag-service/ocr-service/README.md

1.9 KiB

Private OCR Service

The OCR service is an internal RAG dependency. It accepts authenticated PDF jobs, processes them with one PaddleOCR worker, and stores transient artifacts in its private /data/jobs volume.

Runtime guarantees

Area Behavior
Capacity At most three queued or running jobs and one active OCR execution.
Recovery One recovery after interruption; a second interruption or 15-minute total timeout is terminal.
Review images PNG files are persisted during the OCR render and served without rendering the PDF again.
Storage Admission and PNG publication stop safely when the storage budget or free-space reserve is unavailable.
Cleanup Confirmed jobs are deleted by RAG; unconfirmed jobs expire after 24 hours.
Readiness Requires the model, worker, sweeper, and storage to be operational.

The API is private and its interactive documentation is intentionally disabled. Operational limits and deployment checks are documented in ../docs/OPERATIVA.md; the full lifecycle contract is in ../docs/CONTRATO_CICLO_VIDA_Y_OCR.md.

Development environment

Use the repository-local virtual environment ocr-service/.venv. It is intentionally retained between runs and ignored by Git. Do not install these dependencies globally.

Install

From the repository root:

ocr-service/.venv/bin/python -m pip install -r ocr-service/requirements.txt

Activate it with source ocr-service/.venv/bin/activate.

Test

ocr-service/.venv/bin/python -m pytest ocr-service/tests

The container build downloads the pinned PaddleOCR models. Runtime startup loads those baked models before readiness can report healthy. Deploy one replica with a maximum of 3 CPU and 5 GiB RAM; Docker image labels document these platform-enforced limits.

Build from the repository root so the Dockerfile can use the service-scoped paths:

docker build -f ocr-service/Dockerfile -t rag-ocr-service .