rag-service/ocr-service/README.md

44 lines
1.9 KiB
Markdown

# Private OCR Service
The OCR service is an internal RAG dependency. It accepts authenticated PDF jobs, processes them with one PaddleOCR worker, and stores transient artifacts in its private `/data/jobs` volume.
## Runtime guarantees
| Area | Behavior |
|---|---|
| Capacity | At most three queued or running jobs and one active OCR execution. |
| Recovery | One recovery after interruption; a second interruption or 15-minute total timeout is terminal. |
| Review images | PNG files are persisted during the OCR render and served without rendering the PDF again. |
| Storage | Admission and PNG publication stop safely when the storage budget or free-space reserve is unavailable. |
| Cleanup | Confirmed jobs are deleted by RAG; unconfirmed jobs expire after 24 hours. |
| Readiness | Requires the model, worker, sweeper, and storage to be operational. |
The API is private and its interactive documentation is intentionally disabled. Operational limits and deployment checks are documented in `../docs/OPERATIVA.md`; the full lifecycle contract is in `../docs/CONTRATO_CICLO_VIDA_Y_OCR.md`.
## Development environment
Use the repository-local virtual environment `ocr-service/.venv`. It is intentionally retained between runs and ignored by Git. Do not install these dependencies globally.
## Install
From the repository root:
```bash
ocr-service/.venv/bin/python -m pip install -r ocr-service/requirements.txt
```
Activate it with `source ocr-service/.venv/bin/activate`.
## Test
```bash
ocr-service/.venv/bin/python -m pytest ocr-service/tests
```
The container build downloads the pinned PaddleOCR models. Runtime startup loads those baked models before readiness can report healthy. Deploy one replica with a maximum of 3 CPU and 5 GiB RAM; Docker image labels document these platform-enforced limits.
Build from the repository root so the Dockerfile can use the service-scoped paths:
```bash
docker build -f ocr-service/Dockerfile -t rag-ocr-service .
```