Enterprise AIOrganization: CashToken Rewards3 weeks · 2026

Building an AI Document Extraction Pipeline for 10 Formats

10
10 formats

Supported via a single POST endpoint

100%
100% streaming

Server-Sent Events, no blocking JSON responses

Zero-trust
Zero-trust

ClamAV, SSRF, XXE, and ZIP-bomb protection

The challenge

AI pipelines and knowledge bases needed a reliable way to ingest business documents from URLs and get clean, structured Markdown — not raw text dumps from PDF parsers.

Our approach

We built DocCrunch: an agentic REST API where Gemini 2.5 Pro orchestrates deterministic extractors (PyMuPDF, python-docx/pptx, openpyxl, BeautifulSoup) and a Gemini Flash OCR agent handles scanned pages. 100% SSE streaming, stateless, Docker + ClamAV sidecar on tmpfs RAM disk.

Technology stack
PythonFastAPIGemini 2.5 ProGemini 2.5 FlashDockerClamAVSSE

Facing a similar challenge?

We'll tell you how this approach adapts to your problem — and what it would take to ship it.

Book a strategy call