Enterprise AIOrganization: CashToken Rewards3 weeks · 2026
Building an AI Document Extraction Pipeline for 10 Formats
10
10 formats
Supported via a single POST endpoint
100%
100% streaming
Server-Sent Events, no blocking JSON responses
Zero-trust
Zero-trust
ClamAV, SSRF, XXE, and ZIP-bomb protection
The challenge
AI pipelines and knowledge bases needed a reliable way to ingest business documents from URLs and get clean, structured Markdown — not raw text dumps from PDF parsers.
Our approach
We built DocCrunch: an agentic REST API where Gemini 2.5 Pro orchestrates deterministic extractors (PyMuPDF, python-docx/pptx, openpyxl, BeautifulSoup) and a Gemini Flash OCR agent handles scanned pages. 100% SSE streaming, stateless, Docker + ClamAV sidecar on tmpfs RAM disk.
Technology stack
PythonFastAPIGemini 2.5 ProGemini 2.5 FlashDockerClamAVSSE
Facing a similar challenge?
We'll tell you how this approach adapts to your problem — and what it would take to ship it.
Book a strategy call