The least glamorous bottleneck in enterprise AI is not the model. It is the 400-page PDF that someone scanned crooked in 2009, and the analyst who has to retype the table on page 217 into a spreadsheet before the model gets to see it.
Doctly AI, a 2025-vintage company in Saratoga, California, is selling a way out of that room. Its pitch is straightforward: feed it a complex PDF, get back Markdown, JSON, or CSV through an API, and skip the manual data entry [doctly.ai, 2026]. The company is angel-funded, with the round size undisclosed [Crunchbase, 2026].
The wedge is the ugly PDF
Doctly's product targets the documents that defeat conventional OCR: regulatory filings, legal contracts, financial reports, insurance forms, and medical records [doctly.ai, 2026]. These are the files where a misread decimal point is a compliance event, and where the layout is the actual problem.
The company also offers instant custom extractor generation from drag-and-drop uploads. A workflow where a non-engineer points at a sample document and gets back a working extractor is closer to a product than a primitive.
Cofounder Ali Sheikh, who lists himself as CEO, has written that the company arrived at PDF parsing somewhat sideways, having started with a retrieval system for regulatory documents and discovered that the parsing layer was the part nobody had solved well enough to build on [Medium, 2026] [X, 2026]. That origin story explains why the wedge is narrow and why the initial customer profile is regulated-industry developer teams.
A crowded shelf
| Competitor | Distribution wedge |
|---|---|
| LlamaParse | Bundled with LlamaIndex, the default RAG framework |
| Unstructured | Enterprise document pipelines, well-funded |
| Vectorize | Comes in via the vector database layer |
| Doctly AI | Chat-to-API custom extractors for regulated docs |
Doctly's read appears to be that accuracy on genuinely hard documents, plus a faster path from sample to working extractor, is enough to peel off the buyers who care about a single misread number. That is a defensible thesis on a narrow segment.
Where the bet could break
- No public customers yet. The company has zero G2 reviews and no named deployments in public sources [G2, 2026].
- Open-source pressure from below. Mistral, among others, ships strong OCR models that a competent team can wire up in an afternoon. Sheikh has acknowledged the comparison directly in public [Hacker News, 2026].
- Headcount and runway. Crunchbase reports 1 to 10 employees [Crunchbase, 2026].
The math the buyer is doing
The customer is running a quiet calculation in their head. A mid-size compliance team might process, say, 10,000 complex PDFs a year. At fifteen minutes of analyst time per document at a fully loaded $80 an hour, that is roughly 2,500 hours, or about $200,000 a year, spent moving numbers from a PDF into a spreadsheet. If a parsing API can take that down by 80 percent at a software cost of $30,000 to $50,000, the ROI math closes inside a quarter. That is the conversation Doctly needs to be in.