Jev in production › Document and record classification
Benchmark app in which Jev labels Docling-parsed arXiv papers by category and picks each title from candidate lines, raced against two Gemini 3.8 Flash lanes on the same papers. The Jev lane cost $0.0022 in total (author).
A parser and two models race to read documents: Docling → Jev (TypeSafe's decision model), Docling → Gemini 3.8 Flash, and Gemini 3.8 Flash reading the PDF itself. Every lane answers the same questions, and the answers are scored against arXiv's own metadata.
12 arXiv papers, 189 pages, 4 questions each. Every lane got 12/12 categories and 12/12 titles. The difference was speed and cost.
Lane Wall clock Decide only Cost Category Title --- ---: ---: ---: :---: :---: Docling → Jev 24.5 s 6.5 s $0.0022 12/12 12/12 Docling → Gemini 3.8 Flash 26.7 s 48.5 s $0.0364 12/12 12/12 Gemini 3.8 Flash reads the PDF 32.1 s 84.1 s $0.0882 12/12 12/12
For the project's own README, linking back here: