Jev in production › Document and record classification

Kyotofin tax-doc-classifier

Classifier where Jev reads each tax PDF page's text and returns a probability over 261 IRS forms and 7 page kinds, used across the company's tax document corpus. 34× cheaper per page than the prior Sonnet classifier (company).

Open on GitHub ↗

34× cheapermeasured against a baseline, as published by the source
Use
Document and record classification
Industry
Finance
Form
Open-source tool
Stage
In production
Listed
2026-09-19
Found via
github
Repository
kyotofin/tax-doc-classifier
Stars
503
Forks
62
Last push
2026-09-29
Language
TypeScript
License
Apache-2.0

The README opens with

We ingest thousands of tax documents using an LLM pipeline built last tax season. Jev classifies 100% of our tax document corpus at $0.001 per page — 34× cheaper and 6× faster than that LLM setup. This is the classifier, open sourced.

One request per page. The page's text goes to Jev, TypeSafe's decision model, which returns a probability over 261 IRS forms and 7 page kinds instead of text. No model is trained and nothing is hosted: the classifier is a JSON file describing each form, generated from the IRS's own PDFs.

Badge

For the project's own README, linking back here:

Listed in Jev in production

Also used for document and record classification