Jev in production › Evaluation and testing
An independent reproducible benchmark measured Jev at 80.5% accuracy on banking intents (nibzard).
DMB compares APIs that choose one option from a list and return a confidence score. It measures accuracy, latency, cost, failures, and how confidence relates to correctness. The contenders include TypeSafe AI's jev, LLMs with structured output, and deterministic baselines.
On September 29, 2026, jev-latest scored 79.2% on Banking77's full official test set and 88.6% on CLINC150 with out-of-scope queries. On NLU++, micro intent F1 was 48.3%, and only 4.3% of messages had every intent label correct. Median latency ranged from 267 to 316 ms per decision across these suites.
For the project's own README, linking back here: