Jev in production › Evaluation and testing

Decision Model Benchmark (DMB)

An independent reproducible benchmark measured Jev at 80.5% accuracy on banking intents (nibzard).

Open on GitHub ↗

80.5%measured, as published by the source
Use
Evaluation and testing
Industry
Developer tools
Form
Open-source tool
Stage
In production
Listed
2026-09-27
Found via
github
Repository
nibzard/decision-model-benchmark
Stars
11
Forks
0
Last push
2026-10-07
Language
HTML
License
none stated

The README opens with

DMB compares APIs that choose one option from a list and return a confidence score. It measures accuracy, latency, cost, failures, and how confidence relates to correctness. The contenders include TypeSafe AI's jev, LLMs with structured output, and deterministic baselines.

On September 29, 2026, jev-latest scored 79.2% on Banking77's full official test set and 88.6% on CLINC150 with out-of-scope queries. On NLU++, micro intent F1 was 48.3%, and only 4.3% of messages had every intent label correct. Median latency ranged from 267 to 316 ms per decision across these suites.

Badge

For the project's own README, linking back here:

Listed in Jev in production

Also used for evaluation and testing