Jev in production › Evaluation and testing

Jev Benchmark (ywchiu)

Benchmarks Jev against open-source decision models Clef, Cygnet and Gemma in a simulated enterprise RAG agent routing scenario (ywchiu).

Open on GitHub ↗

The source publishes no measured number.
Use
Evaluation and testing
Industry
Developer tools
Form
Write-up
Stage
Announced
Listed
2026-10-04
Found via
github
Repository
ywchiu/jev_benchmark
Stars
6
Forks
0
Last push
2026-10-04
Language
Python
License
none stated

The README opens with

What Is Jev? System One Models, Applications, Open-Source Ecosystem, and Benchmark Comparisons https://www.largitdata.com/blog/jev-system-one-model-open-source-benchmark/

TypeSafe introduced Jev, a hosted model that works differently from a typical text-generation model. You give it a situation and a set of choices, and instead of generating a long response, it selects one of those choices and returns a probability.

Soon after Jev was released, the community started building several open-source alternatives with a similar goal. We wanted to see how these systems compare with Jev in a realistic Agent routing scenario.

Badge

For the project's own README, linking back here:

Listed in Jev in production

Also used for evaluation and testing