Jev in production › Evaluation and testing

Jev Arena

A local dashboard benchmarks 13 decision-model profiles including Jev across 7,671 cases each, recording latency, accuracy and workflow-replay comparisons (author).

Open on GitHub ↗

7,671measured against a baseline, as published by the source
Use
Evaluation and testing
Industry
AI infrastructure
Form
Open-source tool
Stage
Beta
Listed
2026-09-29
Found via
github
Repository
theaiautomators/jev-arena
Stars
33
Forks
4
Last push
2026-09-28
Language
Python
License
MIT
Jev Arena screenshot
docs/arena-dashboard.png in the theaiautomators/jev-arena README, MIT; shown from GitHub.

The README opens with

A local app for comparing decision models: inspect their answers, measure response times and input limits, and replay decisions inside small software workflows.

The video demonstrates two completed assessments on Windows with an RTX 5090 (32 GB):

Assessment Scope Read the findings --------- Arena Full v2 13 profiles, 7,671 cases per profile, serial timing and recorded workflows Results · Task analysis ABCD support decisions Five profiles, 300 conversations, full handbook versus retrieved policy Results

Badge

For the project's own README, linking back here:

Listed in Jev in production

Also used for evaluation and testing