Jev in production › Evaluation and testing
A local dashboard benchmarks 13 decision-model profiles including Jev across 7,671 cases each, recording latency, accuracy and workflow-replay comparisons (author).

A local app for comparing decision models: inspect their answers, measure response times and input limits, and replay decisions inside small software workflows.
The video demonstrates two completed assessments on Windows with an RTX 5090 (32 GB):
Assessment Scope Read the findings --------- Arena Full v2 13 profiles, 7,671 cases per profile, serial timing and recorded workflows Results · Task analysis ABCD support decisions Five profiles, 300 conversations, full handbook versus retrieved policy Results
For the project's own README, linking back here: