Jev in production › Evaluation and testing
Routes agent-turn judgments to Jev before a larger model runs inside an eval harness, benchmarked 51% faster on 200 real turns (vendor).
One Python interface for evaluating what LLM apps and agents say and do. Use a frontier model as the judge, a small local model, or a decision model, and swap between them without rewriting your evals.
CascadeEvaluator asks Jev, a fast decision model, first. Jev reports how confident it is in each answer. Confident answers are kept; only the rest go to the LLM judge of your choice.
On 200 real agent turns (should the agent have called a tool here, or replied?), against Claude Sonnet 4.6 judging everything:
For the project's own README, linking back here: