Jev in production › Evaluation and testing
Python toolkit that sends RAG, response and agent-trace metrics to Jev in one judge request and checks proposed tool calls against policy before they run, with optional threshold calibration against human labels.
Evaluate responses. Guard actions. A Python toolkit for evaluating LLMs, RAG, and Agents using Decision Models with optional calibration against human labels
Quickstart · Agent evaluation · Tool guards · Calibration · Examples & guides
Use Jev by default, Microsoft Decision-1 on OpenRouter through the same backend, or native OpenAI Decisions API, to judge generated responses and recorded agent executions. Check proposed tool calls before they run. Start with a preset, then bring your own metrics, thresholds, or judge backend as your application grows.
For the project's own README, linking back here: