Jev in production › Evaluation and testing

Typed Evals

Python toolkit that sends RAG, response and agent-trace metrics to Jev in one judge request and checks proposed tool calls against policy before they run, with optional threshold calibration against human labels.

Open on GitHub ↗

The source publishes no measured number.
Use
Evaluation and testing
Industry
Developer tools
Form
Open-source tool
Stage
In production
Plugs into
LangChain, CrewAI, Microsoft Agent Framework
Listed
2026-09-24
Found via
github
Repository
TrustifAI/typed_evals
Stars
16
Forks
1
Last push
2026-10-10
Language
Python
License
MIT

The README opens with

Evaluate responses. Guard actions. A Python toolkit for evaluating LLMs, RAG, and Agents using Decision Models with optional calibration against human labels

Quickstart · Agent evaluation · Tool guards · Calibration · Examples & guides

Use Jev by default, Microsoft Decision-1 on OpenRouter through the same backend, or native OpenAI Decisions API, to judge generated responses and recorded agent executions. Check proposed tool calls before they run. Start with a preset, then bring your own metrics, thresholds, or judge backend as your application grows.

Badge

For the project's own README, linking back here:

Listed in Jev in production

Also used for evaluation and testing