Jev in production › Evaluation and testing
Python library that sends all of a trace's agent evals and guardrails (tool choice, groundedness, scope, indirect injection, PHI) to Jev as typed questions in one request, offline or inside the agent loop. Openlayer measured p50 244ms per request (company).
Evals and guardrails for agents, using Jev-style decision models instead of an LLM judge. All the evals for a trace go out as one request that costs a few thousandths of a cent and comes back in a few hundred milliseconds, so you can run them on every trace and inside the agent loop.
Works with Jev through the TypeSafe or Vercel APIs, with Kev or Laya running locally on a Mac, with Eikos on your own GPU, or with a regular chat LLM if that's all you have (slower, costs more).
For the project's own README, linking back here: