Jev in production › Agent context and memory
Benchmarks Jev as a context-compaction judge against retrieval and summarization on LoCoMo, across 486 questions over 10 conversations (jevtrim).

A comparative analysis on context compaction driven by calibrated judgements instead of summarization. Jev scores every chunk of a conversation for relevance to what is being asked, ordinary Python keeps the chunks that fit a token budget, and the result is auditable, deterministic and replayable offline.
Built on Jev (typesafe/jev-1.13), TypeSafe's System One decision model, reached through OpenRouter.
Left: accuracy against the token budget, conditioned track, 486 questions over 10 conversations. Right: answering with only the evidence (turn or chunk) scores below the best method.
For the project's own README, linking back here: