Jev report · co-dm

Worth asking?

A typed-decision model was wired into five places in a Dungeons & Dragons co-DM. This page measures it against the code it replaces, against a frontier model doing the same job, and against the open-weight model marketed as its replacement.

19→88%table-rule violations caught, incumbent vs Jev
37→88%messages routed correctly
46×wasted spend prevented, per dollar of asking
31×faster than a frontier model, like for like
95%precision, where the frontier classifier managed 57%
01

What was actually tested

The test asked, for each judgement, what did the code do before, what does Jev do, and what does a frontier model do, on the same items, scored against the same answer key.

Four arms:

A — incumbent

The rule that shipped before: a regular expression, a path rule, or in two cases nothing at all.

B — Jev

Called through the product's own question builders and batching.

C — frontier

A frontier model called as a classifier — batched, thresholded, given the same information and nothing more. Run that way on the table rules only; on routing and contradictions its one run was the discarded hand-judged one, so those tables carry no frontier row.

D — Laya

ConvAI's open-weight typed-decision model, on a 12-thread CPU behind the same builders: as shipped, in a shape its 512-token window holds, and its 1024-token checkpoint.

E — Qwen 3.8

Qwen 3.8-flash read as a System 1 sampler: behind the same builders and batching, each question is one prefill-only call and the answer is read from the next-token probabilities, nothing generated. Also prompted for text on the table rules, beside GLM-5.3-flash and Haiku 4.5.

The answer keys were written blind. Whoever labelled a set was forbidden from calling Jev while doing it, and from importing the provider at all.

All four sets are published, each item with the request Jev was sent, its label and the labeller's note, in jev-eval-dataset-2026-10-03.zip (198 KB, JSON Lines with a README), with the skill-selection pairs.

Each set was then audited by a separate reviewer whose standing assumption was that the benchmark had been rigged, and whose first job was to compute what a trivial baseline scores. On the table-rule set, answering “clean” to everything scores 90.8% accuracy, so the figures reported below are recall and precision, not accuracy.

02

Qwen 3.8, read as a System 1 sampler

The claim tested: a general model read the way Jev is read — one pass over the prompt, the answer taken from the probability of the next token — does the same job. The arm scripts that measured Jev and Laya ran unchanged against a local shim that answers TypeSafe's API from qwen/qwen3.8-flash: a yes/no question is P(yes) over P(yes) + P(no); a choice is read two ways, as one letter (A, B, C…) and as one yes/no per option, because a single letter carries position bias. Same items, same answer keys, same scorers.

Every set, Jev against the Qwen 3.8 readout
SetJevQwen 3.8, one letterQwen 3.8, per optionIncumbent
Table rules, every stored run, per judgement — recall / precision92.8% / 96.3%
16 runs
91.8% / 94.9%
7 runs
19.2% / 100%
Contradictions, 60 statements — accuracy96.7%95.0%51.7%
Routing, 90 messages — correct lane87.8%51.1%67.8%36.7%
Asset kinds, 177 rows — correct kind82.5%73.5%87.6%57.6%
Asset kinds, 52 hard rows55.8%65.4%69.2%61.5%

Jev wins routing; Qwen 3.8 wins the art library; the other two are close. On the table rules, every stored run counted per judgement the way the table below counts, Jev catches 92.8% of violations at 96.3% precision over 16 runs and the readout 91.8% at 94.9% over 7: within a point on each. Both move from run to run, Jev between 80.8% and 100% recall and the readout between 80.8% and 100%. Contradictions differ by one statement. Routing is Jev's clearly: the readout's best reading routes 61 of 90 to Jev's 79. On asset kinds the readout, asked one yes/no per kind, scores 87.6% against Jev's 82.5%, and 69.2% against 55.8% on the hard rows.

Read as one letter, the readout loses to itself. The same model and the same questions score 51.1% on routing and 73.5% on assets when the answer is a letter, against 67.8% and 87.6% asked one option at a time. That costs one call per option: a routing question becomes five calls and an asset question six, 7.4 s median per question.

The whole Qwen 3.8 readout, every set and every repeat, 4,402 questions, was billed $0.30 by OpenRouter. No question went unanswered.

Table rules, the same model prompted for text — one run, 16 paragraphs per call, flag at 0.70, 282 judgements
ArmRecallPrecisionF1False alarmsWall clock
B — Jev, the same run scored the same way96.2%100%98.0%04.2 s
GLM-5.3-flash, reasoning low100%89.7%94.5%391.6 s
Qwen 3.8-flash, no reasoning69.2%52.9%60.0%1641.1 s
Haiku 4.561.5%48.5%54.2%1720.1 s

Prompted for text with reasoning off, Qwen 3.8 scores 60.0% F1 against Jev's 98.0% on the same scoring, with 16 false alarms to Jev's none. With reasoning on it did not finish the set: at reasoning_effort: "low" each batch of sixteen took 150 to 290 s, two of six passed the client's 300 s limit, and one more returned malformed JSON. Left on its default it, and GLM, spent all 8,000 tokens reasoning and returned no answer, the same failure the skill-selection study found.

03

Table rules: 19% to 88% of violations caught

94 paragraphs — 68 clean ones that shipped to a real table, 26 written to break exactly one of the three table rules, 51 of the 94 marked hard by the labeller. The rules: never narrate a player's character, never tell the DM when to roll, never end a sentence explaining what a look means.

94 paragraphs, 26 violations, 256 negative judgements, flag threshold 0.70
ArmRecallPrecisionFalse alarms
A — incumbent regex 19.2%
5/26
100% 0
B — Jev, mean of 8 runs 88.0%
±6
95.4% 9 in 2,048
C — gpt-5.6-sol, called as a classifier 96.2%
25/26
56.8% 19 in 256
D — Laya base, as shipped 0.0%
0/26
no flags 0 in 256
B — Jev, one paragraph and one rule per call 96.2%
25/26
92.6% 3 in 256
D — Laya base, one paragraph and one rule per call 34.6%
9/26
37.5% 22 in 256
D — Laya 1024, one paragraph per call, as shipped 96.2%
25/26
26.9% 135 in 256
D — Laya 1024, sixteen per call, as shipped 38.5%
10/26
12.8% 133 in 256
D — Laya 1024, one paragraph and one rule per call 100%
26/26
27.7% 253 in 256

A hand-judged arm was excluded. An Opus agent judging each item by hand scored 26 of 26. The audit found it had read the answer key's labelling notes and each item's provenance tag before judging, and provenance alone separates the 56 guaranteed negatives. Its score is excluded and is not reported as a ceiling.

These are averages over eight shuffled runs, not a single draw. A single run gave Jev 96.2% recall and 100% precision, but every run had reused one item order, so that result was one draw rather than an expected value. Re-run eight times with the order shuffled: precision averages 95.4% (92.3 to 100, reaching 100 in one run of eight) and recall averages 88.0% (80.8 to 92.3). Re-running the original order returned 88.5% rather than the published 96.2%.

Nine false positives across 2,048 negative judgements, concentrated in three paragraphs that straddle the 0.70 line. Jev is not deterministic: across the eight runs, precision ranged from 92.3 to 100 and recall from 80.8 to 92.3.

Eight more runs from the same evening score higher. The harness also stored four shuffled runs and four in the original order, made in the half hour before these eight with the same batching, threshold and scoring. They give 96.2% to 100% recall and 6 false positives in 2,048. Pooling all sixteen stored runs: recall 92.8% (386 of 416) and precision 96.3%, 15 false positives in 4,096. The table keeps the lower eight-run figure it first published; the pooled one is the better estimate.

Arm C calls a frontier model batched, thresholded and asked for a number. Its recall is 96.2% against Jev's 88.0%, with 19 false alarms in 256, thirteen of them on the hardest rule, against Jev's 9 in 2,048.

Against the batched, thresholded frontier model, Jev's recall is about eight points lower and its precision roughly forty points higher.

Laya cannot read the paragraph under the shipped call. It reads 512 tokens per question, head included, which leaves 384 for the state. The shipped state opens with the three rule texts, 416 tokens before the first paragraph, so no paragraph is ever inside its window, at sixteen per call or one per call. Its 282 scores are the same either way, identical across three repeat runs, and in three shuffled orders every one of the 846 judgements equals the shipped-order score of whichever paragraph held that index. All of them fall between 0.30 and 0.59, none reaches 0.70, so it flags nothing: recall 0, no false alarms, 251 s of CPU for the 94 paragraphs. That row is the shipped call, run unchanged. The rows beneath it ask the same shipped question in a shape the window holds: the paragraph first, then the one rule it is judged against, one question per call. Jev on that shape keeps its recall and gives up seven points of precision, three false alarms in 256 against none. Laya on it, with the whole state inside its window on 269 of the 282 calls, flags 9 of the 26 violations, 2 of them under the rule that was broken, and raises 22 false alarms in 256, at 605 ms a call. Laya's 1024-token checkpoint holds the whole shipped one-paragraph state, 434 to 757 tokens, in every call, and flags all 94 paragraphs: recall 96.2%, precision 26.9%, 135 false alarms in 256, at 1.5 s a call. On the shipped batches of sixteen it reads 893 of each call's 1,024 to 2,427 tokens, about the first five paragraphs, and flags 78 of the 94: recall 38.5%, precision 12.8%. On the fitted shape it answers yes to 253 of the 256 negative judgements, median score 0.95: every paragraph flagged under every rule.

The reason is not that the regex is badly written. It is that for one of the three rules no regex was ever written. Nobody could think of one. “Does this paragraph move a player's character?” has no surface form to match on, so that rule went unchecked from the day it was made a rule until the day it was asked as a question. The incumbent scores 0 of 9 on it by construction, not by accident.

04

Routing: the incumbent puts half of everything in one bucket

90 messages typed at the assistant, labelled with the lane a reasonable co-DM should run.

90 messages, five lanes — correct lane chosen
ArmOverall“chat” precisionNot-a-game-message
A — incumbent regex36.7%
33/90
0.26cannot express it
B — Jev87.8%
79/90
0.7411 of 11
D — Laya base31.1%
28/90
0.403 of 11
D — Laya 102422.2%
20/90
0.007 of 11

The incumbent has 45 false positives on a single lane. When no keyword matches, everything lands in “talk”. And it has only four lanes, so a message that is not about the game at all — a billing question, an attempt to talk the assistant out of its own instructions — has no lane of its own. On the four-lane subset, which leaves those messages out, it scores 41.8% against Jev's 86.1%.

Jev strictly dominates here: across all 90 messages there is not one the incumbent gets right and Jev gets wrong (46 the other way, p = 2.8×10-14). Re-run live it returned the same 90 of 90 per-item predictions; re-runs on the table rules varied. With every ambiguous item relabelled to its declared alternative, the incumbent scores 34.4% against Jev's 74.4%.

Injection handling. The incumbent routed 0 of the 11 account and injection messages into a lane that writes anything. It put them in chat and lookup, which are the wrong lanes and do not write anything.

Laya reads every routing message whole and still lands under the regex. Each message is 12 to 503 tokens, inside its window on all but one. It routes 28 of 90 correctly, 31.1% against the incumbent's 36.7% and Jev's 87.8%, and 31.6% on the four-lane subset. It sends 40 of the 90 to lookup. Of the eleven not-a-game messages it recognises three and puts two, both prompt injections, into world, the lane that writes canon; the incumbent put none in a writing lane. 470 ms per call on the CPU. The 1024-token checkpoint routes 20 of 90 correctly: it puts 53 of the 90 messages in the not-a-game lane, so it recognises 7 of the 11 that belong there and sends one of them to guide, the lane that writes a session. 216 ms per call.

05

Contradictions: nothing was checking before

60 statements written against the real canon of eight factions, each scored against the eight facts already on file for that faction. Here there was no incumbent check: the database returned the adjacent facts and a human read them. The benchmark built a lexical baseline instead: token overlap plus a polarity cue, with its threshold tuned on this same set.

60 statements — does this collide with anything on file?
ArmAccuracyPrecisionHard subset
incumbent — no check existed51.7%
base
n/a52.8%
tuned lexical baseline66.7%
0.66
0.65561.1%
Jev96.7%
27/29
1.0094.4%
Laya base, as shipped51.7%
0/29
no flags52.8%
Laya 1024, as shipped55.0%
29/29
0.51855.6%
Jev, one fact per call93.3%
27/29
0.93188.9%
Laya base, one fact per call46.7%
2/29
0.28647.2%
Laya 1024, one fact per call63.3%
25/29
0.58161.1%

The 51.7% is the base rate of never ruling, which is the current behaviour. The colliding fact comes back first 86.2% of the time and in the top three 96.6%, against an expected 4.3 rows read in arbitrary order.

Where it is weak. At the level of which fact collides, precision falls to 0.684 — at the item level (whether something here contradicts) precision is 1.00, and it names an extra non-colliding fact roughly a third of the time. The item-level 1.00 precision was measured on a set that is about half contradictions by construction; real proposals are mostly compatible. It is used as a ranker, not as a gate.

Laya, as shipped, rules on nothing here either. The statement and its eight facts run to 400 to 1,530 tokens; Laya's window holds the statement and the first five or six facts; the statement and its eight facts fit whole in 8 of 60 calls. Its highest score on any item is 0.59, so no item crosses 0.70: recall 0 of 29, and the 51.7% is the same base rate as having no check. Each call took 6.9 s against the shipped 4-second judge budget. Its 1024-token checkpoint holds 59 of the 60 states whole and flags 56 of the 60 items, every contradiction and 27 of the 31 compatible statements: 55.0%, precision 0.518, at 3.0 s per item inside the budget. One fact per call puts the whole state inside the window; Jev on that shape drops from 96.7% to 93.3%, two false alarms where the batched call raised none. Laya on it, every call whole inside its window, flags seven items of which two are contradictions: 46.7%, under the base rate, at 3.7 s per item across its eight calls. The 1024-token checkpoint on the same calls reaches 63.3%, precision 0.581, its highest result on any of the four sets, below the tuned lexical baseline's 66.7%.

06

What it costs

Jev's published price is $42 per billion input tokens, $0.042 per million, with no charge for the answers it returns (typesafe.ai). co-dm's meter records Jev calls in tokens rather than dollars, so the dollar figures below price the LLM turns from the meter and price Jev's own calls from that published rate.

For routing, the incumbent is a local pure function at microsecond latency and zero marginal cost. The table below prices the turn that runs afterwards in the wrong lane.

Routing errors priced from co-dm's own metered turns, over the same 90 messages
ArmGuide turns wastedGuide turns redoneCost of being wrong
A — incumbent regex217$2.62
B — Jev04$0.56
D — Laya base319$3.03
D — Laya 1024317$2.75

In the cost model, Jev's misroutes cost $0.023 less per message routed than the incumbent regex's, against roughly $0.0005 per message to ask Jev at the published rate — about 46 times its own cost.

That dollar column is a model, not a meter reading. No turn was executed in this run. It prices a missed guide request as one wasted cheap turn plus one full rebuild, at the $0.1235 an Opus guide turn actually cost across 29 metered calls. The model does not price a DM rephrasing after a miss, or a wrong guide reaching the table.

In the benchmark's own unit, with no pricing assumption: the incumbent leaves 36 turns short of their proper output budget, against Jev's 8, and the tokens Jev spends deciding are about 3.7× less than the budget shortfall it avoids.

07

Speed: the whole guide in under a second

4.2sJev — 94 paragraphs, 282 questions, 6 calls, serial
131sgpt-5.6-sol, same paragraphs, same batches, serial
184sgemini-3.8-flash, and 32 items lost to bad JSON
251sLaya base, same paragraphs, same batches, on a 12-thread CPU
3.5msthe regex, which catches a fifth as much

Quoted as serial call time on both sides, so the arms are compared under the same conditions. Run with six calls in flight, Jev finishes the guide in 773 ms of wall clock, but that is throughput rather than latency and the frontier arms were not given the same concurrency. Like for like it is 4.2 s against 131 s, about 31×. Latency for one batch is 751 ms.

The same 94 paragraphs and 282 judgements, every arm
ArmAnsweredWall clockOutput tokensCost
Jev94/944.2s6,198$0.0020
gpt-5.6-sol94/94131.5s8,209$0.0851
gemini-3.8-flash62/94184.2s28,845$0.1147
deepseek-v4.1-flash32/32 (subset)284.5s11,231$0.0072
Laya base, on a 12-thread CPU94/94 (no paragraph in its window)251.4s0$0 self-hosted
Laya 1024, on a 12-thread CPU94/94 (about five per batch in its window)247.5s0$0 self-hosted

Every frontier model wrote its reasoning out before answering: 20 to 33 seconds per batch of sixteen for the fastest of them, 191 seconds for a single batch of the cheapest. Jev writes a number, and a batch of sixteen comes back in about three quarters of a second. Laya writes a number too and spends the time in its forward pass instead: 42 seconds per batch of sixteen on the CPU.

deepseek-v4.1-flash costs a twelfth of what gpt-5.6-sol costs and took more than twice as long for a third of the work.

The shipped code runs both: the regexes run first, cost 3.5 ms for all 94 paragraphs and catch five real violations, and the judge is asked only about what is left.

08

Where it loses

  • It is not deterministic

    Eight runs over the same 94 paragraphs, shuffled, moved recall between 80.8% and 92.3% and precision between 92.3% and 100%. A single run can land on the best draw in that range.

  • The frontier ceiling was never actually measured

    The arm meant to establish it — claude-opus-5 reading each item and reasoning — scored perfectly on both judgements and turned out to have seen the answer key's own notes and each item's provenance before judging. It has been discarded, not quoted. So the ceiling for a careful frontier model is unknown here; what is measured is how one behaves as a deployable classifier, which is 56.8% precision against Jev's 95.4%. Establishing the real ceiling needs an uncontaminated run that has not been done.

  • The incumbent is partly being charged for a rule nobody wrote

    9 of the 26 violations test rule 1, which has no incumbent instrument at all, so those are automatic misses and are counted in the 19.2% headline. Restricted to the 85 items whose rules the regexes actually implement, the incumbent reaches 29.4% recall. The gap narrows and does not close.

  • The answer key shares an operationalisation with the thing it grades

    The labeller wrote the key against the literal wording of each rule — which is also close to the wording Jev is given in the question. Where that reading of a rule is wrong, both are wrong together and the benchmark scores it correct. Fifteen negatives are marked contested for this reason, and two of the three paragraphs that ever produced a false positive are among them. The routing set has the same shape: its labels were written from the same lane descriptions Jev reads, so 87.8% is agreement with those labels. The DM who owns these rules has reviewed neither key.

  • Two no-content shortcuts beat the incumbent

    On the table-rule set, “flag anything under 160 characters” scores 81.9% and “flag anything that was authored rather than taken from the guide” scores 87.2% — both above the incumbent's 77.7%, both reading nothing. The authored violations are shorter than the real paragraphs. On a length-matched slice Jev scores 97.9% and the trivial baseline scores 44.7%.

  • The cheap models are slow, not broken

    A single-paragraph probe had deepseek-v4.1-flash spending 1,500 output tokens and emitting nothing. Re-measured over 32 items it answered all 32 cleanly. It took 284 seconds for 32 paragraphs, about 8.9 seconds each, which puts one session guide at a quarter of an hour. gemini-3.8-flash took 184 seconds and 28,845 output tokens, and two of its six replies were malformed JSON, so 32 of the 94 paragraphs got no answer.

    A model that answers in prose has to be parsed, and parsing can fail. Jev's answer is typed, so there is nothing to parse. Across every run here it returned 94 of 94.

  • The answer key shares a blind spot with the thing it grades

    The labeller flagged it: 56 of the 68 clean paragraphs were written by this same assistant under these same three rules. Any rule it systematically misreads is baked into the key as correct. Fifteen labels were marked contested for exactly this reason, and the incumbent, Jev and the frontier arm all score 100% on that slice, because the slice cannot separate them.

  • On the art library, the evidence currently argues against using it

    The 287-of-300 figure below is a sample measurement, not a scored benchmark, and its ground truth comes from file names — which the incumbent path rule never sees, so that comparison is not neutral and is not claimed as one. A separate 177-row set was scored, and what it found is unflattering: a three-line rule that reads only the file extension beats the shipped path rule, and once the set is reweighted to the library's real composition — about 83% props — simply answering “other” to everything scores 0.83 against Jev's 0.77. On that judgement, at that prevalence, the model does not beat the baseline. It stays in proposals-only. On the same 177 rows, unweighted, the shipped path rule scores 57.6%, Jev 82.5%, the frontier model 96.6%, and always answering “other” 45.8%. Laya, which reads each 116-token state whole, scores 37.3%: it calls 95 of the 177 rows a token, where 24 are, at 589 ms a row. Its 1024-token checkpoint scores 17.0% and calls 162 of the 177 a token.

  • Fact triage is not benchmarked at all

    A set exists; the run is unfinished. No result is reported.

09

A sorting bug the classifier exposed

The art-library judgement (row four below) was sampled before it was trusted. The library holds 181,805 files sorted by the folder they sit in, 151,132 of them filed as tokens, the cut-out figures that stand on a battlemap. Three hundred were put to the model. Eight more, pulled by name from the same pile:

A wooden wall bracket for a banner
a bracket
A wooden chair lying on its side
a fallen chair
A wooden crate
a crate
A cloth display cushion
a cushion
A barrel tap
a barrel tap
Folded napkins in a silver ring
folded napkins
A single bed with a ladder
a bunk
A black piano bench
a piano bench

287 of the 300 were not tokens, at 0.91 average confidence. The paths behind the answers showed why: the library's root folder is named Assets and Tokens, and the sorting rule handed that word down to every bracket and cushion inside it. Maps had a guard against this; tokens did not.

The fix was one line in the sorting rule and corrected 59,345 files. The pattern appeared in the classifier's sample of 300.

10

What is actually switched on

Five judgements, and how far each is trusted today
JudgementShapeBenchmarkedStatus
Table rulesyes/no ×3yes Live. Regex first, judge on the rest, flags at 0.70.
Canon contradictionsyes/noset built, run unfinished Live. Sits on the database's answer before the writing model sees it.
Message sortingchoiceyes Shadow only. Logs both answers, changes no routing.
Art library kindschoicenot scored Proposals only. Committing them is a separate step.
Fact triagescore ×2not scored Ranking only. It promotes nothing to canon.

Without the API key, all five switch off and the assistant runs as it did before they were added. In every judgement the incumbent code still runs first, and the model only sees what it passes on.