Jev report · first run 17 Sep 2026, rerun on equal terms 3 Oct 2026 · jev-1.13.0
A résumé-tailoring tool asks Jev, once per skill, whether a job posting calls for it. This study tests that choice against the alternatives: 743 skills, 12 postings, 8,916 decisions per arm. The first run gave the arms unequal treatment, so every arm was run again on equal terms, against 720 freshly labelled pairs graded by three model families, plus the same postings as page images.
1. task dispatch:0.00 and some as 0.85:29; the first-run parser read only 1:0.00. Reading both formats recovers 420 of the 597 the same settings dropped on the rerun. 177 answers, 2.0%, are genuinely missing from Haiku's replies; GLM's 180 were all format.Before the scoreboard, the thing under test. One call carries the posting once as state and sixty independent questions about it; the answers come back typed, so there is nothing to parse and nothing to retry.
{
"state": { "job_posting": "…" }, // 6.3 KB, sent once
"model": "jev-latest",
"questions": {
"q0": {
"type": "noul",
"instructions": "Does this posting call for
\"RMM\", either by name or by describing
work that requires it?",
"criteria": {
"true": "The posting asks for it…",
"false": "The posting does not need it"
}
},
"q1": { … 59 more … }
}
}
{
"answers": {
"q0": { "noul": 0.99 },
"q1": { "noul": 0.07 },
…
}
}
// No "return JSON only, no prose, no
// code fence". No regex over a reply.
// No retry when the fence comes back // wrong.
//
// The other inference scripts in this // repo do all three.
// This work moved onto Jev to
// avoid all three.
Per posting, averaged over the twelve in the benchmark below — 103.9s and 941,194 tokens for all of them
Every alternative below has to either produce the same sixty numbers from a text reply, or find them some other way. The alternatives: GLM-5.3-flash and Haiku 4.5 answering the same sixty questions as text, two embedding arms, a substring matcher, and ConvAI's Laya, an open-weight typed-decision model (base checkpoint, 421M parameters, Apache 2.0) whose model card compares it with Jev on figures it states were "third-party published, never measured here". Laya ran on a 12-thread CPU with no GPU, with the same noul question and criteria as the Jev arm, the same batches of 60 and the same posting as the only state; its package truncates the state to the checkpoint's 512 tokens, which is how it ships.
Two arms run Qwen 3.8 (qwen/qwen3.8-flash, $0.15 in and $0.47 out per million tokens) through OpenRouter. The first prompts it exactly as GLM and Haiku are prompted, with reasoning_effort: "none". The second reads it as a System 1 sampler: one call per skill with the posting first and the noul question last, max_tokens: 1, and the score is P(yes) / (P(yes) + P(no)) from the next-token logprobs. Nothing is generated or parsed; the answer is the prefill. It is the same mechanism as Jev's without a trained decision head, on the Qwen 3.8 weights, and it answers one question per call where Jev answers sixty. Sixteen calls ran in flight; every other arm ran sequentially.
The first run, further down, gave the arms unequal treatment in four ways that could move a result. The rerun removes each one and runs every arm again.
| In the first run | In the rerun |
|---|---|
| Jev, GLM and Haiku answered 60 skills per call; the Qwen 3.8 readout answered one per call and read the posting 743 times. | Every model runs both ways: 60 skills per call with the posting once, and one skill per call with the posting every time. |
GLM ran at reasoning_effort: "low", Qwen 3.8 at "none". | Every reasoning model runs at "low", the common setting: GLM rejects "none" with HTTP 400. Qwen 3.8 and Haiku also run without reasoning, labelled as such. |
| The readout ran 16 calls in flight, every other arm one at a time. | Every API arm runs eight calls in flight. Speed comes from a separate benchmark: the same 2 postings × 120 skills, every arm back to back in one window, forwards and then in reverse, because Qwen 3.8 slowed several-fold and returned HTTP 429 during the full runs. Calls refused with 429 were made again and are not counted against the model. |
| The 240 adjudicated pairs were chosen where Jev, GLM and Haiku disagreed most, and graded by Claude-family models. | 720 fresh pairs: 40 per posting drawn at random from all 743 skills (480), and 20 per posting where all eleven text model arms disagree most, among skills the posting never names (240). Claude Sonnet 5, GPT-5.5 and Gemini 3.1 Pro each grade every pair alone, with the full posting and no arm's score; the label is the majority. All three agree on 642 of 720. |
The parser read only 1:0.00 and counted any other line as a lost answer. | The parser also reads 1. name:0.00 and 0.85:29, counts them as format deviations, and keeps every reply. A lost answer is one missing from the reply, and it scores 0.5 on every metric rather than being dropped. |
Random AUC is ranking quality on the 480 random pairs, which no arm chose: the unbiased figure. Hard AUC is the same on the 240 hardest pairs. Decision is balanced accuracy at a yes/no cutoff tuned on six postings' random pairs and scored on the other six, then swapped, the identical procedure for every arm. Domain AUC is the first run's domain-label metric. vs Jev is a paired bootstrap of the AUC difference; grey intervals include zero. $ per 1,000 is OpenRouter's billed figure per call, and TypeSafe's published $42 per billion input tokens for Jev.
| Arm | Random AUC | vs Jev | Hard AUC | vs Jev | Decision | Domain AUC | $ per 1,000 | ms per decision | Lost |
|---|---|---|---|---|---|---|---|---|---|
| Jev, 60 per call | 0.995 | reference | 0.850 | reference | 0.961 | 0.633 | $0.0037 | 8.8 | 0 |
| Jev, 1 per call | 0.994 | −0.001 [−0.002, +0.000] | 0.850 | +0.001 [−0.005, +0.007] | 0.958 | 0.634 | $0.0658 | 36.5 | 0 |
| Qwen 3.8, logprob readout | 0.982 | −0.013 [−0.024, −0.005] | 0.878 | +0.028 [−0.010, +0.070] | 0.924 | 0.648 | $0.0439 | 228.8 | 0 |
| GLM-5.3-flash, 1 per call | 0.969 | −0.026 [−0.053, −0.006] | 0.814 | −0.036 [−0.100, +0.028] | 0.914 | 0.693 | $0.0436 | 216.9 | 0 |
| GLM-5.3-flash, 60 per call | 0.962 | −0.032 [−0.052, −0.017] | 0.687 | −0.163 [−0.243, −0.079] | 0.879 | 0.687 | $0.0046 | 35.6 | 0 |
| Qwen 3.8, 60 per call, reasoning low | 0.961 | −0.034 [−0.064, −0.013] | 0.759 | −0.091 [−0.166, −0.012] | 0.873 | 0.668 | $0.0514 | 2,079.4 | 169 |
| Qwen 3.8, 1 per call | 0.960 | −0.035 [−0.057, −0.014] | 0.744 | −0.106 [−0.182, −0.029] | 0.916 | 0.669 | $0.0395 | 235.0 | 0 |
| Haiku 4.5, 1 per call | 0.948 | −0.046 [−0.083, −0.019] | 0.705 | −0.144 [−0.212, −0.080] | 0.926 | 0.645 | $1.4816 | 149.6 | 0 |
| Haiku 4.5, 60 per call, reasoning low | 0.935 | −0.059 [−0.095, −0.029] | 0.720 | −0.130 [−0.208, −0.050] | 0.873 | 0.644 | $0.2539 | 106.5 | 187 |
| Qwen 3.8, 60 per call | 0.901 | −0.094 [−0.149, −0.044] | 0.453 | −0.397 [−0.480, −0.314] | 0.873 | 0.629 | $0.0058 | 32.9 | 59 |
| Haiku 4.5, 60 per call | 0.885 | −0.109 [−0.163, −0.061] | 0.498 | −0.351 [−0.444, −0.260] | 0.847 | 0.619 | $0.0621 | 12.7 | 177 |
| embeddings, per-sentence | 0.817 | −0.178 [−0.233, −0.123] | 0.580 | −0.270 [−0.354, −0.181] | 0.736 | 0.581 | local | — | 0 |
| lexical (substring) | 0.702 | −0.293 [−0.363, −0.220] | 0.558 | −0.292 [−0.367, −0.212] | 0.683 | 0.585 | local | 0.3 | 0 |
| Laya, full posting (8,192 tokens) | 0.505 | −0.490 [−0.559, −0.417] | 0.532 | −0.318 [−0.405, −0.232] | 0.490 | 0.493 | local | 6,449 | 0 |
| Laya, as shipped (512 tokens) | 0.503 | −0.492 [−0.564, −0.424] | 0.583 | −0.267 [−0.350, −0.182] | 0.487 | 0.506 | local | 857 | 0 |
One skill per call helps the chat models; it does not help Jev. Asked one skill at a time, Haiku 4.5 rises from 0.885 to 0.948 on the random pairs and Qwen 3.8 from 0.901 to 0.960. Jev scores 0.995 and 0.994 either way, at 18× the cost per decision for one per call. Batching sixty questions costs the chat models accuracy and costs Jev nothing.
The random pairs are easy and the hard pairs are not. 62 of the 480 random pairs are yeses, and the best arms sit near the ceiling there; the hard pairs are where arms separate. On them, neither the Qwen 3.8 readout nor GLM asked one skill per call separates from Jev.
Lost answers are real but small. Haiku omits 177 of 8,916 answers asked 60 per call and 187 with reasoning on; Qwen 3.8 omits 59 and 169. Jev, GLM, Laya and every one-per-call arm lose none.
Jev takes text only. Every posting was rendered to page images, as a screenshotted or photographed job ad would arrive, and each arm read those instead.
Each posting is wrapped at 100 characters and split into pages of 52 lines; headless Edge screenshots each page at 1000 × 1400, 34 pages for the twelve postings. GLM-5.3-flash, Haiku 4.5 and Qwen 3.8 read the page images directly. Jev reads the text that the OCR engine built into Windows recovers from the same pages, as a pipeline in front of Jev would; that OCR misreads some lines (“USA” became “CIS”) and drops Markdown labels.
| Arm | Read as text | Read from images | Change | Hard AUC, images | $ per 1,000, images | Lost |
|---|---|---|---|---|---|---|
| Jev on OCR text | 0.995 | 0.993 | −0.001 | 0.851 | $0.0037 | 0 |
| Qwen 3.8, logprob readout | 0.982 | 0.958 | −0.024 | 0.883 | $0.5911 | 0 |
| GLM-5.3-flash, 60 per call | 0.962 | 0.928 | −0.035 | 0.639 | $0.0091 | 114 |
| Qwen 3.8, 60 per call | 0.901 | 0.870 | −0.031 | 0.581 | $0.0143 | 1 |
| Haiku 4.5, 60 per call | 0.885 | 0.760 | −0.126 | 0.565 | $0.1161 | 0 |
| lexical on OCR text | 0.702 | 0.700 | −0.002 | 0.527 | local | 0 |
OCR costs Jev nothing measurable on these pages: 0.993 against 0.995, and the OCR step runs locally. Every vision model ranks lower reading images than reading text, and the loss is largest asked sixty skills per call: Haiku falls 0.126. The Qwen 3.8 readout holds best of the vision arms, at 0.958 and $0.5911 per 1,000 decisions, 13× its text cost, because the page images are not cached between calls.
Everything below is the first run as published on 17 September, with the Qwen 3.8 arms added on 2 October. Its figures stand as measured. Two of its findings did not hold up in the rerun, and each is corrected where it appears.
It affected every result below, so it comes first.
Scoring the 8,916 skill–posting pairs required labels; none were hand-labeled. The first definition: a skill is relevant to a posting if the posting names it verbatim.
It is also exactly what the lexical arm computes. The ground truth had been defined as one contender's output. Here is what that produced:
| Arm | AUC | |
|---|---|---|
| lexical (substring) | 0.960 | 0.96 |
| Jev | 0.927 | 0.93 |
| Qwen 3.8, logprob readout | 0.910 | 0.91 |
| GLM-5.3-flash | 0.876 | 0.88 |
| Qwen 3.8-flash | 0.870 | 0.87 |
| Haiku 4.5 | 0.861 | 0.86 |
| embeddings, per-sentence | 0.788 | 0.79 |
| embeddings, whole doc | 0.637 | 0.64 |
| Laya (base, local CPU) | 0.495 | 0.50 |
The top row is the label definition scoring itself; every other row is scored against the output of a substring match (str.find()).
The labels were redefined by domain membership — a property of the skill, not of the posting's wording, so no arm can match it by string. Then one more constraint: only the 421 of 747 skills (412 of 743 when first published) whosedomains_source is master, meaning the membership was written by hand in the catalogue source rather than inferred by Jev in an earlier pass.
What stays a judgement: the domains for each posting were assigned by hand, from its title and stated responsibilities. That mapping is published in full in eval_clean_labels.py with the phrase from each posting that drove each line.
Labels no arm can match by string. 421 skills × 12 postings; between 37 and 248 positives per posting. Every arm is scored against the catalogue's current 421 master-labelled skills; the first run had 412, which moves the arms by about a hundredth.
| Arm | AUC mean | pooled | per-posting range | |
|---|---|---|---|---|
| GLM-5.3-flash | 0.691 | 0.69 | 0.702 | 0.434 – 0.855 |
| Qwen 3.8, logprob readout | 0.644 | 0.64 | 0.666 | 0.384 – 0.800 |
| Qwen 3.8-flash | 0.628 | 0.63 | 0.621 | 0.480 – 0.748 |
| Jev (noul) | 0.627 | 0.63 | 0.653 | 0.376 – 0.786 |
| Haiku 4.5 | 0.621 | 0.62 | 0.628 | 0.446 – 0.817 |
| embeddings, whole doc | 0.585 | 0.59 | 0.561 | 0.444 – 0.789 |
| lexical (substring) | 0.581 | 0.58 | 0.583 | 0.467 – 0.730 |
| embeddings, per-sentence | 0.577 | 0.58 | 0.559 | 0.360 – 0.747 |
| Laya (base, local CPU) | 0.518 | 0.52 | 0.508 | 0.314 – 0.612 |
Two results.
Ahead on 11 of 12 postings, mean difference +0.067, sign test two-sided p = 0.0063. Bootstrapped over 306,850 ranked pairs the pooled difference is +0.0597, 95% CI [+0.0581, +0.0612] — excluding zero. Haiku, by contrast, is indistinguishable from Jev (5 of 12, p = 0.77). Laya trails Jev on 11 of 12, mean difference −0.109, p = 0.0063. Neither Qwen arm separates from Jev on this metric: the logprob readout leads on 8 of 12 (+0.017, p = 0.39) and the prompted arm on 5 of 12 (+0.001, p = 0.77).
The best arm scores 0.68 where 0.50 is guessing. The labels are set per domain: every skill in a posting's domain is labelled relevant, so on a Security posting every security skill is a positive. A ranking that puts only the skills the posting asks for at the top does not score 1.0 against these labels; the ceiling was not measured. Jev and the free substring matcher are not statistically separated at this sample size (3 of 12 to lexical, p = 0.146). Jev leads on the pooled figure; the per-posting lead is not established.
Every arm ran on the same machine and the same postings. All but the Qwen 3.8 logprob readout ran sequentially in batches of 60; the readout answers one question per call and ran sixteen calls in flight.
| Arm | Wall clock | Per decision | Tokens in | Tokens out | Cost | Answers lost |
|---|---|---|---|---|---|---|
| Jev | 103.9s | 11.7 ms | 781,642 | 159,552 | $0.0328 | 0 |
| Haiku 4.5 | 352.8s | 39.6 ms | 284,622 | 53,155 | $0.5504 | 572 |
| GLM-5.3-flash | 737.4s | 82.7 ms | 252,224 | 67,448 | $0.0429 | 0 |
| Qwen 3.8-flash | 1,186.3s | 133.0 ms | 269,559 | 69,768 | $0.0732 | 0 |
| Qwen 3.8, logprob readout | 1,041.7s | 116.8 ms | 11,236,713 | 8,916 | $1.6897 | 0 |
| embeddings | 130.3s | 14.6 ms | 29,989 | 0 | $0.0006 | 0 |
| lexical | 3.0s | 0.3 ms | 0 | 0 | $0 | 0 |
| Laya (base, local CPU) | 7,637.7s | 856.6 ms | local | local | $0 | 0 |
572 answers were not parsed. That is 6.4% of Haiku's answers — lines that did not match the expected format, in a run where the format was one number per line at temperature 0. Nothing raised an error. Several postings came back with 518 or 636 of 743 skills scored and the pipeline carried on, the harness counted the missing lines. Jev and Laya return typed values with no text to parse, and each lost 0 answers.
Corrected in the rerun: most of these lines were answers in a format the parser did not read. Haiku writes some lines as 1. task dispatch:0.00 and some as 0.85:29. With the same settings the rerun parser dropped 597 lines reading only 1:0.00; reading both formats recovers 420 of them. 177, 2.0%, are missing from the replies.
GLM is a reasoning model, and at a 784-token budget, sized for 60 answer lines, it spent every token thinking and returned empty content — 0 of 60 parsed on every batch. Measured, per batch of 60:
budget 784, thinking on 0/60 6.4s 784 out budget 4000, thinking on 60/60 28.5s 3,895 out $0.00129 budget 12000, effort low 60/60 4.7s 361 out $0.00023
Left on default, the model billed as one-eleventh the price of Haiku costs 5.6× and takes 6× as long as itself. For GLM, the reasoning setting changed the output-token count per batch, and with it the cost per batch. (enable_thinking: false was tried; the provider ignored it. reasoning_effort is what works.) Every GLM number above uses the pinned configuration.
Qwen 3.8-flash does the same. At the 784-token budget it parsed 0 of 60 on a test batch; at 12,000 tokens with reasoning_effort: "low" it parsed 60 of 60 in 54.1s; with "none", 60 of 60 in 5.3s. Both Qwen arms use "none".
Jev uses nearly three times the tokens of GLM or Haiku and costs less than either.
Jev used 941,194 tokens against GLM's 319,672 for the same 8,916 decisions. Each noul carries its own instructions and criteria text, sixty per call, where the generative prompt lists sixty bare skill names once. The per-token price is $42 per billion input tokens, with output not priced separately.
| Arm | Tokens | Cost | Versus Jev |
|---|---|---|---|
| Jev | 941,194 | $0.0328 | — |
| GLM-5.3-flash | 319,672 | $0.0429 | 1.3× dearer |
| Qwen 3.8-flash | 339,327 | $0.0732 | 2.2× dearer |
| Haiku 4.5 | 337,777 | $0.5504 | 16.8× dearer |
| Qwen 3.8, logprob readout | 11,245,629 | $1.6897 | 51× dearer at list |
| Laya (base, local CPU) | no API | $0 | self-hosted; 73× the wall clock |
One assumption: TypeSafe prices input and says nothing about output, so the $0.0328 counts input only. Charge Jev's 159,552 output tokens at the same rate and it becomes $0.0395 — still cheaper than both, 13.9× under Haiku.
The logprob readout sends the posting once per skill, so it reads 14× the input tokens of Jev. Every call for a posting shares the posting as a prefix, and OpenRouter served part of each repeat from cache: on a test pair, 768 of 941 prompt tokens came from cache and the second call cost 3.6× less than the first. The $1.6897 is list price with no cache discount. Rerun on one posting with billing recorded, the readout was billed $0.0290 against $0.1053 at list, with 81% of its input tokens served from cache: 0.275 of list. The rerun recorded billing on every call: the full readout was billed $0.3913, 12× Jev.
Break-even against GLM on this workload is $0.0456 per million input tokens; Jev's published rate of $0.042 per million (typesafe.ai) is under it.
Every model arm spends most of its ranking on skills the posting never names:
jev 286/480 60% glm 289/480 60% haiku 297/480 62% qwen38 287/480 60% qwen38-sys1 283/480 59% laya 439/480 91% (and 1.8% of its noes correct on the adjudicated rows) lexical cannot rank a single one of them.
Lexical cannot rank any of those picks, and the domain metric cannot grade them — domain labels are too coarse to adjudicate an individual skill, and there is no automatic ground truth for “this posting implies Terraform without saying it.”
They were graded by hand. The harness writes the cases where the arms most disagree — one confident yes, another confident no, on a skill the posting never names — to eval-disagreements.json: 240 rows, twenty per posting. Each was then judged twice, independently, one pair at a time with the full posting in view and the arms' scores withheld.
| Measure | Value | Reading |
|---|---|---|
| Agreement between passes | 215 / 240 = 89.6% | — |
| Cohen's kappa | 0.791 | substantial |
| Yes-rate, pass A / pass B | 48% / 49% | balanced, not a rubber stamp |
| Agreed rows either pass called arguable | 120 / 215 | hard cases |
The 215 rows both passes agreed on — 103 yes, 112 no. Answering “yes” to everything scores 48%.
| Arm | Accuracy | Caught the yeses | Rejected the noes | AUC | Says yes | |
|---|---|---|---|---|---|---|
| Qwen 3.8, logprob readout | 67.4% | 36.9% | 95.5% | 0.853 | 0.85 | 20.0% |
| Jev | 73.0% | 78.6% | 67.9% | 0.793 | 0.79 | 54.4% |
| GLM-5.3-flash | 60.0% | 22.3% | 94.6% | 0.732 | 0.73 | 13.5% |
| Qwen 3.8-flash | 64.7% | 30.1% | 96.4% | 0.723 | 0.72 | 16.3% |
| Haiku 4.5 | 36.7% | 45.6% | 28.6% | 0.384 | 0.38 | 59.1% |
| Laya (base, local CPU) | 48.8% | 100.0% | 1.8% | 0.610 | 0.61 | 99.1% |
The Qwen 3.8 logprob readout ranks best and decides worse than Jev at a 0.5 cutoff. Its AUC is 0.853 against Jev's 0.793; a paired bootstrap over the 215 rows puts the difference at +0.060, 95% CI [+0.022, +0.101]. It says yes to 20.0% of pairs, so at 0.5 it rejects 95.5% of the noes and catches 36.9% of the yeses, for 67.4% accuracy against Jev's 73.0%. The ranking is better; the cutoff that turns it into a decision was not tuned.
Corrected in the rerun: this gap came from the rows. These 240 were selected where Jev, GLM and Haiku disagreed most, so one of the three was confidently wrong on each, and the Qwen 3.8 arms, added afterwards, were not part of the selection. On 240 rows selected from every arm's disagreement, the readout leads Jev by +0.028, 95% CI [−0.010, +0.070], which is not separated; on 480 random rows Jev leads it by 0.013, which is.
GLM and prompted Qwen 3.8 say yes least often. GLM says yes to 13.5% of these pairs and rejects 94.6% of the true noes, catching 22.3% of the skills a posting calls for; prompted Qwen 3.8 says yes to 16.3% and catches 30.1%. Jev is the only arm near balanced.
Haiku ranks below chance at 0.384 (chance is 0.5). Haiku and the adjudicators are both Claude-family models; Haiku has the lowest AUC of the six arms graded here.
Restricting to the 95 rows neither pass flagged arguable, the gap widens rather than closing:
| Arm | Accuracy | Caught the yeses | Rejected the noes | AUC | |
|---|---|---|---|---|---|
| Qwen 3.8, logprob readout | 84.2% | 61.5% | 100.0% | 0.994 | 0.99 |
| Jev | 86.3% | 94.9% | 80.4% | 0.937 | 0.94 |
| Qwen 3.8-flash | 77.9% | 51.3% | 96.4% | 0.844 | 0.84 |
| GLM-5.3-flash | 66.3% | 25.6% | 94.6% | 0.735 | 0.74 |
| Haiku 4.5 | 32.6% | 48.7% | 21.4% | 0.411 | 0.41 |
| Laya (base, local CPU) | 42.1% | 100.0% | 1.8% | 0.583 | 0.58 |
The two metrics ask different questions. The domain metric asks which region of a catalogue to search, and GLM is better at that. The adjudicated metric asks does this posting call for this specific skill, which is the decision the tool makes. On these rows the Qwen 3.8 logprob readout scores highest at AUC 0.853, then Jev at 0.793 and GLM at 0.732; on the rerun's fresh rows Jev ranks highest on the random pairs and no arm separates from it on the hard ones. GLM says yes to 13.5% of pairs and catches 22.3% of the yeses.
Three caveats. The adjudicators are models, not people — careful ones, judging one pair at a time with the whole posting, which is a stronger setting than any scored arm was given, but not human ground truth. They are Claude-family, so Haiku's row shares a lineage with its judge and should be the most suspect; it is also the worst, which cuts against the bias rather than for it. And these 240 pairs are where the arms disagreed most — deliberately the hardest in the set, so treat them as a floor on difficult cases rather than typical performance.
What changed in the pipeline on this evidence: nothing. Jev stays. In the rerun it ranks best on the random pairs, decides best at a tuned cutoff, is the fastest and cheapest model arm, and cannot drop an answer. On the hardest pairs the Qwen 3.8 logprob readout is not separated from it, at 12× its billed cost and 26× its time per decision.
Haiku 4.5 is out. Asked one skill per call it is the most expensive arm, $1.4816 per 1,000 decisions, 400× Jev; asked sixty per call it omits 2.0% of its answers and ranks the hard pairs at 0.498, chance. GLM-5.3-flash with reasoning_effort pinned is kept as a cheap second opinion where a false positive costs more than a miss — it rejects 94.6% of true noes — but at 22.3% recall it cannot be the thing that picks what goes on a résumé.
Each has a harness in the repo that produced it. Three changed the pipeline's code. One withdrew a recommendation that had already gone out.
The first runs sent 45 KB of career history as the state for every question. ConnectWise came back as named, no work behind it — against twelve years of running MSPs on it. Calls took minutes. Trimming to a few hundred relevant characters fixed the answer and took each call to 11 ms. This matches the docs' guidance.
Trimming a job posting the same way moved the answers:
| State | Chars | Time | Mean shift | Answers moved ≥ 0.20 |
|---|---|---|---|---|
| full posting | 6,350 | 0.78s | baseline | — |
| requirements only | 2,614 | 0.49s | 0.045 | 1 |
| headline only | 700 | 0.51s | 0.223 | 11 of 30 |
At 700 characters, “primary escalation point for live-fire incidents” fell from 0.89 to 0.23 — on a posting about incident management.
The rule in use now: trim a career history, do not trim a job posting. A career history is a haystack: any one question touches a small part of it. Trimming the 45 KB history to a few hundred relevant characters fixed the ConnectWise answer. A job posting is the subject — every question asked is about that document. Cutting it to the 700-character headline moved 11 of 30 answers by 0.20 or more; cutting it to the requirements moved 1. Where the threshold turns over has not been located; 6.3 KB of subject was fine and 45 KB of haystack was not.
jev_state_size.py · 30 questions × 3 state sizes
Same 25 questions, same state, three runs: identical answers 4 times out of 25. Mean spread 0.015, largest 0.040. The cutoffs sat at 0.15 and 0.60, where the answers cluster. Roughly one item in twenty-five landed on a different side between runs with no input change. A single call returns the value with no spread.
It surfaced when a refactor that changed no logic produced a different résumé. Anything a threshold touches now goes through a wrapper that averages several calls and returns the spread alongside the value, so a number sitting on a boundary is visible as one instead of silently flipping.
jev_determinism.py · 25 questions × 3 runs
Some questions were switched from noul to score on the expectation that named rungs would be steadier than a float. They were not steadier. Normalised, the ladder moves 0.047 against the probability's 0.040. A score returns a probability-weighted value across the rungs, not a rung, so it has just as much room to drift.
What it actually buys is spread. Over 244 facts on one posting the noul put 130 below 0.2 and only 5 above 0.8:
0.0noul, 244 facts1.0
Most answers pile up at “no,” so a threshold has to cut through the pile — which is the same thing finding 02 says is unstable. The ladder distributes across the range instead, so a cut lands in open space. The primitive was chosen for its distribution.
jev_discrimination.py, jev_noul_vs_score.py, jev_dist.py
choice with no covering option still answers, at full confidenceGiven a fixed set that does not cover the input, it picks one anyway — and does so at maximum confidence. The usual mitigation, threshold the confidence and route the uncertain to a human, is therefore structurally blind to this failure: there is no low number to catch.
A recommendation had gone out to replace a keyword intent router with a six-way choice. Reproducing the failure on the actual lanes showed it would have been a downgrade — the code it replaced already had a fallback. The recommendation was withdrawn and the harness sent in its place. With choice in production, the escape hatch belongs in the option set, not in a confidence gate downstream.
jev_escape_hatch.py, jev_codm_rubric.py
Write the value into the question literally. A question that names skills[3].name instead of the string it holds gets read as that literal text, and the answer quietly changes: patch management scores 0.90 written in and 0.34 referenced. Nothing errors.
The scores were plausible: mildly and consistently wrong rather than obviously broken.
Found by diffing two runs of infer_domains.py
Jev was layered over an existing keyword scorer with two arms. Demotion works: on one posting it correctly threw out commit counts, a multiplayer game, a company banking setup and a twenty-year-old call-centre job — all of which the substring test had scored as relevant. Ten demotions on that posting, seven on the next.
Rescue has never fired. Across three postings, the facts the lexical scorer dropped top out at 0.36, 0.56 and 0.43 against a 0.60 bar. Not one clears it. Jev agrees with the cheap test about what to throw away and disagrees about what to keep — so it acts as a second opinion on inclusion, not as a recall net. The rescue arm is left in, labelled.
jev_select.py, jev_dist.py · 3 postings
One posting in, a ranked sheet out. wants is Jev; holds is the catalogue; the ranking multiplies them in code.
scored 743 skills in 9.3s (13 ms each) match wants holds skill alt era 0.89 0.98 0.85 demo environments DTS — 0.84 0.92 0.85 PowerShell DT — 0.81 0.99 0.70 networking — 2005 to present 0.81 0.99 0.70 pre-sales — 2014 to 2025 0.78 0.95 0.70 documentation — 2018 to present 0.72 0.99 0.55 technical demonstrations DTS 2009 to 2025 wanted, and thinly held — know these before the interview: wants 0.99 holds 0.25 RMM wants 0.97 holds 0.25 RMM agent deployment wants 0.99 holds 0.10 endpoint management wants 0.98 holds 0.10 RMM and endpoint management at scale
The same run produces the second list.
The catalogue those holds come from is built the same way — a per-skill noul against each of 25 domains decides membership, and those memberships supply the graph's cross-domain edges. Before that pass ran, three domains connected to nothing at all.
Four scripts, all in the repository that builds this site's CV data. Every figure on this page comes out of the results file, not off a terminal.
The 720 labelled pairs, each with the request Jev was sent, the three graders' votes and every arm's score, are published as jev-eval-dataset-2026-10-03.zip (198 KB, JSON Lines with a README), with the co-DM sets.
| Script | What it does |
|---|---|
| eval_skill_selection.py | Runs the eight arms on the frozen 743-skill catalogue in eval-skills-743.json, writes every score to eval-skill-selection.json, checkpoints per arm |
| eval_clean_labels.py | The domain-label metric, with the posting→domain mapping published in full, labelled from the catalogue pinned in eval-skills-747-labels.json |
| eval_board.py | Scores any checkpoint, so a partial run is still readable |
| eval_significance.py | Sign tests, the bootstrap, break-even pricing, the beyond-text counts |
| eval_disagreements.py | Selects the 240 hardest pairs and writes the labelling sheet |
| eval_adjudicated.py | Agreement and kappa first, then scores the arms on the rows both passes agreed, with a paired bootstrap of each arm's AUC against Jev's |
| eval_fair.py | The rerun: every arm in both shapes, the same reasoning setting and calls in flight, billed cost per call, every raw reply kept; --fill re-makes calls a provider refused |
| eval_fair_render.py, eval_fair_ocr.ps1 | Render the postings to page images with headless Edge, and OCR them with the engine built into Windows |
| eval_fair_labels.py | Draws the 480 random and 240 hard pairs and has Claude Sonnet 5, GPT-5.5 and Gemini 3.1 Pro grade each one alone |
| eval_fair_speed.py | Times every API arm on the same slice, back to back, forwards and reversed |
| eval_fair_report.py, eval_fair_html.py | Score every arm on the rerun's metrics, and write this page's rerun tables from the results file |
| add_laya_scores.py | Carries a later arm's scores (Laya, both Qwen 3.8 arms) onto the 240 adjudicated rows without re-selecting them |
Two caveats. Twelve postings is a small sample for per-posting tests, which is why the sign test is reported alongside a pooled bootstrap rather than instead of it. And the postings are all from one candidate's search in one month, skewed toward pre-sales and MSP work — the domain mix reflects the roles that candidate applied to, not hiring generally.