More decisions right than any
published decision model
Decision Model Benchmark60dB Judge (60db-judge-model-v1), our SystemOne-compatible decision model, answers 214 of JevBench's 231public decisions correctly — head to head with TypeSafe's Jev on the same day and runner, and next to every decision model in JevBench's published v1.4 results.
60dB Judge vs Jev on identical requests
Both systems received the same 231 requests through the same SystemOne-compatible API and were scored by JevBench's own code. Fast mode is the same model answering every question in one pass — callers get it by sending a budget under 3 seconds. Where the two disagree, 60dB Judge is right 17 times and Jev twice (p = 0.001).
| Measure | 60dB Judge | Fast mode | Jev 1.13.0 |
|---|---|---|---|
| Correct231 public decisions | 214 | 201 | 199 |
| Easy tier48 decisions | 48 | 48 | 48 |
| Standard tier72 decisions | 71 | 71 | 71 |
| Hard tier111 decisions | 95 | 82 | 80 |
| Brier scoreProbability error, lower is better | 0.119 | 0.168 | 0.179 |
| Expected calibration errorLower is better | 0.034 | 0.045 | 0.029 |
| Ordinal mean absolute errorScore questions, lower is better | 0.19 | 0.25 | 0.14 |
| Paraphrase consistencySame answer on reworded pairs | 35 / 36 | 35 / 36 | 35 / 36 |
The gap is all in the hard tier: 95 of 111 against Jev's 80. Jev's confidence is slightly better matched to its hit rate and it places Score answers closer to the expected level; 60dB Judge's probabilities are closer to the right answer overall (lower Brier score).
Where the lead comes from
Share correct in each of JevBench's 18 decision families, sorted by 60dB Judge's lead over Jev in percentage points.
Dates and numbers
+6
Long policies
+3
Hard judging
+3
Probability
+2
- Dates and numbersn=15+6
- Long policiesn=19+3
- Probabilityn=10+2
- Hard judgingn=17+3
- Ambiguous requestsn=7+1
- Policyn=12+1
- Adversarialn=6tie
- Extractionn=24tie
- Fact checksn=12tie
- Hard routingn=5tie
- Intentn=24tie
- Multi-step reasoningn=18tie
- Ordinal scoresn=12tie
- Routingn=12tie
- Tool selectionn=12tie
- Trapsn=8tie
- Trade-offsn=6tie
- Answer adequacyn=12-1
Lead = decisions 60dB Judge got right minus Jev's, per family. Hover or focus a row for counts.
Next to JevBench's published v1.4 results
JevBench v1.4.0 (23 September 2026) reports each system's accuracy on the same 231 public decisions. Below are the twelve top-ranked decision models as published, unchanged, with 60dB Judge's own measurement added and marked. No published decision model scores above 200 on these items.
| Official rank | System | Public correct | JevBench score | Sealed accuracy |
|---|---|---|---|---|
| — | 60dB JudgeMeasured by 60db, 30 Sep 2026 · not a JevBench result | 214 | not measured | not measured |
| #1 | Jev 1.13.0 (TypeSafe AI) | 200 | 63.29 | 36.7% |
| #4 | Winnow-12B Q8 | 198 | 55.58 | 33.1% |
| #2 | JevK5 v0.2.0 | 197 | 62.04 | 33.1% |
| #6 | djev (Maisa, diffusion-gemma) | 194 | 52.23 | 29.9% |
| #3 | Hopper | 190 | 59.43 | 34.1% |
| #8 | SemIf, formerly OpenJev (Qwen3.5-4B) | 187 | 47.69 | 26.3% |
| #9 | Jobe Qwen3.5-4B (frozen) | 187 | 46.94 | 25.6% |
| #10 | local-jev Qwen3.5-4B | 186 | 46.80 | 26.0% |
| #12 | jqv (Qwen3-32B zero-shot) | 185 | 44.35 | 28.2% |
| #7 | metask-jev-4b | 184 | 47.78 | 27.6% |
| #5 | reflex 4B (kshetrajna12) | 183 | 53.99 | 28.2% |
| #11 | system-one-open (Gemma 4 E2B LoRA) | 169 | 45.11 | 27.6% |
- The official JevBench score combines Intelligence (including 308 sealed decisions), Calibration, Speed and Cost. 60dB Judge has no official score or rank yet — those need JevBench's own sealed-set run.
- Jev's published public result (200) matches our own run of Jev on 29 September (199), so the two measurements line up.
What this report does and does not show
Shows
- Accuracy, calibration and consistency on JevBench's 231 public decisions, scored with JevBench's own code.
- A same-day, same-runner comparison with Jev 1.13.0, with a paired significance test (p = 0.001).
- Where that public-set accuracy sits against JevBench's published v1.4 results.
Does not show
- An official JevBench score or rank. That requires a submission to JevBench.
- Speed comparable to the leaderboard. Our client ran on the host serving 60dB Judge, so timings leave out the internet round trip that JevBench's include.
- Accuracy on JevBench's 308 sealed decisions, which we have not seen.
How the numbers were produced
- Benchmark: JevBench (MIT) at commit 2fa63fa — public items easy, original and hard (dataset hash dc3995d8…), adapter typesafe, scored with JevBench's summarize(). A failed answer counts as wrong; there were none.
- 60dB Judge: 60db-judge-model-v1, called serially over HTTPS on 30 September 2026. Its configuration was selected on separate development data before this single public-set run. Fast mode is the same model's one-pass readout (budget under 3 seconds), run on 29 September.
- Jev: jev-latest on api.typesafe.ai (answered as jev-1.13.0), called serially over HTTPS on 29 September 2026.
- Published results: results/v1.4/jevbench-v1.4-results.json in the JevBench repository (v1.4.0, 23 September 2026). Public correct counts are published public accuracies × 231. General-purpose LLMs in JevBench's table are not decision models and are left out.
- Significance: exact two-sided sign test on the items where the two systems disagree — 60dB Judge right 17 times, Jev twice.
Benchmark source: JevBench on GitHub
