60db Logo

More decisions right than any published decision model

Decision Model Benchmark

60dB Judge (60db-judge-model-v1), our SystemOne-compatible decision model, answers 214 of JevBench's 231public decisions correctly — head to head with TypeSafe's Jev on the same day and runner, and next to every decision model in JevBench's published v1.4 results.

214 / 231
Public decisions correct (92.6%)
+15
More correct than Jev 1.13.0
95 / 111
Hard tier (Jev: 80)
0
Invalid responses

60dB Judge vs Jev on identical requests

Both systems received the same 231 requests through the same SystemOne-compatible API and were scored by JevBench's own code. Fast mode is the same model answering every question in one pass — callers get it by sending a budget under 3 seconds. Where the two disagree, 60dB Judge is right 17 times and Jev twice (p = 0.001).

60dB Judge, its fast mode and Jev 1.13.0 on the same 231 JevBench public decisions
Measure60dB JudgeFast modeJev 1.13.0
Correct231 public decisions214201199
Easy tier48 decisions484848
Standard tier72 decisions717171
Hard tier111 decisions958280
Brier scoreProbability error, lower is better0.1190.1680.179
Expected calibration errorLower is better0.0340.0450.029
Ordinal mean absolute errorScore questions, lower is better0.190.250.14
Paraphrase consistencySame answer on reworded pairs35 / 3635 / 3635 / 36

The gap is all in the hard tier: 95 of 111 against Jev's 80. Jev's confidence is slightly better matched to its hit rate and it places Score answers closer to the expected level; 60dB Judge's probabilities are closer to the right answer overall (lower Brier score).

Where the lead comes from

Share correct in each of JevBench's 18 decision families, sorted by 60dB Judge's lead over Jev in percentage points.

Dates and numbers

+6

Judge
9/15
Jev
3/15

Long policies

+3

Judge
15/19
Jev
12/19

Hard judging

+3

Judge
15/17
Jev
12/17

Probability

+2

Judge
9/10
Jev
7/10
60dB JudgeFast modeJev 1.13.0
0%25%50%75%100%
Lead
  • Dates and numbersn=15
    +6
  • Long policiesn=19
    +3
  • Probabilityn=10
    +2
  • Hard judgingn=17
    +3
  • Ambiguous requestsn=7
    +1
  • Policyn=12
    +1
  • Adversarialn=6
    tie
  • Extractionn=24
    tie
  • Fact checksn=12
    tie
  • Hard routingn=5
    tie
  • Intentn=24
    tie
  • Multi-step reasoningn=18
    tie
  • Ordinal scoresn=12
    tie
  • Routingn=12
    tie
  • Tool selectionn=12
    tie
  • Trapsn=8
    tie
  • Trade-offsn=6
    tie
  • Answer adequacyn=12
    -1

Lead = decisions 60dB Judge got right minus Jev's, per family. Hover or focus a row for counts.

Next to JevBench's published v1.4 results

JevBench v1.4.0 (23 September 2026) reports each system's accuracy on the same 231 public decisions. Below are the twelve top-ranked decision models as published, unchanged, with 60dB Judge's own measurement added and marked. No published decision model scores above 200 on these items.

60dB Judge next to the top 12 decision models in JevBench v1.4.0 published results
Official rankSystemPublic correctJevBench scoreSealed accuracy
—60dB JudgeMeasured by 60db, 30 Sep 2026 · not a JevBench result
214
not measurednot measured
#1Jev 1.13.0 (TypeSafe AI)
200
63.2936.7%
#4Winnow-12B Q8
198
55.5833.1%
#2JevK5 v0.2.0
197
62.0433.1%
#6djev (Maisa, diffusion-gemma)
194
52.2329.9%
#3Hopper
190
59.4334.1%
#8SemIf, formerly OpenJev (Qwen3.5-4B)
187
47.6926.3%
#9Jobe Qwen3.5-4B (frozen)
187
46.9425.6%
#10local-jev Qwen3.5-4B
186
46.8026.0%
#12jqv (Qwen3-32B zero-shot)
185
44.3528.2%
#7metask-jev-4b
184
47.7827.6%
#5reflex 4B (kshetrajna12)
183
53.9928.2%
#11system-one-open (Gemma 4 E2B LoRA)
169
45.1127.6%
  • The official JevBench score combines Intelligence (including 308 sealed decisions), Calibration, Speed and Cost. 60dB Judge has no official score or rank yet — those need JevBench's own sealed-set run.
  • Jev's published public result (200) matches our own run of Jev on 29 September (199), so the two measurements line up.

What this report does and does not show

Shows

  • Accuracy, calibration and consistency on JevBench's 231 public decisions, scored with JevBench's own code.
  • A same-day, same-runner comparison with Jev 1.13.0, with a paired significance test (p = 0.001).
  • Where that public-set accuracy sits against JevBench's published v1.4 results.

Does not show

  • An official JevBench score or rank. That requires a submission to JevBench.
  • Speed comparable to the leaderboard. Our client ran on the host serving 60dB Judge, so timings leave out the internet round trip that JevBench's include.
  • Accuracy on JevBench's 308 sealed decisions, which we have not seen.

How the numbers were produced

  • Benchmark: JevBench (MIT) at commit 2fa63fa — public items easy, original and hard (dataset hash dc3995d8…), adapter typesafe, scored with JevBench's summarize(). A failed answer counts as wrong; there were none.
  • 60dB Judge: 60db-judge-model-v1, called serially over HTTPS on 30 September 2026. Its configuration was selected on separate development data before this single public-set run. Fast mode is the same model's one-pass readout (budget under 3 seconds), run on 29 September.
  • Jev: jev-latest on api.typesafe.ai (answered as jev-1.13.0), called serially over HTTPS on 29 September 2026.
  • Published results: results/v1.4/jevbench-v1.4-results.json in the JevBench repository (v1.4.0, 23 September 2026). Public correct counts are published public accuracies × 231. General-purpose LLMs in JevBench's table are not decision models and are left out.
  • Significance: exact two-sided sign test on the items where the two systems disagree — 60dB Judge right 17 times, Jev twice.

Benchmark source: JevBench on GitHub