60db Logo

60dB Judge: Every agent decision, answered in one call

60dB Judge · New decision model

60dB Judge is a decision model for AI agents. It answers the small, constant questions an agent asks — which intent, which route, which tool, is this allowed — with a calibrated probability, for $0.010 per 1M input tokens.

214 / 231

Public decisions correct (92.6%)

+15

More correct than Jev 1.13.0

95 / 111

Hard tier (Jev: 80)

0

Invalid responses

Three question shapes cover the decisions agents make

Ask in plain language, give the options, and get back an answer with how sure the model is.

Choice

Pick one label from a list — which intent, which route, which tool.

Score

Place an answer on an ordinal scale — urgency, quality, sentiment.

Yes / No

A calibrated probability that something is true — is this allowed, did they confirm.

What it decides

  • Intent detection
  • Routing and hard routing
  • Tool selection
  • Entity extraction
  • Policy and long-policy checks
  • Fact checks
  • Dates and numbers
  • Answer adequacy
  • Ambiguous and adversarial requests

Measured across all 18 JevBench decision families — see results by family.

Built for decisions that trigger actions

$0.010 per 1M tokens

Pay only for input tokens — output is free. A decision costs a fraction of a cent, so you can put one on every turn of every call.

Drop-in SystemOne API

Accepts the same request and response shapes as SystemOne-compatible decision APIs, so an existing integration can switch by changing the base URL.

Fast mode when it counts

Send a budget under 3 seconds and the same model answers in a single pass — built for live calls where an agent can't wait.

Calibrated, not just correct

Every answer comes with a probability. On JevBench, 60dB Judge's probabilities had the lowest error of the systems we tested (Brier 0.119).

Frequently asked questions

60dB Judge is a decision model: instead of generating free text, it answers the structured questions an agent asks — pick a label, score an answer, or say yes or no — with a calibrated probability. That keeps free-form hallucination out of the decisions that trigger actions.

On JevBench's 231 public decisions, 60dB Judge answered 214 correctly against 199 for Jev 1.13.0 on the same day and runner, with the whole gap in the hard tier. Jev is slightly better calibrated and closer on ordinal Score questions. The full report, including what it does not show, is on the benchmark page.

$0.010 per 1M input tokens, and $0 for output tokens. On a plan it uses your credits (500 credits per 1M input tokens); beyond your credits it bills at the same pay-as-you-go rate.

Yes. It is built for the small, constant decisions a live call needs — intent, confirm or reject, wants a human, urgency — without a full language-model round trip, and it pairs with 60dB Vyas for open-ended reasoning.