Hillclimbing a smart router on agent workloads
EVO found a setup that beats every single model we measured on ITSMBench, at 21× lower cost than the nearest one.
what is ITSMBench (by Vibrant Labs)
ITSMBench (by Vibrant Labs) measures agent work: the agent runs an IT service desk. A person asks it to do something, such as resolve an incident, raise a problem record, or notify the caller. The agent asks questions, searches the database, and writes changes back.
ITSMBench evaluates agents in multi-turn IT environments on tasks associated with L1/L2 engineers, managers, and admins in real-world settings. The environment includes a policy file, 93 tools, 900+ interconnected records across 16 entity types, a few hundred unique entities, and a user simulator. Rewards are based on the state of the final database at the end of the task.
what EVO does
EVO is a smart router for the workloads you already run. It learns each workload from your traffic and traces, then searches for a cheaper way to serve it — which model, when to escalate, how to structure the call — without dropping below the quality you already ship.
It starts as a drop-in endpoint: same output as the model you name today. Traffic only moves after an offline search clears your bar on held-out requests. Full product writeup.
results
71.4% at $0.0224 a task.
Best single model (Claude Opus 4.8): 69.3% at $0.4703.
Same harness. ~21× lower cost. Higher score.
Every model we measured on ITSMBench, and the router EVO found. The dashed line is the frontier: nothing below and to the right of it is worth buying. The router ends the frontier, so no model is both cheaper and more accurate.
| setup | accuracy | cost / task | note |
|---|---|---|---|
| EVO router | 71.4% | $0.0224 | no frontier model |
| Claude Opus 4.8 | 69.3% | $0.4703 | best single model |
| Grok 4.6 | 66.4% | $0.2743 | |
| DeepSeek V4 Pro | 64.3% | $0.0606* | new peak list |
getting it right once, and getting it right every time
Accuracy averages over four tries at each task. It hides the difference between a setup that solves a task reliably and one that solves it sometimes.
pass@4 is the share of tasks a setup got right at least once in four tries. pass^4 is the share it got right all four times. A service desk that closes an incident correctly one time in four is not usable, so pass^4 is the number that decides whether you would deploy something.
| setup | cost / task | accuracy | pass@4 | pass^4 |
|---|---|---|---|---|
| gpt-5.6 luna | $0.0047 | 19.3% | 60.0% | 5.7% |
| gpt-oss 120b | $0.0109 | 41.4% | 71.4% | 5.7% |
| deepseek v4 flash | $0.0114 | 55.7% | 74.3% | 28.6% |
| gpt-5.4 nano | $0.0170 | 16.4% | 51.4% | 2.9% |
| EVO router | $0.0224 | 71.4% | 88.6% | 45.7% |
| gemma 4 31b | $0.0352 | 47.1% | 80.0% | 17.1% |
| qwen3.7 plus | $0.0408 | 64.3% | 91.4% | 28.6% |
| deepseek v4 pro | $0.0606* | 64.3% | 94.3% | 31.4% |
| glm 5.2 | $0.1033 | 50.7% | 82.9% | 22.9% |
| gemini 3.5 flash | $0.1198 | 55.0% | 80.0% | 28.6% |
| qwen3.7 max | $0.2009 | 56.4% | 82.9% | 20.0% |
| kimi k3 | $0.2619 | 51.4% | 82.9% | 25.7% |
| grok 4.6 | $0.2743 | 66.4% | 91.4% | 42.9% |
| claude opus 4.8 | $0.4703 | 69.3% | 88.6% | 42.9% |
Four trials per task, same harness and same turn budget for every row.
The router has the highest pass^4 we measured: 45.7%, against 42.9% for Opus and Grok. It is not top on pass@4 — DeepSeek V4 Pro hits 94.3%. So this is not “find answers nobody else can find.” It’s “land the same correct final state more often.”
Several models clear about four fifths of the tasks at least once and only about a quarter every time. The capability is there. The consistency is not. Last-mile tuning on the grade is what closes some of that gap.
why this matters
Even the best frontier model in our sweep is not something you’d put on a desk queue if you need the correct final state every time. Reliability is the last mile — and that mile is specific to your grade and your traces, not a public leaderboard.
We hill-climbed that last mile in public on ITSMBench. On your stack it’s the same loop: learn the workload, search the setup, move traffic only when the number you care about moves.
get started
Request access. Early users have gotten 30–60% in cost savings.
ITSMBench is by Vibrant Labs (repo). * DeepSeek V4 Pro cost uses DeepSeek’s new peak/off-peak list (peak shown; off-peak is about $0.030).