Request access
Menu
← All posts · Research

Hillclimbing a smart router on agent workloads

EVO found a setup that beats every single model we measured on ITSMBench, at 21× lower cost than the nearest one.

what is ITSMBench (by Vibrant Labs)

ITSMBench (by Vibrant Labs) measures agent work: the agent runs an IT service desk. A person asks it to do something, such as resolve an incident, raise a problem record, or notify the caller. The agent asks questions, searches the database, and writes changes back.

ITSMBench evaluates agents in multi-turn IT environments on tasks associated with L1/L2 engineers, managers, and admins in real-world settings. The environment includes a policy file, 93 tools, 900+ interconnected records across 16 entity types, a few hundred unique entities, and a user simulator. Rewards are based on the state of the final database at the end of the task.

what EVO does

EVO is a smart router for the workloads you already run. It learns each workload from your traffic and traces, then searches for a cheaper way to serve it — which model, when to escalate, how to structure the call — without dropping below the quality you already ship.

It starts as a drop-in endpoint: same output as the model you name today. Traffic only moves after an offline search clears your bar on held-out requests. Full product writeup.

results

71.4% at $0.0224 a task.

Best single model (Claude Opus 4.8): 69.3% at $0.4703.

Same harness. ~21× lower cost. Higher score.

ITSMBench: accuracy against cost per task Cost per task on a log scale against accuracy, for the measured models and the EVO router. The router reaches 71.4 percent at $0.0224 a task. The nearest model above 69 percent costs $0.4703. Grok 4.6 is 66.4 percent at $0.2743. DeepSeek V4 Pro is 64.3 percent at $0.0606 on the new peak first-party list. 20% 30% 40% 50% 60% 70% $.005 $.01 $.03 $.10 $.30 cost per task, log scale accuracy gpt-5.6 luna gpt-oss 120b deepseek v4 flash gpt-5.4 nano deepseek v4 pro* gemma 4 31b qwen3.7 plus glm 5.2 gemini 3.5 flash qwen3.7 max kimi k3 grok 4.6 claude opus 4.8 EVO router · 71.4% · $0.0224 21x cheaper than the nearest model above 69%

Every model we measured on ITSMBench, and the router EVO found. The dashed line is the frontier: nothing below and to the right of it is worth buying. The router ends the frontier, so no model is both cheaper and more accurate.

setupaccuracycost / tasknote
EVO router71.4%$0.0224no frontier model
Claude Opus 4.869.3%$0.4703best single model
Grok 4.666.4%$0.2743
DeepSeek V4 Pro64.3%$0.0606*new peak list

getting it right once, and getting it right every time

Accuracy averages over four tries at each task. It hides the difference between a setup that solves a task reliably and one that solves it sometimes.

pass@4 is the share of tasks a setup got right at least once in four tries. pass^4 is the share it got right all four times. A service desk that closes an incident correctly one time in four is not usable, so pass^4 is the number that decides whether you would deploy something.

setupcost / taskaccuracypass@4pass^4
gpt-5.6 luna$0.004719.3%60.0%5.7%
gpt-oss 120b$0.010941.4%71.4%5.7%
deepseek v4 flash$0.011455.7%74.3%28.6%
gpt-5.4 nano$0.017016.4%51.4%2.9%
EVO router$0.022471.4%88.6%45.7%
gemma 4 31b$0.035247.1%80.0%17.1%
qwen3.7 plus$0.040864.3%91.4%28.6%
deepseek v4 pro$0.0606*64.3%94.3%31.4%
glm 5.2$0.103350.7%82.9%22.9%
gemini 3.5 flash$0.119855.0%80.0%28.6%
qwen3.7 max$0.200956.4%82.9%20.0%
kimi k3$0.261951.4%82.9%25.7%
grok 4.6$0.274366.4%91.4%42.9%
claude opus 4.8$0.470369.3%88.6%42.9%

Four trials per task, same harness and same turn budget for every row.

The router has the highest pass^4 we measured: 45.7%, against 42.9% for Opus and Grok. It is not top on pass@4 — DeepSeek V4 Pro hits 94.3%. So this is not “find answers nobody else can find.” It’s “land the same correct final state more often.”

Several models clear about four fifths of the tasks at least once and only about a quarter every time. The capability is there. The consistency is not. Last-mile tuning on the grade is what closes some of that gap.

why this matters

Even the best frontier model in our sweep is not something you’d put on a desk queue if you need the correct final state every time. Reliability is the last mile — and that mile is specific to your grade and your traces, not a public leaderboard.

We hill-climbed that last mile in public on ITSMBench. On your stack it’s the same loop: learn the workload, search the setup, move traffic only when the number you care about moves.

get started

Request access. Early users have gotten 30–60% in cost savings.

ITSMBench is by Vibrant Labs (repo). * DeepSeek V4 Pro cost uses DeepSeek’s new peak/off-peak list (peak shown; off-peak is about $0.030).