Introducing EVO Router
Frontier model performance at a fraction of the cost, on the workloads you already run.
Most of what you send to a frontier model does not need one. Almost nobody checks, because checking means building evals for every workload and testing every alternative against them, and getting it wrong means shipping worse output to users.
So the model gets picked once, at the start, and never revisited. New models ship every month. The prompt gets tuned, the scaffold gets rewritten, and the model stays whatever it was on day one. You keep paying frontier prices for a model that stopped being frontier months ago. Newer ones ship every few weeks that do the job better and cost less.
EVO Router does that checking. It is a drop-in endpoint: point your gateway at it and it starts by serving the model you already use, returning the output you already get. From there it learns each workload from live traffic, searches offline for a more efficient way to serve it, and moves traffic only once that route clears your quality bar on requests it has not seen.
The quality bar is set from your own workloads, not a leaderboard. EVO builds evals out of your production outputs, so clearing the bar means matching the quality your traffic actually got.
what evo optimizes
Routing is one lever. EVO searches six.
- Model and provider. The right model for each kind of request, re-checked as new ones ship.
- Prompts. The same prompt does not behave the same on every model, so EVO rewrites it to fit.
- Fusion and voting. Several small models answering together, with a tiebreak.
- Cascades. A cheap first pass, escalating only when the answer needs it.
- Harness and tools. Tool definitions, planning loops and retries, tuned per model, since a scaffold built for one model rarely suits another.
- Context and caching. What goes in the window, in what order, and what gets reused.
Every candidate is scored offline against your baseline before anything ships, so nothing moves on a vendor claim or a public benchmark number.
frontier model performance at a fraction of the cost
To show what that search produces, we pointed it at a public benchmark. GPQA Diamond is 198 PhD-level science questions, and the models at the top of it cost five to twelve cents each. EVO ran unattended overnight on $47 of API spend.
Accuracy against cost per question, log scale. Field and effort-ladder data: Artificial Analysis, retrieved 3 August 2026, costed from their published token counts at list prices. EVO's point is our own run at OpenRouter billed rates, overlaid on their chart. The router point is that vendor's own published figure under their own harness. In the scorecard below, GPT-5.6 Sol on GPQA Diamond is Artificial Analysis's run at maximum effort; on IFEval both the router and Sol figures are that vendor's published numbers, since no independent IFEval measurement was available. EVO's bars are our own runs throughout.
The staircase is the cost-quality frontier, the cheapest way anyone has found to reach each level of accuracy. EVO sits on it. Of the 61 models in the competitive band, 49 cost more and score no better, and the cheapest one that scores higher costs 5× as much.
The dark line is a single frontier model across its reasoning-effort settings. EVO beats every setting but the most expensive, at a tenth of the price.
Solve rate above, cost per task below, on a shared axis. GPQA Diamond and IFEval are the two benchmarks EVO has been run against end to end. Sources differ by bar and are listed under the frontier chart above.
On GPQA Diamond EVO beats the alternate router on both axes, and the reference model scores a point higher while costing 13 times as much. On IFEval EVO scores above both and costs 8 times less than the reference. That is the shape the search keeps producing: parity or better on the axis you care about, and a different order of magnitude on the one you pay for.
93.18% on GPQA Diamond at $0.00825 a question. The nearest model that scores higher costs 5× more.
the system it found
Two commodity models. No frontier model is called anywhere in the pipeline.
A cheap open-weight model answers three times. If all three agree, which happens on about 91% of questions, that is the answer, and the question cost about two tenths of a cent. If they disagree, one call to a second model settles it. The second model was not chosen for being strong. It was chosen because it comes from a different lab, so its errors are decorrelated from the first model's.
That is the kind of solution EVO's search tends to find: an optimized path for most of the traffic, with spend concentrated on the few requests where that path is uncertain. Here that concentration is steep. The 9% of questions that escalate consume 74% of the total bill. That is what the search arrived at for this workload, with the models available this month. Next month there are new models and the answer changes, which is the reason to have something that keeps checking.
Reported GPQA figures are the average of two scored trials, 91.9% and 94.4%; a single 198-question run carries roughly plus or minus 3.4 points at 95% confidence. This is a cross-harness comparison on the same 198 questions: Artificial Analysis ran theirs under their protocol at maximum-effort settings and list pricing, and ours is our own run overlaid on their chart rather than measured by them. Prices are a snapshot of 3 August 2026.
get started
EVO Router is in private beta. Request access and we will send you an API key and a base URL. It sits next to the gateway you already run, as a provider or a proxy in front of it, and it does not replace it.
Once you have a key, integration is a base URL and a model prefix:
from openai import OpenAI
client = OpenAI( base_url="https://api.evo-hq.com/v1", api_key=os.environ["EVO_API_KEY"],
)
client.chat.completions.create( model="evo-auto/gpt-5.6", messages=[{"role": "user", "content": "..."}],
)
Your SDK, messages and tool definitions stay as they are. The model you name after
evo-auto/ is whatever you run today, and it stays the fallback, so on day one
every request passes straight through to it. Workloads start moving only once EVO has
evidence they should, and you can pin any workload back to the base model at any time.