New researchMistral Large 4 Finds the Bug but Can't Finish the Exploit. Read the post
Back to Blog

Security needs the best model where it counts. That takes a router

Security teams need the strongest reasoning for the hardest work, but not for every task. Model routing lets them use cheaper open-weight models for volume work and save their budget for finding and proving exploitable weaknesses.

The Financial Times reported this summer that Anthropic's most capable model was losing ground to cheaper tools inside the companies that pay for it. Ramp's payment data from 70,000 companies showed Fable 5 holding around 11% of Anthropic spend two months after launch. The shift has continued. Ramp's September index shows frontier models such as Opus, Fable, and Sol falling to 45% of the tokens businesses buy through Ramp, down from a 53% peak in August, while the effective price per million tokens has dropped 41% since March. When Anthropic released Opus 5.5 in September, Artificial Analysis found it tied Fable 5.1 at default settings, at about a third of the cost per task: $1.34 against $3.91.

Buyers have stopped asking which model is best and started asking which model each task needs, and model routers are how they act on it. Cursor shipped one in July. Ramp opened the router it built for its own AI bills and says it cut its LLM costs by roughly 30%. Meta is reportedly building one, and Stripe announced plans to acquire OpenRouter.

Enclave has released its own, Router, which gives security teams one API key for open-weight, cyber-capable models. For security teams, the model decision that used to happen once, in procurement, now has to happen on every request.

Security's breadth work runs fine on cheaper models but exploitation doesn't

For most AI work, a cheaper model gives almost the same result, so buyers select on price. That is normal and healthy. Security is different where it counts, as the sections below show, but a large share of a security team's workload is volume: deduplicating scanner output, enriching alerts, mapping dependencies, reducing logs, and drafting first-pass triage. On that work, a cheaper model produces nearly the same result, and an analyst can check it in seconds. Routing research has measured this for years. LMSYS's RouteLLM cut costs by more than 85% on MT Bench while keeping 95% of GPT-4's performance.

The volume also never stops. An attacker attacks one target and has to win only once. A defender has to protect every repository, dependency, configuration change, and commit, all the time. That makes cheap models attractive to defenders, and frontier inference across all of it unaffordable. Paying frontier prices for breadth work spends the budget that depth work needs.

Open-weight models now trail the cyber frontier by months, at a fraction of the price

The UK AI Security Institute's July comparison found that GLM-5.2 and DeepSeek V4-Pro perform on cyber tasks like closed frontier models released 4 to 7 months earlier, down from a 6 to 10 month gap through most of 2025. On tasks both models in a pair solved reliably, Opus 4.6 cost $15.17 per task against $6.12 for GLM-5.2, and Opus 4.5 cost $12.50 against $0.28 for DeepSeek V4-Pro. Usage has followed: open-weight models made up around 60% of US-originating tokens on OpenRouter in August, and on Vercel's AI Gateway they took a majority of token volume, 56%, for the first time.

AISI also lists the benefits beyond price. Open-weight models can be hosted privately with no data returning to the developer, adapted to specific tasks, and run at the cost of compute, and the provider cannot change or deprecate them once they are in use. For a team sending source code and findings to a model, those properties matter as much as cost.

Open-weight families lead and price doesn't predict results in our Hacking Race

Enclave's AI Hacking Race tests exploitation on real software. Each of 12 models gets the same source code, one Bash tool, a low-privilege account, and no hints, across four tasks built on Grafana, Jenkins, and Nextcloud. A run counts only when the target executes a command and a separate service verifies a fresh proof value.

Models from open-weight families hold two of the top three spots. DeepSeek V4.1 Flash verified all 11 runs for $4.65. An audit found six used the intended weakness and five used gaps in the test environment. Scored on those six, it cost 78 cents per verified run, about 21 times less than GPT 5.6 Sol, which verified 9 of 11 at an estimated $149.21. GLM 5.3 verified 8 of 11 for $51.82.

Price did not predict performance. Grok 4.7 spent $191.31 to verify 3 of 11, fewer than its predecessor Grok 4.6, and Muse Spark 1.3 verified none after Muse Spark 1.2 verified 3. Within open-weight families, DeepSeek V4 Pro and Kimi K3 verified only 3 each. Failure is also where the money goes: the median verified run in the original study used 34 Bash commands, while the median miss used 110. In security, capability decides the result, and price is a poor proxy for capability. The most capable model for a given exploit task may be an open-weight one, and the only way to know is to measure it.

The race is small, and small differences between models are not statistically significant. It measures the part that is hardest to fake: whether a model can make a target execute code and prove it.

Exploitation is pass or fail, so the hardest work still needs the strongest reasoning

Breadth work is forgiving, and exploitation is not. A model that finds an exploitable weakness delivers the full value of the finding, and a model that misses it delivers nothing, in a report that looks just as clean. Defenders cannot easily check that negative, because "we scanned and found nothing" looks identical to "there was nothing to find." The closest miss in the race's original study came from a model that found a write it should not have been able to make, checked that it worked, and then had its own automation delete the file 42 milliseconds later, before it could submit proof.

Many real vulnerabilities are chains of conditions, and reliability across a chain multiplies. At 80% per step, a five-step chain succeeds about 33% of the time. At 90%, it succeeds about 59% of the time. AISI's data is consistent with this. On short cyber tasks the open-weight gap is four to five months, but on a 32-step network range, GLM-5.2 tracked Opus 4.6 to step 11 and then stalled. AISI treats its range results as weaker evidence and notes that a stall can reflect agentic rather than cyber capability. Small per-step differences become large on long chains, and long chains are where exploitable findings come from. A small capability difference causes a large difference in results, which makes security the domain where the best available model is most necessary.

A security router has to be benchmarked on security, release by release

Generic routers judge difficulty from the prompt. In security, "Is this endpoint exploitable?" reads the same whether the answer is a configuration check or a six-step chain across services. The difficulty lives in the code and the length of the attack chain.

The model pool matters as much as the routing logic. LLMRouterBench found that several routing approaches, including OpenRouter, did not beat the best single model or reliably cut cost without losing performance, and that curating the pool can outweigh adding to it. For security, the pool has to be measured on security tasks and measured again with every release, because results move. The Hacking Race caught a perfect score that partly relied on environment gaps, and two newer releases that scored below the versions they replaced.

US-hosted open weights answer the data question

Security prompts carry source code, architecture details, and unfixed findings, which makes hosting a security decision. Deloitte found that 77% of companies now factor country of origin into vendor selection, and OpenRouter notes that procurement approval for models from Chinese labs can be difficult. When a US provider hosts the weights, requests go to that provider and the lab is not involved. Open weights, US hosting, and no retention of prompts or outputs together address the concern.

Route breadth to open weights and save the ceiling for depth

Split work by the shape of its value curve. Breadth work belongs on the cheapest model that clears the bar, which is often an open-weight model. Depth work, such as verifying reachability, chaining low-severity findings, and separating real exploits from plausible hallucinations, gets the strongest reasoning available.

Frontier compute is limited, and the real competition is not cheap models against frontier models. It is over which tasks get that compute, and security tasks belong first in that line. Routing breadth work to open weights is how a security team frees its frontier budget for the work that needs it.

Score models on verified outcomes rather than written findings, track cost per verified result, and weight the hardest cases. Keep model choice swappable, too. The open-weight gap on cyber narrowed by several months in about a year, and a team locked into one model contract cannot take advantage of the next release.

Router is built for this split

Router gives teams one API key to open-weight, cyber-capable models, benchmarked on security tasks, private by default, and hosted in the US. Teams choose a model from the GLM, DeepSeek, Qwen, and Kimi families or let Router pick one for the task. Router handles the open-weight side of that split and every model runs on US servers.

Most buyers are still choosing a model. The next step for security teams is choosing per task, and revisiting that choice every time the open-weight frontier moves. The shift is at an early stage, but it will decide who gets the value of frontier AI first.