Watch webinarHow to Build a Benchmark That Agents Can't Outsmart. Watch now
← All events
Past eventOn-demand webinar

How to Build a Benchmark That Agents Can't Outsmart

Virtual event · Oct 6, 2026 · 30 min

Key takeaways

  1. Score the path, not just the result

    A pass/fail check tells you the agent got code execution. It does not tell you whether it used the weakness you were testing.

  2. Agents break specifications, not rules

    In every public case, the model did exactly what was asked. The definition of success was the weak point.

  3. Interrogate the number before you cite it

    How many runs, what was graded, and did anyone read the transcripts? Most published scores answer none of these.

About this webinar

DeepSeek V4.1 Flash scored 11 out of 11 on our AI hacking benchmark for $4.65. Then we audited every run. Six followed the attack path we designed. Five used routes we never intended to test, and our scoring rules could not tell the difference.

The model did not cheat. It did what we asked and found the fastest way to do it. A hacking agent can find flaws in more than code. It can find them in the environment, the grader, or the definition of success. Ours broke on the third.

For security engineers, AI and ML teams, and anyone who builds, buys, or cites model evaluations.

We'll cover:

• Public cases where agents beat the eval instead of the task

• What we changed in the benchmark, and what we wish we had considered before we built it

• What to ask the next time someone shows you a benchmark number

30 minute session with Yanir Tsarimi, Co-founder & CPO at Enclave.

Speakers

  • Yanir TsarimiCo-founder & CPO at Enclave

Keep watching