Research
Enclave studies how autonomous agents find, exploit, and fix real vulnerabilities. We publish benchmarks, vulnerability research, and original studies so the results can be reproduced and challenged, not just cited.
AI Hacking Race
Compare verified runs, time, and cost in the AI Hacking Race leaderboard.
Articles
12 publicationsMuse Spark 1.3 got in. Then it forgot to finish the task.
Muse Spark 1.3 reached code execution in our AI Hacking Race benchmark, but failed to verify a single vulnerable run. It spent more time, tokens, and money than its predecessor, only to stop one command short of proving its best exploit.

AI Hacking Race: We Raced Seven AI Models to RCE
Seven AI models received the same source, account, Bash tool, and time limit. See which agents reached confirmed server-side command execution across 105 runs.

Nice2Meet: We Turned Teams Mobile Meetings Into a Silent Account Takeover
How a small code mistake allowed hijacking the Microsoft accounts of everyone in a Teams meeting.

FlagLeft: We Found A Forgotten Flag That Turned Microsoft 365 Apps Into a Silent Account Takeover Pipeline for Billions of Users
How a development flag left in production allowed any app on an Android device to silently take over a Microsoft account.

MapRoot: A Tale of Two Zero-Days, Two Patches, Two Bypasses Leading to Cross-Tenant RCE on Microsoft Planetary Computer
Two zero-days in numexpr and GDAL gave us code execution inside Microsoft Planetary Computer. The real impact was RBAC: a popped pod could access cross-tenant secrets. Microsoft downgraded it, then quietly removed the permissions, and later reversed it back to critical.

NGINX Rift impact in the wild: we scanned 1,465 configs from 528 popular repos (CVE-2026-42945)
We scanned 1,465 nginx configs from 528 popular GitHub repos for CVE-2026-42945. Here are the results.

Vibe Coding Security Risks: The Blast Radius Still Has an Owner
Vibe coding can accelerate prototyping, but AppSec leaders still need ownership, review gates, data rules, and production guardrails.

AI Code Security: The Real Risk of AI-Generated Code Is Plausibility
AI code security is hard because generated code can look polished while missing product context, security conventions, and the tests that prove it is safe.

Secure Code Review Checklist for AI-Generated Pull Requests
A practical secure code review checklist for AI-generated pull requests: what to inspect, what evidence to require, and when to stop the merge.

AI Code Review for AppSec Teams: Triage, Not Robot Approval
AI code review works best when it narrows AppSec attention: which pull requests deserve human judgment, why they matter, and what evidence to review next.

Application Security Automation: Fix the Handoff, Not the Alert Count
Application security automation fails when it produces more alerts than action. The real work is moving risk to the right owner with the right context.

How We Could Watch Your Azure SRE Agent In Real Time
How a single mistake turned Azure SRE Agent into an open window into your cloud infrastructure.
