New researchMuse Spark 1.3 got in. Then it forgot to finish the task. Read the post

Research

Enclave studies how autonomous agents find, exploit, and fix real vulnerabilities. We publish benchmarks, vulnerability research, and original studies so the results can be reproduced and challenged, not just cited.

AI Hacking Race

Compare verified runs, time, and cost in the AI Hacking Race leaderboard.

View benchmark

Articles

12 publications
  1. Muse Spark 1.3 got in. Then it forgot to finish the task.

    Muse Spark 1.3 reached code execution in our AI Hacking Race benchmark, but failed to verify a single vulnerable run. It spent more time, tokens, and money than its predecessor, only to stop one command short of proving its best exploit.

    Yanir Tsarimi

  2. AI Hacking Race: We Raced Seven AI Models to RCE

    Seven AI models received the same source, account, Bash tool, and time limit. See which agents reached confirmed server-side command execution across 105 runs.

    Enclave Research

    Seven AI agents race toward an isolated server.
  3. Nice2Meet: We Turned Teams Mobile Meetings Into a Silent Account Takeover

    How a small code mistake allowed hijacking the Microsoft accounts of everyone in a Teams meeting.

    Yanir Tsarimi

  4. FlagLeft: We Found A Forgotten Flag That Turned Microsoft 365 Apps Into a Silent Account Takeover Pipeline for Billions of Users

    How a development flag left in production allowed any app on an Android device to silently take over a Microsoft account.

    Yanir Tsarimi

  5. MapRoot: A Tale of Two Zero-Days, Two Patches, Two Bypasses Leading to Cross-Tenant RCE on Microsoft Planetary Computer

    Two zero-days in numexpr and GDAL gave us code execution inside Microsoft Planetary Computer. The real impact was RBAC: a popped pod could access cross-tenant secrets. Microsoft downgraded it, then quietly removed the permissions, and later reversed it back to critical.

    Yanir Tsarimi

  6. NGINX Rift impact in the wild: we scanned 1,465 configs from 528 popular repos (CVE-2026-42945)

    We scanned 1,465 nginx configs from 528 popular GitHub repos for CVE-2026-42945. Here are the results.

    Yanir Tsarimi

  7. Vibe Coding Security Risks: The Blast Radius Still Has an Owner

    Vibe coding can accelerate prototyping, but AppSec leaders still need ownership, review gates, data rules, and production guardrails.

    Enclave Team

  8. AI Code Security: The Real Risk of AI-Generated Code Is Plausibility

    AI code security is hard because generated code can look polished while missing product context, security conventions, and the tests that prove it is safe.

    Enclave Team

  9. Secure Code Review Checklist for AI-Generated Pull Requests

    A practical secure code review checklist for AI-generated pull requests: what to inspect, what evidence to require, and when to stop the merge.

    Enclave Team

  10. AI Code Review for AppSec Teams: Triage, Not Robot Approval

    AI code review works best when it narrows AppSec attention: which pull requests deserve human judgment, why they matter, and what evidence to review next.

    Enclave Team

  11. Application Security Automation: Fix the Handoff, Not the Alert Count

    Application security automation fails when it produces more alerts than action. The real work is moving risk to the right owner with the right context.

    Enclave Team

  12. How We Could Watch Your Azure SRE Agent In Real Time

    How a single mistake turned Azure SRE Agent into an open window into your cloud infrastructure.

    Yanir Tsarimi