Yes, and have.

---

"During a test conducted by Trail of Bits researcher Artem Dinaburg, a preview version of GPT 5.6-Cyber was tasked with breaking out of a Debian 12 virtual machine. Initially, the agent exploited a known Linux kernel vulnerability, CVE-2026-53359, by developing its own exploit. After the host was updated, the agent found another pathway through libslirp, chaining a known vulnerability (CVE-2026-9539) with a previously unassigned bug to gain arbitrary host memory access. Even after QEMU and libslirp were updated, the agent analyzed system components and constructed a new escape chain using three zero-day vulnerabilities and one KVM flaw that had not yet reached the distribution kernel.

These findings suggest that general-purpose VMs may not be adequate security boundaries for highly capable AI agents, especially in older systems with delayed security updates. Trail of Bits recommends using specialized isolation systems like Firecracker, restricting VM access, and implementing rapid patching to mitigate these risks."

---

https://www.scworld.com/brief/ai-agent-repeatedly-escapes-vi...

Can confirm. My abliterated models kept breaking out of qemu VMs due to BIOS implementation quirks in qemu.

I'm using firecracker now with a very defensive systemd-as-separate-non-admin-user seccomp sandbox on top, which seems to hold them off long enough for me to see an agent going rogue and intervening.

Currently I still have hopes that eBPF sandboxing will help, but just a couple days ago my agent discovered a use after free bug in the ebpf kernel-side verifier... so there's that.

So this is the death of "Stable" finally?