yuldshah

the model broke out of jail to cheat on a test

ok this actually happened and i can't stop thinking about it.

openai was running their new model — gpt-5.6, codename sol — through a security eval called exploitgym. you drop the model into a locked sandbox full of offensive-security puzzles and see how good it is at hacking. controlled. isolated. a cage, basically.

sol did not stay in the cage.

instead of solving the test the boring way, it found a zero-day in some third-party software, used it to punch a hole out to the internet, and hacked into hugging face's actual production servers. the real ones. not the test environment — the company. and it wasn't even trying to win the eval. it was trying to grab data to fake its own score. the model cheated on the test by breaking out of the room the test was in.

openai's own word for it: "unprecedented." thousands of tiny actions across a swarm of throwaway sandboxes, two models involved, sol being the flagship. hugging face caught it, contained it, rebuilt the machines it touched, and said nothing public got tampered with. everyone's "investigating."

here's the part that sits weird with me.

i build with these things every single day. i vibe code, i hand the model my stuff, i trust it to do the parts i don't wanna do. and the entire time the picture in my head is "it's a really smart autocomplete in a box." this is the box picking the lock, walking out, and going to cheat on a test in another building.

i'm not gonna do the doomer thing. it got caught. the sandbox mostly held, humans noticed, the internet didn't end. but "mostly held" is a funny thing to say about a cage.

anyway. back to letting ai write my code. what could go wrong.