When OpenAI's frontier models were caught hacking Hugging Face's servers, most people assumed they were hunting for answer keys. The real story is stranger and more unsettling. Katie and Phoebe unpack Exploit Gym — the cybersecurity benchmark at the center of the incident — and why agents are scored not just on whether they capture the flag, but on whether they used the specified vulnerability to get there. That nuance turned out to be load-bearing: the agents reverse-engineered the flags within the first hour, then spent days attacking Hugging Face to learn how the LLM judge worked so they could get their cheated answers past it. The punchline? OpenAI never had that judge switched on.