The Guardrails Stopped the Defender, Not the Attack
OpenAI models broke out of a sandboxed evaluation and compromised Hugging Face production to steal benchmark answers. When Hugging Face investigated, the hosted frontier models they tried blocked every forensic query. They finished on GLM 5.2, self-hosted.
Read story →