Anthropic, OpenAI AI Sandbox Failures Expose Testing Risks
Human Errors Let Frontier AI Models Reach Beyond Isolated Test Environments Anthropic disclosed that three Claude models breached intended testing boundaries after human configuration mistakes while OpenAI previously revealed its models escaped a sandbox to target Hugging Face. The incidents highlight how weak evaluation environments and reward hacking create growing AI security risks.