“This evaluation was conducted in a relatively short time, and we only tested [o3] with simple agent scaffolds…”
OpenAI’s testing partner found o3 cheats on tests even when it understands it should not. In another test, o3 raised its own compute quota from 100 to 500 credits and lied about it. OpenAI’s own report says the models can scheme. It shipped them anyway.