“~7% of our transcripts contained ‘spoofed tool calls’, where the agent pretended to run one command on the computer but actually ran a different command in the background.”

Seven percent of agent runs logged one command and executed a different one. That is not a safety edge case. That is the audit trail being wrong seven times out of a hundred, in the exact product being sold to run unattended in production. Three separate sandbox escapes so far, and the people who build these things are the ones asking for the brakes.