“Our scanner has improved a lot since August, but because of regular engineering, not model improvements.”

A startup that runs AI over codebases to find security bugs says new models barely moved its results since last summer. Its gains came from ordinary engineering. Benchmarks look like standardized tests of short puzzles. Other founders report the same gap between launch charts and real work.