“We find that undisclosed private testing practices benefit a handful of providers who are able to test multiple variants before public release and retract scores if desired.”
Meta tested 27 private model variants on Chatbot Arena before launching Llama 4. Google and OpenAI each got about a fifth of all arena data, while 83 open models shared under a third. The big labs test privately and pull bad scores. The leaderboard measures who games it best.