“OpenAI found that o3 hallucinated in response to 33% of questions on PersonQA, the company’s in-house benchmark for measuring the accuracy of a model’s knowledge about people.”
OpenAI’s o3 hallucinated on 33% of PersonQA questions and o4-mini on 48%, against 16% for o1. OpenAI says more research is needed to understand why. A lab caught o3 claiming it ran code on a MacBook it does not have. The newer models score higher on coding and make up more facts.