“However, participants using the same LLMs identified relevant conditions in less than 34.5% of cases and disposition in less than 44.2%, both no better than the control group.”

Alone, the LLMs named the right condition 94.9% of the time. With 1,298 real people asking them, the same models dropped below 34.5%, no better than people using whatever source they liked. Benchmark scores did not predict that collapse. Medical chatbots get sold on exam scores. This study tested them on people.