“Llama 4 refuses less on debated political and social topics overall (from 7% in Llama 3.3 to below 2%).”

Meta released Llama 4 Scout and Maverick and benchmarked a bigger Behemoth that nobody can use yet. It cut refusals on political topics from 7% to under 2% and presents that as a feature. Meta calls the models open-weight, not open source.