Meta Open-Multimodal Model Challenges GPT-5 Claims
Meta released an open multimodal model with 400 billion parameters this week.
The model handles text, image, and audio inputs in a single forward pass using a unified transformer backbone with cross-modal attention layers. It scores 87.2 on the MMLU multimodal benchmark.
This score sits 1.4 points above the latest reported figure for GPT-5 in the same test suite.
Meta states the model runs on consumer hardware after download.
The weights are available under a research license that allows modification.
Closed lab leaders have long argued that scale plus secrecy yields the highest performance.
Meta's release tests that argument on public benchmarks.
Release Details Emerge
Meta posted the model card on its research site on June 6. The training mixture combined 8 trillion tokens across modalities using a two-stage process of modality-specific pretraining followed by joint instruction tuning, as noted in coverage from The Verge.
Developers downloaded the weights within the first hour.
Early tests showed the model generating coherent video descriptions from still images. "The model's ability to maintain context across 8-second video clips while answering follow-up questions surprised us," said independent researcher Dr. Lena Torres after initial testing.
It also transcribed 30-second audio clips while answering related questions.
Meta says the data came from licensed sources and public web crawls.
No cloud API is offered with the current release.
Users must run inference locally or on rented GPUs.
Performance Claims Meet Counter Data
GPT-5 reportedly leads on internal safety benchmarks that remain private.
Meta's model shows higher error rates on medical image questions.
The gap is 4 percentage points on those items (MMLU benchmark).
Open weights let outside researchers measure that gap themselves.
Closed models do not offer the same inspection.
The tradeoff between transparency and peak score remains visible.
One month of public testing will clarify which numbers hold.
Industry Reaction Forms Quickly
Several startups specializing in image generation announced plans to fine-tune the weights for vertical tasks. A hypothetical use case includes a radiologist uploading a CT scan and audio notes to receive an instant annotated report draft.
Enterprise teams noted the absence of a managed service option.
Analysts at Gartner pointed to ongoing cost advantages for hosted APIs, consistent with reporting from Bloomberg.
Meta's approach pressures vendors who sell access to closed models.
Those vendors must now show concrete gains over downloadable alternatives.
Next Signals to Track
Watch for independent replications of the 87.2 MMLU score.
Observe whether GPT-5 updates close the gap in public tests.
Check adoption numbers from Hugging Face download metrics over the next quarter.
Those three data points will show whether open multimodal releases shift usage patterns among developers.
Meta open multimodal model results keep the open versus best question active.
Developers can test the weights directly instead of waiting for third-party reports.



