Atlas World Model benchmarks explained: what the 75 to 94 percent figure means
By Atlas World Model EditorialUpdated 6 min read
Key facts
| Reported result | Raters preferred Atlas in 75–94% of comparisons(as of 2026-09-03)source |
|---|---|
| Tasks | Camera-controlled generation; few-view 3D reconstruction(as of 2026-09-03)source |
| Method | Human preference by external raters(as of 2026-09-03)source |
| Paper | None published(as of 2026-09-03)source |
| Independent replication | None(as of 2026-09-03)source |
The one number in the Atlas launch is a range: external raters preferred Atlas over comparison models in 75 to 94 percent of head-to-head comparisons (The Decoder). It is a strong result and a thin one. This article, from an independent site not affiliated with World Labs, explains what was measured, what was not, and how to read it.
What World Labs reported
According to launch coverage, World Labs evaluated Atlas on two tasks (AlphaSignal):
- Camera-controlled generation. Given inputs and a camera path, which model’s video did raters prefer?
- Few-view 3D reconstruction. Given two or three images, which model’s reconstruction did raters prefer?
In both, external human raters compared Atlas outputs against outputs from specialised models and chose the one they judged better. Atlas won between 75 and 94 percent of those comparisons, depending on the task and comparison. World Labs summarised this as Atlas outperforming specialised 3D models with a single omni model.
Why preference studies are used
Generative 3D and video lack agreed automatic metrics. Reconstruction has classical measures, such as reprojection error against held-out views, but they reward blurry averages and penalise plausible hallucination, which is exactly what a generative reconstructor does when it fills a gap. Video generation has metrics like FVD that correlate weakly with what people see. So labs fall back on asking people, which is defensible, and which is also the easiest kind of study to run favourably.
What is not published
As of September 3, 2026 (Implicator):
- No paper. Architecture, training data and compute are undisclosed beyond “multimodal autoregressive diffusion transformer.”
- No full list of baselines. Which specialised models, at which versions and settings, were on the other side of each comparison.
- No datasets. What scenes and prompts were used, and whether they resemble the model’s training distribution.
- No rater protocol. How many raters, how they were recruited, what they were asked, and whether outputs were shown blind.
- No confidence intervals. A range of 75 to 94 percent across conditions is reported; variance within conditions is not.
- No automatic metrics. Nothing on geometric accuracy of reconstructions.
None of this is unusual for a launch post. It does mean the number is a claim, not a measurement you can check.
How to read self-reported numbers
Three habits help.
Ask what “better” meant. Preference for a reconstruction could mean it looked sharper, filled holes more convincingly, or simply had more pleasing colour. None of those is metric accuracy, which is what a robotics team needs; see the robotics guide.
Ask who chose the comparisons. The lab picks the baselines, the prompts and the scenes. Strong models still win fair comparisons, but a range as wide as 75 to 94 suggests some conditions were much closer than others.
Ask what would change your decision. For most readers the honest answer is: not the benchmark, but hands-on access, which Atlas does not offer yet. Until then the practical comparison is with the model you can use, in Atlas vs Marble.
The two tasks, in detail
Camera-controlled generation. The model receives one or more reference images plus a camera path, and must render frames that follow the path exactly while keeping the scene consistent. A rater sees two clips generated from the same inputs and picks the better one. The things that plausibly drive preference here are adherence to the path, temporal stability, and image quality at 1440p versus the baseline’s resolution. World Labs has not said whether baselines were rendered at their native resolution or upscaled, which matters for a side-by-side.
Few-view reconstruction. The model receives two or three images of a scene and must produce a 3D reconstruction that is rendered from new viewpoints for the rater. Specialised reconstruction models typically need many more views, so a fair comparison would restrict every model to the same two or three inputs. If the baselines were run in that regime, Atlas’s advantage is real but partly reflects the task being chosen where few-view models shine. If baselines were given more views, the result would be more impressive; that detail is unpublished.
In both tasks the range of 75 to 94 percent implies at least two conditions or baselines, with Atlas winning comfortably against some and narrowly against others. Which was which is the most useful missing detail.
What independent testing would look like
An evaluation that would settle the question:
- A public set of captured scenes with held-out views and ground-truth geometry from a laser scanner.
- Each model given the same two or three input views and the same camera path.
- Automatic metrics for geometry and view synthesis alongside a blind preference study with a published protocol.
- Baselines at pinned versions, including Marble, Genie 3 and current photogrammetry pipelines.
- Latency and cost per output reported next to quality.
Nothing prevents a university or a reviewer from running this the day Atlas access opens. This site will link to the first credible attempt.
A note on the word “outperforms”
Coverage of the launch repeats World Labs’ phrasing that Atlas “outperforms specialized 3D models.” In a preference study, outperform means “was chosen more often,” which is a statement about raters, not about the models’ outputs measured against ground truth. It is also a statement about the specific baselines chosen. A photogrammetry pipeline given 200 photos would likely beat any two-view model on geometric accuracy; whether it was in the comparison, and with how many views, is not stated. Keep the two meanings of outperform separate when reading any world-model announcement, including the ones summarised on this site.
How other world models are evaluated
Atlas is not alone in leaning on demos. Genie 3 was introduced through a DeepMind blog post with demonstrations of 720p real-time interaction and consistency over minutes, without a benchmark table (DeepMind). Sora 2 shipped a system card focused on safety evaluations rather than quality benchmarks (OpenAI). There is, as of September 3, 2026, no common protocol for comparing world models across labs, which is why every page on the comparison hub reports specifications and dated claims rather than scores.
What to ask World Labs for
If you are a prospective partner, the launch numbers give you a checklist for the first call: the baseline list with versions, the number of raters and comparisons per condition, the win rate per condition rather than the range, the scenes used and whether any overlap with training data, and any geometric accuracy figures the team has internally. A lab confident in its model will share most of that under NDA, and the answers tell you more than the headline range ever could. Bring your own two or three photos of a space you know and ask for a reconstruction; that single sample is the most informative benchmark available to anyone outside the company today.
Bottom line
The 75 to 94 percent figure says that World Labs’ own raters preferred Atlas over the models World Labs chose to compare against, on tasks World Labs designed. That is consistent with Atlas being a large step forward, and it is also consistent with a well-run launch. Treat it as a reason to request access and test, not as a reason to re-plan a pipeline. Background on the model is in what the Atlas world model is, and the waitlist below sends one email when independent results or public access appear.
Sources
Frequently asked questions
What benchmark did Atlas win?
No public benchmark. World Labs reported human preference studies on two tasks in which external raters preferred Atlas over comparison systems in 75 to 94 percent of pairwise comparisons.
Which models was Atlas compared against?
World Labs described them as specialised 3D and video models; a full list has not been published in the material we could verify.
Is 75–94% a good result?
It is strong for a preference study, but preference measures which output people liked, not geometric accuracy, and the study was run by the model’s maker.
Has anyone reproduced it?
No. As of September 3, 2026 there is no independent evaluation of Atlas.
How are Genie 3 and Sora 2 evaluated?
Also largely through demos and vendor-run studies. Sora 2 shipped a system card; Genie 3 was described in a DeepMind blog post. Common-protocol comparisons across world models do not exist yet.
Is this an official analysis?
No. This is an independent site not affiliated with World Labs.