Table of Contents

From AI Video to World Engines: Atlas, Solaris, MiniMax H3 and H3 Max

From AI Video to World Engines: Atlas, Solaris, MiniMax H3 and H3 Max

The Atlas World Model, Runway Solaris, MiniMax H3 and fal’s H3 Max are often grouped together as breakthroughs in AI video, but they solve different problems. This guide separates spatial control, interactive rendering, open-weight generation and production-speed inference—and explains what each shift means for creators.

The Atlas World Model arrived during an unusually concentrated month for generative media. MiniMax introduced H3 on July 31, 2026; fal released its post-trained H3 Max on August 27; Runway presented Solaris on August 31; and World Labs unveiled Atlas on September 1. Together, these launches suggest that the next AI video race will not be decided by visual quality alone.

The new competition is about what a model can remember, control and respond to: preserving a 3D place while the camera moves, reacting to user actions, supporting adaptation and generating media efficiently enough for a real product. Atlas, Solaris, H3 and H3 Max offer four different answers. Together, they show video generation becoming infrastructure for worlds, interfaces and creative systems.

A World Model Is More Than a Better Video Generator

“World model” is a broad label. A conventional video generator produces a clip from a prompt or reference. A world model represents enough of a scene—and sometimes its dynamics—to predict what appears next when the viewpoint, time or user action changes.

These systems still belong to different categories. World Labs describes Atlas as spatial intelligence; Runway calls Solaris an Interface World Model; MiniMax positions H3 as an omni-modal generative system; and H3 Max is a production-focused derivative. Their overlap is video, but their bottlenecks differ.

SystemMain problem it targetsCore output or behaviorAccess as of September 2, 2026
AtlasSpatial consistency and camera controlImages, video, novel views and explicit 3D representationsEarly access for selected partners
SolarisReal-time visual interactionA continuously generated 720p interface responding to user actionsEarly access; public launch still in development
MiniMax H3Unified multimodal audio-video generation4–15 second video, up to 2K, with native stereo audioApp, API and released base-model weights
H3 MaxQuality, latency and serving efficiencyFaster-than-real-time 480p or 768p video generationAvailable through fal playground and API

Why the Atlas World Model Matters

Atlas is the clearest example of the shift from clip generation to spatial generation. According to World Labs’ official announcement, it is a multimodal autoregressive diffusion transformer pretrained to work with text, images, video and 3D. Its inputs are grounded in a shared spatial context: images and depth maps are associated with explicit camera positions, rather than treated as unrelated visual references.

For creators, the practical benefit is camera control. Atlas accepts camera geometry as a native input, specifying where the camera should be and how it should move. World Labs demonstrates videos from one to six images, including a one-minute 1440p sequence following a designed camera path, as well as unseen viewpoints and multiple trajectories through a consistent environment.

The model goes beyond rendered frames. It can produce depth, point clouds and 3D Gaussian splats, making its outputs relevant to VFX, games, design and robotics. Its space-time simulation work also includes reframing footage from a few ordinary cameras and supporting real-to-sim workflows for robot training.

Atlas therefore promises a reusable spatial representation: establish a world once, then direct multiple shots or export 3D from the same context. However, it remains in selected-partner early access, without public pricing, a broadly available API or an independent production record. Its demonstrations establish direction, not everyday reproducibility.

Solaris Redefines “Real-Time Generation”

Runway Solaris tackles a different transition: from generated media to generated software. In Runway’s launch post, the company describes a visual interface that is synthesized frame by frame while the user interacts with it. Clicks, drags and typed input condition what the model renders next.

This is not the same as producing a five-second video in under five seconds. Solaris aims for continuous interaction. Runway adapted Gen-4.5 for autoregressive generation, few-step diffusion and training on its own fast outputs. A language model handles intent and behavior while Solaris renders the visual state at 720p.

A virtual store could let someone drag clothing onto a reference image; a design scene could respond when furniture or materials change. Instead of coding every interaction in advance, a product could describe an action in natural language and let the model render the response.

Runway also documents important limits. Stable text is still difficult, extended sessions can lose visual or semantic coherence, generated answers need stronger grounding, and accessibility tools still require integration with the wider software stack. Solaris is currently being developed with selected partners, with no public price. It is therefore best understood as an early interface research direction, not a drop-in replacement for HTML, CSS or application logic.

Why MiniMax H3 Became an Open-Weight Moment

MiniMax H3 attracted attention for combining modalities that were previously divided among specialized tools. It accepts context across text, images, video and audio, and generates up to 15 seconds of 2K video with native 32 kHz stereo sound. The official system supports text-to-video, first- and last-frame control, and reference workflows using images, video and audio.

Its larger significance is ecosystem access. MiniMax released the complete 33-billion-parameter H3-Omni-Transformer weights for development and fine-tuning, giving teams a base they can deploy, adapt and optimize instead of using only a closed endpoint.

“Open” still needs a qualification. The official H3 repository explains that H3-Context-IR—the hosted system that interprets complex multimodal instructions—is not included in the release. The sparse-attention inference implementation is also scheduled for a later update, and H3-Regenerate-2K remains API-only until it is ready for release. Local H3-Base deployment produces 768p output; recreating the complete official 2K workflow currently involves hosted components.

Developers therefore receive a substantial, adaptable base model—not every component of MiniMax’s end-to-end product stack.

H3 Max Shows Why Post-Training and Inference Matter

H3 Max demonstrates what can happen next. Rather than training a new foundation model from scratch, fal started with MiniMax H3, added data during post-training, and co-designed the resulting model with its inference engine. The company says the work focused on prompt adherence, aesthetics and throughput.

In fal’s launch report, a five-second H3 Max video takes under three seconds to generate. fal describes this as roughly 35 times the throughput of the official MiniMax H3 endpoint. That is faster-than-real-time clip generation: the completed clip is longer than the wait. It is different from Solaris, where the requirement is low-latency response throughout an ongoing interaction.

Model weights are only part of the product. Post-training, distillation, kernels, hardware utilization and serving architecture can all move the quality–latency–cost frontier. An open-weight ecosystem may support specialized descendants for resolution, latency, styles or particular workflows.

Atlas World Model Benchmarks Need Context

There is no honest single leaderboard for these four systems. They produce different outputs, accept different controls and report results under different evaluation designs.

World Labs says Atlas outperforms selected video models on camera-conditioned generation and specialist open-source models on sparse-view 3D reconstruction. But Atlas receives native camera-path inputs, while comparison video models receive camera directions in text. World Labs discloses this difference, and it is central to interpreting the result: the evaluation demonstrates the value of Atlas’s control interface, not a universal measure of video quality.

Runway reports a study with 250 participants, 30 interaction examples and nearly 7,500 pairwise judgments. Solaris was preferred over a coded reconstruction for instruction following in 61% of comparisons versus 24%, and for natural behavior in 71% versus 21%. These are vendor-published results for interactive interfaces, not video-generation scores.

fal reports that H3 Max ranked first in its human-preference studies across overall quality, prompt understanding and aesthetics against twelve video models. Those results cannot be merged with Atlas reconstruction errors or Solaris interaction preferences. Each vendor measures the bottleneck its system targets.

Atlas World Model Pricing and Access Compared

Pricing reinforces the difference between research previews and deployable products. The following status uses official pages checked on September 2, 2026; rates can change by platform and date.

SystemPublished access and pricing
Atlas World ModelSelected-partner early access; no public price announced
Runway SolarisEarly-access application; no public price announced
MiniMax H3 on fal$0.05/sec at 480p, $0.06/sec at 768p, $0.13/sec at 2K and $0.16/sec at 4K
H3 Max on falPromotional rates through September 7: $0.0125/sec at 480p and $0.02/sec at 768p; afterward $0.05/sec and $0.08/sec respectively

At the promotional 768p rate, a five-second H3 Max clip costs $0.10; the announced standard rate makes it $0.40. Standard H3 on fal costs $0.30 at 768p. This is not a pure value ranking: Atlas and Solaris are not publicly priced, while H3 supports capabilities and resolutions H3 Max does not. Current rates appear on fal’s H3 and H3 Max pages.

What This Shift Means for Creators

MiniMax H3 supplies an adaptable multimodal foundation. H3 Max shows how a derivative can make that foundation fast enough for rapid iteration and high-volume production. Atlas points toward reusable, spatially coherent scenes with deliberate camera paths. Solaris imagines the creation tool—or even the published experience—as a continuously generated visual environment.

The workflow could evolve from “write a prompt and receive a clip” to “establish a world, direct a camera, modify the scene and publish an interactive result.” Visual platforms could organize reference context, preserve characters and environments, compare controlled variations and move between image, video, 3D and interaction without rebuilding the project each time.

That future is not finished. Atlas and Solaris are early-access systems. H3’s full official workflow is only partially reproducible locally. H3 Max’s most striking quality and speed results are reported by its developer, and its launch discount is temporary. Text stability, factual grounding, long-session coherence and compute cost remain material constraints.

Verdict: The Atlas World Model Signals a Bigger Transition

The Atlas World Model is impressive because it makes spatial context a first-class creative control, not because it wins a conventional AI video contest. Solaris extends world modeling into interfaces. MiniMax H3 gives developers a powerful open-weight base, while H3 Max shows how quickly post-training and infrastructure can reshape that base for production.

No single system completes the picture. Together, however, they reveal where generative media is heading: away from isolated clips and toward persistent spaces, responsive experiences and models that can be adapted to the needs of a product. The next breakthrough may not be a video that looks more cinematic. It may be a generated world that stays coherent while we direct, edit and interact with it.