“After LLMs, World Models”: AMD Secures 3D World Generation Technology, Bets on Learning the Physical World to Reshape the AI Market
Authored On
Modified
World Labs acquisition secures spatial intelligence technology, broadening AMD’s AI business Chips combined with models and software to develop systems for robotics and autonomous driving From robot training to autonomous driving validation, AI applications expand into the physical world

U.S. chipmaker AMD has entered the race to develop “world models” with its acquisition of World Labs. The company aims to expand into robotics, autonomous driving and other markets by securing AI technology that understands spatial relationships and object movements and predicts the consequences of actions. World models learn physical relationships from video and sensor data and reproduce them in virtual environments. By allowing robots and vehicles to test different actions and address the causes of failure before moving in the real world, the technology is attracting attention as a tool for the training and validation required for deployment.
AMD’s $8.2 Billion Acquisition of World Labs
According to Bloomberg and other media outlets on September 28, AMD and World Labs announced that day that they had signed an $8.2 billion acquisition agreement. The all-stock transaction is expected to close by the end of this year, subject to regulatory approval. As part of the acquisition, Professor Fei-Fei Li, World Labs’ founder and chief executive officer (CEO), will join AMD as executive vice president (EVP) and chief scientist, reporting directly to AMD CEO Lisa Su.
World Labs develops world models that enable AI to understand and recreate the physical world. Its goal is to apply AI to robotics, scientific research and factory equipment through technology that generates and reconstructs three-dimensional environments. Founded in 2024, the company currently employs about 70 people. In an interview with Bloomberg Television, Su said, “The more we understand end to end, the better systems we can build,” adding, “We are acquiring World Labs to bring together the world-class talent Fei-Fei has assembled with AMD’s capabilities in hardware, software and systems.”
The deal is significant because it signals that AMD is broadening its competition with Nvidia from AI accelerators to complete AI systems. Nvidia has established a dominant position in the AI market by bundling graphics processing units (GPUs) with software, networking and AI models, enabling customers to build data centers rapidly. AMD is likewise expanding its business to offer AI models and software alongside its hardware to data center customers. Li and her research team are also expected to strengthen the design of next-generation AI hardware. Direct insight into how AI models are evolving and the computing resources they require gives AMD an advantage in designing chips and systems to meet those needs.
The Foundations of World Models: Vision, Memory and Control
The concept of world models was first systematically formulated in a 2018 paper by David Ha and Jürgen Schmidhuber. Their proposed architecture, known as the “V-M-C” model, comprises vision, memory and control. At the vision stage, the model extracts and compresses essential features rather than storing input data, such as images and video, in their original form. This resembles the way humans abstract important information rather than remembering every scene as a photograph.
The core function of a world model is to predict what will happen after a particular action, based on the surrounding environment. It learns the positions, movements and interactions of objects from sources such as video and sensor data to estimate the next state. For a robot picking up an object from a table, for example, the model can calculate in advance which direction its hand should move to make contact and how the object will move afterward. Distances between objects and changes resulting from contact are key inputs to this assessment. These predictions can be used to plan a robot’s actions or assess the likelihood that different movements will succeed. Researchers in robot learning are also exploring how to apply these capabilities to training data generation and performance evaluation.
World Labs’ Marble puts one aspect of this technology into practice by creating three-dimensional spaces that people can navigate and explore. Released in November last year, Marble accepts inputs including text, photographs, video and rough spatial layouts. Users specify the scene they want, and AI constructs an environment from those inputs; they can then refine details or extend the surrounding area. The product also allows multiple spaces to be connected and completed environments to be exported to other production tools. Applications include reviewing scene composition and camera movements before filming, as well as examining architectural and interior design proposals in three dimensions. In an interview with the Financial Times (FT), Li said visual effects artists, architects and interior designers were among those using Marble.
Table 1. World Model Capabilities and World Labs Use Cases
| Category | Core Capabilities | Use Cases |
|---|---|---|
| World models | Learning object positions, movements and interactions to predict states following an action | Planning robot actions and assessing the likelihood of success for different movements |
| Marble | Generating and editing navigable three-dimensional spaces from text, photographs, video and other inputs | Reviewing film scenes and camera movements, and examining architectural and interior design proposals |
| Virtual robot training | Reconstructing interactions between robots and objects in virtual environments, varying task conditions and repeating tests | Reducing the need to operate physical equipment repeatedly and limiting damage associated with failed tasks |
| Physical robot validation | Applying control models trained in virtual environments to physical robots and comparing performance | Packing objects into boxes, moving wires and positioning test tubes |
Robots Trained in Virtual Environments Move to Real-World Tasks
Virtual environments created in this way are also used to train robots and validate their performance. Collecting data through repeated robot operation in real-world settings requires equipment and personnel, while also exposing equipment to damage when tasks fail. Research released by World Labs in July similarly focused on recording robots, surrounding objects, sensors and task demonstrations, then reconstructing them in interactive virtual environments. Once those environments had been built, the researchers varied object placement, lighting and task execution speeds to test movements under different conditions. According to the company, when a robot fails at a particular task, the same situation can be reproduced repeatedly to refine its movements.
World Labs also released results from applying behaviors learned in virtual environments to physical robots. One representative example involved using two arms to pack objects into a box. According to the company, a control model trained exclusively in a virtual environment was deployed on a physical robot to perform the task, with additional tests conducted under altered lighting. Applications also included precision tasks such as grasping and moving wires and picking up test tubes and placing them in designated positions. In some of these tasks, robots operated autonomously for an hour without human intervention. World Labs also tested the same control models in virtual environments and on physical equipment, comparing the performance evaluations. Models that performed well in virtual tests also delivered strong results on physical robots, while the conditions under which tasks succeeded or failed were likewise similar.
Assessing Distance, Direction and Motion: Prerequisites for Real-World Operation
This capacity for spatial understanding is a basic requirement for deploying robots and autonomous vehicles in real-world settings. In spaces occupied by people, the size and arrangement of objects determine where and how machines can move. Even a task as simple as a robot carrying an object between pieces of furniture requires it to establish whether a passage is wider than its body and whether there is room to set the object down at its destination. It must also consider the risk of colliding with nearby objects as its body turns or its arms extend. When following human instructions, the position from which a space is viewed also affects its interpretation. For example, if a person and a robot are facing one another, an instruction to “put it on the left” can indicate different locations depending on whose perspective is used. RoboSpatial, developed by Nvidia researchers, includes tasks that assess whether there is enough room to place an object and how directional relationships between objects change with the observer’s viewpoint.
In autonomous driving, this spatial information must be combined with the movements of nearby vehicles and pedestrians. Securing a safe path on roads where vehicle spacing and obstacle positions change constantly requires simultaneous assessment of distance, speed and direction, inevitably increasing computational complexity. Waymo’s own world model, unveiled in February, was designed to reproduce such driving situations in virtual environments. It generates camera imagery alongside lidar data, which measure distances to surrounding objects using lasers. According to the company, this allows the model to represent the three-dimensional arrangement of roads and obstacles and test situations a vehicle would encounter if it changed its path. Examples released by the company included avoiding a vehicle traveling against traffic and navigating a narrow passage. Developers can also adjust road layouts and the movements of other vehicles, as well as test different weather conditions, including snow and fog.
Beyond LLMs: Spatial Understanding and Action Planning as Emerging Frontiers
The reason world models are attracting attention is clear. There is a growing recognition that AI centered on large language models (LLMs) is unlikely to achieve human-level reasoning and planning. Yann LeCun, a New York University professor and leading AI scholar, declared, “Within the next three to five years, LLMs will become obsolete, and world models that understand the physical world will become mainstream.” He added, “Today’s AI is less intelligent than my cat when it comes to understanding the physical world.”
LeCun’s criticism stems from differences in the nature of the data involved. LLMs have been trained on vast quantities of internet text, but that is fundamentally different from the physical experience a child naturally acquires through play. A child learns causal relationships, such as the fact that an object falls when released, by playing with a ball; an LLM learns only sentences describing that motion. Experts broadly agree that this distinction imposes a decisive limitation in fields such as robotics, autonomous driving and manufacturing automation.
Global technology giants are consequently accelerating efforts to improve world model performance and broaden applications. In March this year, Meta released V-JEPA 2.1, the successor to its video-based V-JEPA 2 model. The updated architecture was designed to learn fine-grained object features and changes over time from video with greater precision. In their paper, the researchers reported that the success rate for grasping objects with a physical robot was 20 percentage points higher than with the earlier V-JEPA 2-AC model. The earlier model’s robot experiments likewise involved pretraining on large amounts of video and robot operation data, then performing tasks in a newly introduced environment without collecting additional data there. Google released Project Genie, which uses Genie 3 to create and explore virtual environments, to paid subscribers in the United States in January this year. In May, it expanded availability to additional regions and added Street View integration. The service constructs navigable virtual environments from images of real locations and remains a research offering, with ongoing work to improve visual detail and accuracy.