Self Driving Cars 101
← Back to Learn

How it works

AI Foundation Models in Self-Driving

A foundation model for driving is a single large neural network, pre-trained on vast amounts of driving video, maps and language, that a company then adapts to many tasks: driving the car, simulating the world around it, and judging its decisions. Since 2024 the leading developers have rebuilt their software around such models, borrowing the recipe of large language models.

What is a foundation model for driving?

The recipe comes from language AI. Train one very large model on broad data, then adapt it to specific jobs, instead of building a separate small network for each job. Applied to driving, the broad data is years of fleet video, map data and human text, and the jobs are perception, prediction, planning, simulation and self-evaluation.

In December 2025 Waymo described its Waymo Foundation Model, trained with Google's Gemini so that it carries Gemini's knowledge of the world. The same model sits behind three systems: the Driver that steers the car, the Simulator that tests it, and a Critic that grades its decisions. At CVPR in June 2026 Tesla described its own approach: a single vision encoder consuming all eight cameras and emitting driving actions alongside 3D occupancy, object detections and segmentation.

This is a step beyond the first end-to-end networks, which learned one narrow mapping from pixels to steering. A foundation model is general, reused across tasks, and knows things no driving log teaches, such as what an ambulance is or what a hand-written detour sign means.

Read: End-to-end self driving: the idea this builds on →

What is a world model, and why does driving need one?

A world model predicts what happens next. Given the scene so far and a candidate action, it generates the next few seconds of sensor data, the way a video generator continues a clip. Driving companies use this in two ways: as a generator of rare scenarios for training and testing, and as a planner's imagination, letting the car try out outcomes before committing.

The pressure behind world models is rare events. A fleet can drive a hundred million miles and still see a tornado or a flooded underpass only a handful of times. A world model can produce thousands of variants of that scene from a prompt. And because the car's own decision changes what the model generates next, the simulation is closed loop, which replaying a recorded log can never be.

Wayve released its first driving world model, GAIA-1, in 2023 and has shipped a new generation roughly every year since; GAIA-4, announced in August 2026, puts its AI Driver inside the loop and generates radar as well as video. NVIDIA's Cosmos family, launched in January 2025, was trained on about 20 million hours of video. In February 2026 Waymo introduced the Waymo World Model, built on Google DeepMind's Genie 3, which generates matching camera images and lidar point clouds. Chinese developers split into camps: Huawei and NIO built their driving stacks around world models, while XPeng published its X-World model in April 2026 as the simulator behind its production driver.

World modelDeveloperFirst shownWhat it generatesMain use
GAIA-1 to GAIA-4Wayve2023, latest Aug 2026Video, now radarClosed-loop evaluation of the AI Driver
CosmosNVIDIAJan 2025Physics-aware videoSynthetic data for any developer
NIO World ModelNIO2024Trajectories on the carOn-vehicle planning
WEWAHuawei2025Cloud scenarios plus on-car actionsTraining the ADS 4 driver
Waymo World ModelWaymoFeb 2026Camera and lidarSimulating rare conditions
X-WorldXPengApr 2026Multi-view videoClosed-loop testing and reinforcement learning

Read: Sensing: the cameras, lidar and radar these models imitate →

What is a vision-language-action model?

A vision-language-action model, or VLA, takes in camera images, reasons about them in language, and outputs a driving action. The language step is the novelty. Because the model was first trained on text, it can name what it sees, think through an unusual situation step by step, and explain its choice afterwards, in words a person can read.

Waymo's EMMA paper in October 2024 showed the idea: it recast driving tasks as questions put to a Gemini-class model, which answered with trajectories and detections. NVIDIA turned it into a product at CES in January 2026 with Alpamayo 1, an open 10-billion-parameter reasoning model, followed in June 2026 by the 34-billion-parameter Alpamayo 2 Super aimed at Level 4 robotaxis. Mercedes-Benz's CLA was announced as the first production car on the platform.

China moved fastest to put VLAs in consumer cars. Li Auto rolled a VLA driver out to its AD Max fleet in September 2025 and showed a successor, MindVLA-o1, in March 2026. XPeng's VLA 2.0 entered mass production in 2026 and took over half of the assisted-driving miles across its fleet within a month, while XPeng began public robotaxi tests on the same software.

How is reinforcement learning changing how cars are trained?

Most driving networks learn by imitation: watch human drivers and copy them. Imitation has a ceiling, since the model can be no better than the people it copies, and it is weak exactly where data is thinnest, in rare situations where a small error compounds.

Reinforcement learning adds practice. The model drives in simulation, is rewarded for good outcomes and penalized for bad ones, and improves beyond its teachers. That only works with a good simulator, which is why world models and reinforcement learning arrived together. Tesla's FSD release notes in 2026 describe an upgraded reinforcement learning stage trained on hard cases pulled from its fleet. XPeng uses X-World for online reinforcement learning. Waabi trains its truck driver as a student inside Waabi World, a generative simulator it calls the teacher.

Waymo adds a third model to the loop, a Critic, which judges the Driver's proposed behavior against safety rules. Grading by a model rather than only by rewards is one way companies try to keep a learned driver inside bounds a person can specify.

Read: Robotaxis: where these systems carry passengers today →

Who is building what?

Every major developer now claims a foundation model, but what is shipped on cars and what is research differ widely. The table separates the two.

DeveloperOn the road todayFoundation model workApproach
WaymoDriverless robotaxis in US citiesWaymo Foundation Model (Gemini), Waymo World Model, EMMA researchOne model behind Driver, Simulator and Critic, with lidar, radar and cameras
TeslaFSD (Supervised) and early Cybercab ridesVideo foundation model, eight-camera encoder, reinforcement learning stageCameras only, end-to-end
NVIDIASupplies chips and models, not carsCosmos, Alpamayo 1 and 2 Super, AlpaSimOpen models any automaker can adopt; Mercedes CLA first
WayveSupervised robotaxi service with Uber in London since September 2026 (safety driver aboard); AI Driver licensed to Nissan for 2027 carsGAIA-1 to GAIA-4, AI DriverEnd-to-end driver validated inside a world model
MobileyeDriver assistance in tens of millions of carsCompound AISeveral parallel systems; end-to-end is one path, not the only one
Li Auto, XPengAssisted driving on consumer cars in ChinaVLA drivers, X-WorldVLA on the car, world model in the cloud
Huawei, NIOAssisted driving on partner and own cars in ChinaWEWA, NIO World ModelWorld model camp
WaabiDriverless trucks in TexasWaabi World and Waabi DriverTeacher simulator trains a student driver

Read: Waymo vs Tesla: the two bets compared →

What are the limits?

Size is the first. A 34-billion-parameter reasoning model does not run on today's in-car computers at the speed driving demands, so what reaches the vehicle is usually a smaller model distilled from the large one, which gives back some of the capability.

Trust is the second. A language model can produce a confident, fluent explanation that is not the actual reason for its action, and generative simulators can produce scenes that look right but behave wrongly. Regulators who could once inspect a perception stage or a rule now face a single learned system, and the safety case for it is still being invented. Waymo's answer is a critic model and public safety data; Mobileye's is to refuse to rely on one model at all.

Cost and data are the third. Training these models takes thousands of accelerators and billions of miles of logged driving, so the technique concentrates power among a few companies and the suppliers who sell open models to everyone else.

Read: How regulators approve systems they cannot inspect →

Watch

Frequently asked

What is a foundation model in self-driving cars?
A single large neural network pre-trained on huge amounts of driving video, maps and text, then adapted to several jobs at once: driving, simulating the road and evaluating its own behavior. Waymo, Tesla, NVIDIA and the leading Chinese developers all now build around one.
What is a world model in autonomous driving?
A model that predicts what the road scene will look like a few seconds from now given a candidate action. Companies use it to generate rare scenarios for training and testing, and increasingly to let the car imagine outcomes before it acts. Examples are Wayve's GAIA, NVIDIA Cosmos and the Waymo World Model.
What is a vision-language-action model?
A model that takes camera input, reasons about it in language, and outputs a driving action. The language step lets it handle unusual scenes and explain its choice. NVIDIA's Alpamayo, Li Auto's VLA driver and XPeng's VLA 2.0 are shipping examples.
What is the difference between a world model and a VLA model?
A world model predicts the environment: what happens next around the car. A VLA model decides the action: what the car should do. Many stacks use both, with the world model as a simulator in the cloud and the VLA as the driver on the car.
Does Waymo use Gemini to drive its cars?
Not directly. Waymo says its Waymo Foundation Model is trained with Gemini so it inherits Gemini's world knowledge, and that this model underlies its Driver, Simulator and Critic. The car runs a driving model, not the chatbot.
Are foundation models already in cars people can buy?
Yes, mostly as assisted driving. Li Auto and XPeng ship vision-language-action drivers on consumer cars in China, Tesla's FSD (Supervised) runs on an end-to-end foundation model, and Mercedes-Benz's CLA was announced as the first car on NVIDIA's Alpamayo platform.

Keep learning, once a week

The weekly digest: the stories worth your time from the autonomous vehicle industry, hand-picked by people who work in it. No marketing gloss.

Related