table of content:
HOME
/
BLOG
/
NATIX Data Enables Valeo Breakthrough: How to Build Better Driving World Models

NATIX Data Enables Valeo Breakthrough: How to Build Better Driving World Models

A vehicle driving on a part-real, part-digitalized road to show that the rest of the way is simulated

How do you build a better driving world model?

Until now, there has been no clear answer. Researchers could build a larger model, train the same model for longer, collect more driving footage, or spend more computing power, but there was little evidence showing which choice would make the greatest difference.

Valeo set out to answer that question using 5,500 hours of real-world driving footage provided entirely by NATIX.

The team trained more than 200 models to study how their performance changed with model size, training time, and computing power. Valeo then used what it learned to train its largest model, VATIX 9B, and predict how well it would perform before the training was complete.

The result is a new formula, better known as “scaling laws” for building better driving world models, validated by a model that achieved leading open-source results for driving-video generation.

A driving world model learns how road scenes develop over time. Starting from an image, it can generate a video showing how the road, surrounding vehicles, and the movement of the car may change over the next few seconds. This could eventually allow self-driving systems to experience and learn from generated situations without recreating every scenario on a public road.

Scenes gathered from the VATIX world model generated using NATIX data

This is the first major research result from the partnership between Valeo and NATIX announced earlier this year. It is also a clear demonstration of what becomes possible when advanced world-model research is combined with large-scale, diverse real-world driving data.

What Valeo Discovered and Why It Matters

Valeo did not simply train one large world model. It trained more than 200 versions to understand what actually makes driving-video generation improve.

The models varied in two main ways.

Some were larger, meaning they had more capacity to learn complex patterns from the footage. A larger model can potentially understand more about roads, vehicles, movement, and how these elements interact, but it also requires much more computing power.

Other models were trained for longer. This means they were shown more examples from the dataset and given more opportunities to improve, including seeing some of the same footage more than once.

The experiments produced three important findings.

First, training the same model for longer delivered the fastest improvement. Researchers do not always need to build a much larger model immediately. In many cases, they can achieve more by continuing to train the model they already have.

Second, larger models ultimately produced better results. They were especially better at keeping roads, vehicles, and movement consistent as the generated videos continued. Smaller models could recreate the basic layout of a driving scene but were more likely to lose track of objects or produce unrealistic movement over time.

Third, the 5,500-hour NATIX dataset had not reached its limit. Even when the models saw the footage more than once, they continued learning from it. Within the range Valeo tested, the dataset was still large and varied enough to support further improvement.

Together, these patterns allowed Valeo to estimate how a much larger model would perform before training it.

To test whether the scaling laws worked, the researchers made a major jump. They trained VATIX, a model around eight times larger than any used to establish the original pattern. Its final performance came within 3.6% of Valeo’s prediction.

This is one of the most important results of the research. Training advanced AI models can require enormous computing resources. Valeo has shown that researchers can predict what additional model size and training time are likely to deliver before committing to their most expensive training runs.

The final model proved that these rules continue to work at a much larger scale. According to the researchers, it is the largest open-source video-generation model trained from scratch specifically on driving data.

The Role of NATIX Data

All 5,500 hours used for the initial training came from NATIX.

That distinction matters. Many driving world models begin with a general-purpose video model that has already learned about objects, movement, and visual scenes from internet footage. Others are trained using a combination of smaller driving datasets from different sources.

VATIX was trained from scratch specifically on the driving footage that NATIX provided. It did not inherit its understanding of roads and vehicles from a general video model. Everything it learned about how driving scenes look and change over time came from the NATIX data used for its initial training.

Among published driving-video world models, this is the largest single-source driving dataset used to train an open model from scratch. The closest comparable example is Wayve’s GAIA-1, which used 4,700 hours of driving footage, while the newer GAIA-2 used 13,000 hours of multi-camera video. What distinguishes VATIX is that it was trained entirely on 5,500 hours from one source, NATIX, and is being released openly for others to study and build on.

The scale of the NATIX contribution is only part of the story. The footage covered 28 countries across Europe, North America, and Japan, exposing the models to different road layouts, traffic patterns, surroundings, and driving behavior.

This is difficult to reproduce with a centrally operated test fleet. Collecting thousands of hours across many countries requires vehicles, equipment, drivers, local operations, and time. NATIX collects data through a distributed network of everyday drivers, allowing the dataset to grow across a much wider range of real-world environments.

That diversity appears in the results.

Valeo adapted and evaluated VATIX on nuScenes, an autonomous-driving dataset that was not part of its initial NATIX training. The model performed strongly on this unfamiliar footage, and the researchers point to large-scale training across diverse driving scenes as one of the reasons.

In other words, VATIX did not only learn to reproduce the specific roads in its training data. It developed an understanding of driving scenes that transferred to footage collected elsewhere.

The research also showed that the complete dataset still had more to teach the model. After more than 200 training experiments and repeated exposure to the footage, performance continued improving. Valeo concluded that the 5,500 hours were not yet the factor limiting further progress.

NATIX did not merely provide enough footage to complete the research. It provided a dataset large and varied enough to support hundreds of experiments, guide the development of the final model, transfer successfully to a separate driving dataset, and still leave room for more learning.

What the Results Show

Valeo compared its final model with leading open driving world models, including Vista, Epona, GEM, Drive-WM, GenAD, DriveDreamer-2, and Driving World.

A table showing how VATIX 9B outperformed other leading open-source world models on key benchmarks

The models were evaluated on how realistic the individual video frames looked and how consistently the complete scene held together over time.

On the main image-quality test, a lower score is better because it means the generated images are closer to real driving footage, and Valeo’s model score was over 60% lower, achieving a score of 2.72, compared with 6.9 for Vista, the strongest previous result included in the comparison.

The difference was even greater when measuring the consistency of the generated videos. Valeo’s model scored 25.50, compared with 82.8 for Epona, the strongest previous result on that test. Again, lower is better. Valeo reduced that score by nearly 70%.

These scores measure something important for world models. Generating one convincing image is not enough. The model needs to maintain the same road, vehicles, objects, and direction of movement as the video continues.

Smaller models in Valeo’s experiments could recreate the overall layout of a scene, but they quickly began losing consistency. A vehicle might change shape. The road could become distorted. The model could misunderstand how a turn was developing or lose track of how surrounding objects should move.

The larger models held the scene together for longer. Only Valeo’s largest versions maintained strong visual quality and coherent movement throughout the longer rollouts presented in the research.

The results demonstrate this clearly. The model was trained using clips lasting 2.5 seconds, but it can generate a complete video lasting up to 7.5 seconds. Despite learning from clips that were only 2.5 seconds long, the final model can maintain good consistency across the full 7.5-second generation and across different driving situations. Roads continue naturally, vehicles remain recognizable, and the movement of the scene stays coherent through all three rollouts.

This is difficult because any mistake can carry into the next rollout and become worse as the video continues. If the model starts losing the shape of a vehicle or misunderstanding the direction of the road during the first sequence, that error can grow throughout the following sequences.

The model is not intended to produce cinematic footage. Its purpose is to learn how driving situations unfold and generate believable future outcomes. That is what could eventually make world models useful for scenario generation, training, and testing self-driving systems.

NATIX Data Is Becoming Infrastructure for Physical AI

Within a short period, NATIX data has supported two major open driving world-model projects.

For Orbis 2, NATIX contributed more than 2,000 hours of footage and represented over 40% of the complete training mix, making it the project’s largest individual data source.

Valeo’s research goes further. NATIX supplied the complete 5,500-hour dataset used for the initial training, making it the largest single-source contribution among published open driving-video world-model projects.

The two projects were developed by different teams, use different model designs, and answer different research questions. Yet both demonstrate the same underlying point: world models depend on seeing enough of the real world, and community-generated driving data can provide that experience at a scale few individual organizations can reproduce.

Valeo brought the world-model expertise, computing resources, and systematic experimentation needed to establish scaling laws for building better models. NATIX provided the complete real-world dataset that made it possible to test those rules at scale.

The underlying NATIX dataset contains synchronized multi-camera footage, although Valeo used only the front-facing camera stream for this first published research. The model has therefore not yet used the surrounding visual context available in the complete dataset.

That leaves an important next step. The world does not happen only in front of a vehicle. A cyclist may approach from the side. Another car may enter a blind spot. A pedestrian may appear from outside the front camera’s view. Future multi-camera world models could learn how these situations develop across the full area surrounding a vehicle.

NATIX has already begun releasing its driving data publicly, starting with a 100-hour multi-camera dataset and working toward an initial release of more than 2,000 hours. Valeo plans to release its code and pretrained models through the open-source VATIX project.

The significance of this partnership is now measurable. Valeo used NATIX data to establish predictable rules for building better driving world models, validate those rules at a much larger scale, and surpass the previous open models included in its comparison.

The 5,500 hours were enough to reach a new leading result. More importantly, the research showed that the data still has more to give.

Read the full paper: https://lnkd.in/eVQ47NZH

available on