

A self-driving system can lead every public benchmark and still be nowhere near approved to carry passengers without a safety driver. That gap surprises people outside the industry, and it should. If a model tops the leaderboard on perception accuracy or planning quality, it feels reasonable to assume it is ready for the road. Regulators do not see it that way, and the reason comes down to a simple but easy-to-miss distinction: a benchmark measures performance on a fixed set of test cases someone else picked in advance. Autonomous driving validation data has to prove something much larger: that a system behaves safely across the full range of conditions it will actually meet once it is driving among people.
That distinction is the whole article, so it is worth stating plainly here. A benchmark score is a sample. Regulatory approval asks about a domain. Understanding what closes that gap, and why a good score alone cannot close it, is what it actually takes to validate a self-driving car for regulators.
When a regulatory body evaluates a driving system, it is not grading a model. It is evaluating a documented, evidence-based argument that the system is acceptably safe for a defined set of real-world conditions. This kind of argument has a name in safety engineering: a safety case. Think of it less like a test score and more like a structured brief a lawyer would build, where the client is a piece of software, and the audience is a body deciding whether the public should be exposed to it.
A safety case usually rests on three pieces of evidence. The first is a precise description of where and how the vehicle is meant to operate, covering geography, road type, weather, speed range, and traffic conditions. This is called the Operational Design Domain, or ODD, and every other piece of evidence gets measured against it. The second is a record of what has been tested inside that domain, including the rare and dangerous situations, not just the common ones. The third is a track record of what actually happened once the vehicle drove in the real world, captured through metrics like disengagements, near misses, and incidents. A disengagement is simply a moment when a human safety driver had to take control back from the automated system, and regulators treat these events as some of the most informative data a company can report.

The details of what counts as sufficient evidence vary by jurisdiction, and each regulator sets its own thresholds, but the underlying structure keeps repeating everywhere: a defined domain, a tested scenario set inside it, and a track record of real driving behavior. That is what a safety case really means in practice, and it explains why a single number, however impressive, was never going to be the whole story.
Public perception and planning benchmarks were built mostly from common driving conditions: clear weather, daylight, predictable traffic. They have pushed the field forward in real ways, but they were never designed to represent the situations that actually cause crashes. Erratic pedestrian behavior, debris in the road, sudden weather changes, and unusual lighting are underrepresented by design, because rare events are rare. A system can score well on average case metrics and still fail exactly where it matters most.
This is the same problem a kitchen runs into if it only practices its top five dishes. Everything runs smoothly until a customer asks for something unusual, and then the whole operation falls apart. A benchmark trains and measures a system against the dishes it already knows how to make. A regulator wants to know what happens the first time something unusual walks through the door.
Regulators know this, which is why their own frameworks lean on different measurements than a benchmark leaderboard. Safety validation research for automated driving systems increasingly depends on real crash rates and disengagement counts rather than model accuracy scores, precisely because those numbers reflect what happened, not what a curated test set predicted would happen. California's DMV, for example, requires autonomous vehicle permit holders to file annual disengagement reports summarizing real-world miles driven and every case where a safety driver had to step in. In the most recent reporting period, permit holders logged more than 9 million miles in autonomous mode. That is a small fraction of what a single benchmark dataset contains, but it is a fraction that actually happened on public roads, and that is exactly why regulators weight it so heavily. The real world does not cooperate by scheduling rare events conveniently, which is precisely what makes real-world exposure so valuable as evidence.
Once an ODD is defined, the next step is deriving the scenarios that matter inside it, then testing against them systematically. Some of that testing happens in simulation. Some happens with real vehicles logging real miles under real conditions. The frameworks currently taking shape make this explicit rather than leaving it to individual companies to decide. In early 2026, a United Nations regulatory body finalized a Global Technical Regulation on automated driving systems that formally adopted the safety case approach as the standard way to justify that a system is safe enough for market introduction, rather than relying on a single bright-line performance number.
In the United States, NHTSA's proposed AV STEP framework would formalize data reporting and independent safety assessment for companies seeking exemptions or deployment approval, and a 2026 congressional discussion draft goes further, proposing that manufacturers submit a documented safety case demonstrating their system meets defined safety performance criteria before entering interstate commerce. The common thread across all of these efforts is that none of them ask for a single leaderboard score. They ask for a body of evidence: a defined domain, a tested scenario set inside that domain, and a track record of what happened when the system actually drove.

A widely used standard called SOTIF, short for safety of the intended functionality, gives a useful way to think about why this matters. It sorts every situation a vehicle might face into four categories, based on whether the situation is safe or hazardous, and whether the team even knows about it yet. The easy category is known and safe, ordinary driving under well-understood conditions. The hardest category is unknown and hazardous: situations nobody has thought to test for because nobody has encountered them yet. Simulation and structured test plans are good at finding the known hazardous cases. Broad, diverse real-world exposure is the main tool for surfacing the unknown ones, which is exactly why regulators keep asking for it.
Training benefits from amplification. Validation depends on fidelity. That distinction matters more here than almost anywhere else in autonomous driving. During training, it is genuinely useful to take a rare event and expand it into many synthetic variations, because that gives a model more exposure to situations it might otherwise never encounter. A world model, a system that learns how a driving scene evolves and can generate physically coherent variations of a rare event, is valuable for exactly this reason. It expands coverage without waiting for the world to hand a fleet three complicated near misses before lunch. NATIX has written before about how world models support both training and validation of driving systems, and the distinction carries directly into this discussion.
A safety case, though, cannot be built entirely on synthetic scenarios, because simulation reflects the assumptions its designers made, and a regulator's job is precisely to question those assumptions. Synthetic data is excellent for stress testing known failure modes. It is not a substitute for evidence of how a system behaved when it met the real, messy, physically imperfect world. That is why regulators keep returning to real miles, real disengagements, and real incident data even as simulation tools improve. Coverage and fidelity are different jobs, and only one of them can close out a safety case.
Put together, a working validation pipeline looks less like a single test and more like a loop that keeps running. The ODD sets the boundary. Scenario generation identifies what needs to be tested inside that boundary. Simulation covers volume and lets teams probe rare combinations cheaply. Real-world driving, tracked through disengagement and incident data, grounds the whole argument in what has actually occurred. As the system changes, or as it moves into a new domain, the loop runs again.
This is also why validation never really finishes. An ODD that covers a sunny city grid does not automatically cover a snowy highway three states away. Every expansion into a new domain reopens the same question a regulator is really asking: what evidence do you have for these conditions specifically, not conditions like them.
This is also why many companies expand carefully, one geography at a time, rather than switching on nationwide coverage overnight. Every new city or region means new weather patterns, new road layouts, and new driver behavior, so the ODD has to be redefined and the evidence rebuilt before the system is trusted there. It is a slow, deliberate process by design, not a limitation regulators impose out of caution alone.

This is where the data problem becomes visible. A safety case is only as strong as the real-world evidence behind it, and most of that evidence is limited by how much of the world a company has actually observed. Front-facing dashcam footage from a handful of test cities is not the same as structured, multi-camera coverage across a wide range of geographies, weather, and traffic behavior. Rare events also do not happen only in front of the car. They unfold across angles, and a dataset built entirely from forward-facing sensors misses a lot of what a regulator would want to see.
NATIX approaches this from the data side. Its VX360 hardware collects multi-camera footage from everyday drivers around the world, and the NATIX and Valeo partnership is built specifically around using that real-world, surround-view data for both training and validation of driving world foundation models. That distinction matters here. The same real-world grounding that strengthens a model during training is also the kind of evidence a safety case needs when it has to show a system was tested against the true diversity of conditions it claims to handle, not just the conditions a lab happened to have on hand. Tools that make that footage searchable matter too, because finding the rare, safety-relevant clip inside millions of hours of driving is its own problem, separate from collecting the footage in the first place.
For a team trying to expand into a new domain, that head start matters. Instead of waiting months to accumulate enough of its own miles in a new city or climate, a team can draw on footage that already exists there, and focus its own testing on the specific gaps that remain. None of this replaces the regulatory process itself. It supplies raw material that safety cases are built from: structured, geographically diverse, real-world driving data that goes beyond what any single company's own fleet could gather alone.
The industry's hardest problem right now is not building a system that performs well on a fixed test set. It is producing evidence that a regulator can trust, spanning the full range of situations a system will actually meet on the road. Benchmarks will keep improving, and they should, but they measure a sample chosen in advance. Regulators are asking about a domain, tested and observed as completely as anyone can manage. Closing that gap is not a matter of a better score. It is a matter of showing more of the real world, and showing it well enough that the argument holds up.