For the first time, there is a global regulation that requires coverage evidence to certify an autonomous vehicle for public roads. UNECE WP.29 R185 sets the certification standard for any and all Automated Driving Systems (ADS) (equivalent to SAE Level 3, 4, and 5) and it takes effect in January 2027.
This blog explains in plain terms what R185 requires for coverage. Coverage-driven validation methodology has long been championed by Foretellix and seeing it reflected in R185 confirms the direction we have pushed for years. Meeting the requirements of R185 requires purpose-built methodology and tooling that can evaluate and measure coverage systematically. Engineering teams can start preparing for the January 2027 deadline using a toolchain that already supports R185, like Foretellix’s customers have.
Scope note: R185 applies to driver-out autonomous systems (equivalent to SAE Level 3, 4, and 5). It does not cover driver assistant systems (SAE L2++), those are addressed separately under UN R171.
What Is New?
R185 is the first regulation to require manufacturers to prove ODD (Operational Design Domain) coverage as part of certification. According to the standard, miles driven is not evidence. Scenario coverage, tied to a documented methodology, with metrics and targets, is evidence. Without it, there is no safety case, and without a safety case, there is no certification.
What Is Coverage?
Coverage is used to quantify the completeness of testing across both the conditions represented by the test scenarios and the behavioral competencies and KPIs demonstrated by the ADS. In practice, that means two things. First, how well your set of test scenarios represents the full range of situations the ADS will face inside its ODD, including the edge cases at the boundaries, not just the easy middle of the distribution. Second, how completely the testing demonstrates that the ADS can perform the required behavioral competencies and meet the associated KPIs across that scenario space. Coverage is not a count of test miles or test cases, it is a measure of how completely those tests represent the conditions the vehicle needs to handle safely.

The Coverage Checklist
Below are the coverage related requirements as stated in the regulation (full document here), with what each one means for your team.
1. You need a documented scenario identification and generation method.
Per paragraph 5.3.2.13, “the safety concept must describe the scenario identification and generation approach” and show that it covers “the appropriate nominal, critical, and failure situations” using data-driven, knowledge-driven, and stochastic methods.
What it means: a spreadsheet of test cases is not enough. You need a repeatable, documented method for scenario identification, and classification for scenario generation, and it has to explicitly account for all three situation types the regulation defines: nominal (normal driving, no critical or failure conditions present), critical (a situation requiring prompt action to avoid or mitigate a crash), and failure (a situation where a system fault, such as a sensor or computer failure, compromises the ADS’s ability to perform the driving task at all).
2. You need a documented scenario selection approach tied to the ODD.
Per paragraph 7.3.2.14, the safety concept must describe the approach to scenario selection to cover reasonably foreseeable situations and conditions, including selection of scenarios where the ADS needs to initiate a fallback response, and use of appropriate techniques to explore the parameter space.
What it means: you have to show your work on how scenarios map back to the ODD, including edge cases at ODD boundaries, not just the easy middle of the distribution.
3. You need defined coverage metrics and targets, not just test counts.
Per paragraph 7.3.2.16(a), the safety case must include verification and validation plans that explain “how scenarios and situations are selected… to provide reasonable coverage of the ODD and its boundaries,” along with the “methodology, metrics, and targets used to determine reasonable ODD coverage.”
What it means: you need a defensible coverage metric, a target for that metric, and evidence that you hit the target. “We ran a lot of simulations” is not a metric.
4. Coverage has to be demonstrated across all three testing pillars together, not separately.
Per paragraph 8.3.2.1.2, the assessment verifies that “the combined coverage of the testing results from all pillars (virtual, track, real-world) is sufficient to support the ADS safety case claims.”
What it means: simulation coverage, track coverage, and real-world coverage are assessed as one combined picture. Gaps in one pillar have to be justified or closed by another. That also means you need a documented way of showing how coverage was collected and calculated across all three pillars, not just a claim that it happened. Without that documentation, assessors have no way to verify whether a gap in one pillar was actually closed by evidence from another.
5. Your test evidence has to be assessed for coverage, consistency, and relevance, explicitly.
Per paragraph 8.3.2.4.1.5, assessors verify “the suitability of the set of tests carried out as evidence to support the safety case, in particular in terms of coverage, consistency, and relevance.”
What it means: coverage is a named, assessable criterion. It is not a side effect of running enough tests, it is something you have to demonstrate on its own terms.
6. Reasonable coverage of foreseeable conditions is a pass/fail criterion for the safety case.
Per paragraph 8.3.1.4(g), the safety case is only considered robust if the testing evidence “provides reasonable coverage of foreseeable operating conditions and events in the intended area of operation,” including conditions at and beyond ODD boundaries.
What it means: this is not a nice-to-have. Insufficient coverage is a documented reason the safety case can be rejected.
7. Coverage must measure both the scenarios used as test inputs and the behavioral competencies and KPIs demonstrated during testing.
Annex 7, Section 2.2 states that “coverage” refers to the degree to which scenarios sufficiently incorporate driving situations to validate the relevant requirements, and that it “can be measured both on the test scenarios serving as input to the test as well as on the behavioural competencies and KPIs demonstrated during testing.”
What it means: you need coverage metrics on your input scenario set and on the competencies and KPIs your ADS actually demonstrated. A complete coverage argument should address both dimensions, rather than treating scenario coverage alone as sufficient.
What AV Developers Should Do Now
R185 turns these requirements into a set of questions every AV program now has to answer.
Key questions include:
- How is the ODD defined and decomposed into testable conditions?
- How are behavioral competencies mapped to scenarios?
- How are nominal, critical, and failure scenarios generated?
- How is coverage measured across simulation, track, and real-world testing?
- How are metrics and targets defined for reasonable ODD coverage?
- How is evidence connected to safety claims?
- How are gaps identified and closed?
- How are in-service learnings fed back into the validation process?
These are not only regulatory questions. They are engineering questions. They are safety questions. And they are now central to the industry’s path toward scalable deployment.
Foretellix Has a Toolchain For This
The Foretellix Physical AI toolchain supports this process by helping teams define and decompose the ODD, including the specific driving capabilities the ADS is expected to demonstrate within it. These behavioral competencies can include maintaining lane position, changing lanes safely, responding appropriately to pedestrians and cyclists, and performing a safe fallback. The toolchain maps those competencies to scenarios and generates nominal, critical, and failure scenarios in a structured, repeatable way. It then measures not only whether the right scenarios have been tested, but also whether the ADS demonstrated those required driving capabilities and met the associated KPIs during testing. By unifying this evidence across simulation, track, and real-world testing, the toolchain can connect results to the safety case, identify coverage gaps, and help teams determine what to test next.
A Turning Point for the Industry
The UNECE WP.29 R185 ADS framework is more than a new regulation. It is a signal that the global conversation around autonomous vehicle safety is maturing.
The industry is moving beyond broad claims and toward structured evidence. Beyond miles driven and toward meaningful coverage. Beyond isolated testing and toward unified validation across simulation, track, and real-world environments. Beyond one time approval and toward continuous safety management throughout the ADS lifecycle. This is the direction Foretellix has long believed the industry must take.
Our customers have had access to ODD coverage measurement, scenario-based validation, and safety case evidence generation for some time. With these capabilities they are well positioned to meet the requirements of the R185 regulation. The specific methodology and coverage targets will vary by program and ODD, but the toolchain gives teams the framework to address and answer each of these questions with evidence, and we would be glad to walk through them with your team.
Contact Foretellix to learn more.