If you can’t trust AI to verify an autonomous driving stack, can you still use AI in the verification and validation (V&V) process? Yes, but not where most people first assume.
Working with AV programs at scale, we’ve found that one of the most valuable ways to apply AI in V&V today is not in deciding whether a system passes or fails, but in helping engineers interpret verification data that has grown beyond human-scale analysis.
Modern AV development generates enormous amounts of data across real-world driving, simulation, replay, and other test environments. Understanding which conditions matter, where meaningful gaps exist, and what actions to prioritize requires engineers to analyze datasets that continue to grow in both scale and complexity.
AI-assisted analysis helps teams navigate that complexity by accelerating the exploration, interpretation, and action on verification data while keeping the underlying verification methodology deterministic and reproducible.
This blog explores how Foretellix applies AI on top of structured V&V workflows to help engineers identify coverage gaps, prioritize testing, and build stronger, evidence-backed safety cases.
The Problem: Verification Data Has Exceeded Human-Scale Interpretation
V&V engineers today face a widening gap. On one hand, they have massive amounts of raw, noisy data: real-world fleet logs, simulation results, replay systems, HIL testing, coverage reports, and KPI outputs, each holding a piece of the overall verification picture. On the other hand, what they actually need is analysis: a coherent understanding of how the system behaves across the Operational Design Domain (ODD). Closing that gap starts with processing the raw data into structured scenario intelligence, automatically labeling every scenario and its parameters and organizing them under a well-thought-out verification plan. Only on top of that structured foundation can AI meaningfully help analyze the data, find patterns, and answer questions.
A modern AV verification workspace contains millions of test runs spanning thousands of coverage dimensions, with verification plans nested ten or more levels deep. Even highly experienced engineers struggle to manually trace patterns across datasets of this size.
The challenge becomes even harder because autonomous driving failures rarely depend on a single variable. A lane-change scenario alone can vary across illumination, weather, road geometry, traffic density, ego speed, and contender behavior. Each additional dimension multiplies the number of combinations engineers must evaluate.
The result is a combinatorial problem that quickly exceeds what humans can reliably interpret through dashboards, spreadsheets, and manual investigation.
This is where many V&V workflows begin to slow down. Engineers spend significant time searching for relevant data, correlating failures across systems, and determining which gaps are actually important. In many organizations, extracting insight from verification data now consumes more engineering effort than generating the data itself.
AI-Assisted Analysis on Top of Deterministic Verification
Foretellix addresses this challenge through the Foretify Physical AI Toolchain, which enables data-driven development, validation, and safety evaluation for Physical AI systems, such as advanced autonomous vehicles.
Within this toolchain, Foretify Evaluate unifies real-world and simulation data into a shared framework so teams can measure coverage and performance consistently across the ODD. This creates a structured foundation for large-scale verification analysis.
The verification workflow itself remains deterministic and reproducible. Coverage-driven verification methodologies, semi-formal verification plans, KPIs, coverage maps, coverage models, and high-accuracy temporal scenario labeling generate the underlying structured verification framework. The AI operates on top of that data to support analysis and navigation.

The AI doesn’t determine whether a system is safe, whether a test passes, or whether a scenario qualifies as an event of concern. Those decisions remain grounded in the verification framework and its deterministic methodologies. The role of AI is helping engineers interpret verification data at a scale that manual workflows can no longer support.
Every insight remains traceable to specific scenarios, runs, coverage buckets, metrics, and log segments. Engineers can drill directly from a high-level observation into the underlying evidence that produced it.
A Different Way to Work With Verification Data
Traditionally, engineers investigating verification results move between dashboards, reports, scripts, and manually constructed queries. Understanding a coverage gap or identifying the conditions associated with a failure can require exporting data into notebooks or building custom analysis flows.
AI-assisted analysis changes that interaction model. Instead of manually navigating tools, engineers can ask questions in natural language:
- Where are the largest coverage gaps?
- Which conditions correlate most strongly with failures?
- Where should additional testing be prioritized?
- Which parts of the verification plan are currently unreachable, and how should the plan be adapted?
The system responds with structured, evidence-backed answers while preserving the context of the investigation. Engineers can continue refining questions, drilling deeper into scenarios, and exploring relationships across datasets without restarting the analysis process each time.
This creates a much more iterative workflow where analysis becomes an ongoing exploration of the verification data rather than a sequence of disconnected reporting tasks.
Making Combinatorial Complexity Manageable
One of the biggest difficulties in AV verification is understanding how conditions interact. A system may appear to struggle with a certain maneuver, but the actual issue often emerges only under a very specific combination of circumstances. A failure might occur during a lane change under low-light conditions, in dense traffic, on curved road geometry, and within a narrow ego-speed range.

These relationships are difficult to identify manually because they span multiple interacting dimensions simultaneously.
AI-assisted analysis helps teams identify statistically meaningful patterns across those dimensions. Instead of looking at isolated events one at a time, engineers can understand where failures cluster, which combinations contribute most strongly to events of concern, and where gaps exist relative to target ODD coverage.
This allows teams to distinguish meaningful signals from background noise much more efficiently.
What This Looks Like in Practice
Consider an engineer reviewing a verification workspace built from millions of miles of driving data alongside several million simulation runs, all processed into structured coverage information.
The workspace returns a verification grade of 65.45%. That number alone does not explain what should happen next.
Instead of opening multiple dashboards and manually correlating reports, the engineer can simply ask the AI: “Where are the biggest coverage gaps?”
The system identifies several underrepresented coverage areas, including nighttime lane changes in high-density traffic conditions.
The engineer follows up: “In the conditions we have encountered, where are we seeing the most events of concern?”
The system identifies a statistically significant concentration of events under a specific combination of traffic density, road geometry, and lighting conditions occurring well above the fleet baseline.
The results are presented through structured visualizations that make complex relationships easier to interpret, allowing engineers to quickly understand where failures cluster, how coverage is distributed across the ODD, and which conditions contribute most strongly to events of concern.
The engineer then requests the associated log segments and receives direct links to the relevant scenarios, timestamps, routes, and sensor traces. Those insights can also feed directly into the next stage of the workflow. Rather than sending out the testing fleet to try and close the identified gaps in the ODD, by using Foretellix’s synthetic data generation (SDG) capabilities, teams can generate additional variations of the identified scenarios. This will enable an expansion of the ODD coverage, exploration edge cases, and creation of new training and validation data focused on the conditions associated with problematic behavior.

What previously required hours of exporting CSVs, building scripts, and cross-referencing dashboards now takes a few iterative questions.
Critically, every conclusion still comes from the underlying verification methodology and structured verification data. The AI didn’t decide what qualified as an event of concern. The verification framework defined that logic. The AI just accelerated the process of finding patterns inside the data.
Faster, More Consistent Decisions
Beyond accelerating individual investigations, this workflow changes who can participate in verification analysis. Tasks that previously required deep familiarity with specific verification structures or internal tooling become more broadly accessible and allow teams to extract meaningful insight from the data.
The result is faster and more consistent decision-making across V&V teams, with analysis methodology shared rather than locked inside individual expertise.
Safety Cases Built on Traceable Evidence
Building a safety case requires clear, reproducible evidence showing how the system performs across relevant conditions and scenarios. Foretify Evaluate supports this by maintaining a direct connection between analysis results and the underlying verification data. Coverage metrics, run outcomes, scenario labels, and verification results remain traceable to specific scenarios and test executions.
That traceability is essential for safety engineering and regulatory review. The system doesn’t produce conclusions such as “the AI believes the vehicle is safe.” Instead, it presents the conditions tested, the outcomes observed, the coverage achieved, and the gaps that still remain.
The AI accelerates how that evidence is explored, assembled, and interpreted, while the evidence itself remains grounded in deterministic verification methodologies.
Scaling Verification Analysis for Physical AI Systems
Returning to the question that we opened with: where can AI be trusted in the verification process?
Whether AI should eventually play a role in verification decisions, determining whether a system is safe, whether a test passes, whether a scenario qualifies as an event of concern, is one the industry will continue to debate as the underlying technology matures.
As of today, those judgments need to remain anchored in deterministic verification methodology, where every conclusion is traceable to specific scenarios, runs, and metrics. The maturity and auditability required for AI to make those calls at the rigor a safety case demands isn’t there yet.
What is clear is where AI already brings efficiency: helping engineers understand verification data that has grown beyond human-scale interpretation and finding patterns hidden across millions of test runs and thousands of dimensions. Moving teams from “we have the data” to “we know what it means.”
That’s a problem AI can already solve, with full traceability and without compromising the integrity of the underlying verification methodology. As autonomous systems continue to grow in capability and complexity, that role becomes essential, not optional. The rigor stays in the verification methodology. The AI helps engineers understand the results.