Whether an unsafe consequence can be predicted, with calibrated confidence, before the last moment it can still be prevented.
The problem
Can a physical AI system identify unsafe consequences early enough for an effective response, while remaining calibrated about uncertainty and avoiding unusable false alarms? A continuously maintained estimate of risk is one candidate framework for studying this question, not a settled definition of safety.
Why it remains unresolved
Long-horizon forecasts compound model error and depend on how robots and people react to one another. Earlier warnings create more intervention time but also more uncertainty and false positives. A predictor can therefore appear accurate while still being operationally useless.
Initial research questions
- Which future outcomes must be forecast, and at what spatial and temporal resolution?
- How should uncertainty over plans, dynamics, human behavior, and external events be propagated?
- How should detection lead time be measured when several interventions or failure boundaries exist?
- When should uncertainty itself trigger slowing, deferral, or information gathering?
What would count as progress
- Closed-loop evaluations in which actions affect subsequent observations and hazards.
- Metrics for detection lead time, calibration, false-alarm burden, missed-hazard severity, and intervention feasibility.
- Benchmarks with delayed consequences, branching futures, and locally safe actions that form unsafe sequences.
- Prospective evidence that earlier warnings reduce the frequency or severity of failure.
Who is needed
World-modeling and planning researchersControl and uncertainty researchersHRI specialistsRuntime-monitor developersRobotics companies with real interaction traces
Which guarantees, if any, survive the move from simulation to hardware, from one body to another, and from one machine to many.
The problem
Which safety claims, measurements, and mechanisms survive changes of body, task, environment, operator, scale, software stack, and deployment regime—and which must be revalidated?
Why it remains unresolved
Bodies differ in mass, reach, compliance, latency, sensing, and failure modes. Simulation omits contacts, wear, timing effects, and human responses. A metric may keep its name while changing operational meaning across embodiments; generalization claims are empty unless both the transformation and preserved property are named.
Initial research questions
- Which properties should remain invariant, and under exactly which transformations?
- When is recalibration enough, and when is a new specification or evaluation required?
- How can sim-to-real gaps be decomposed into perception, dynamics, control, timing, and interaction?
- How much hardware evidence is required before making a cross-embodiment claim?
What would count as progress
- Matched evaluations across simulation and multiple physical embodiments.
- Transfer matrices showing which claims hold, degrade, or fail across named changes.
- Standardized reporting of body properties, sensing, latency, payload, environment, and operating envelope.
- Explicit triggers for revalidation when a guarantee does not transfer.
Who is needed
Robotics companies across embodiment classesSimulation and digital-twin teamsDomain-adaptation researchersHardware and controls engineersBenchmark and metrology specialists
Candidate framework · Problems 03 and 05
Anticipation and recoverability
A maintained estimate of risk is one candidate way to study whether danger can be recognized early enough for intervention. It is not PASC’s final safety definition. Its value must be judged by calibration, detection lead time, intervention success, and residual harm.
Context + historyWhat has happened?
Maintained estimateWhat may happen next?
Intervention windowWhat can still be prevented?
Problems → Research → Evidence → Benchmarks → Standards
PASC does not begin by declaring a standard. It begins by establishing the evidence a standard would require.
Contribute to the research agenda →