PASC — Physical AI Safety ConsortiumPASC

PASC Research Agenda · Version 0.1

Seven open problems for physical AI safety.

Physical AI safety cannot be established at the model layer alone. This agenda identifies the open problems that emerge when models receive memory, sensors, bodies, and the ability to act over time.

01

Specification

Whether harm can be defined precisely enough to enforce, without stripping away the context that makes it harmful.

The problem

Physical AI needs safety requirements that are precise enough to evaluate or enforce while preserving the context that determines whether an action is harmful. A useful specification must account for the system, task, people, environment, time horizon, uncertainty, and available interventions—not only label an isolated action safe or unsafe.

Why it remains unresolved

Physical harm is contextual and unfolds over time. The same motion may be acceptable or dangerous depending on proximity, intent, speed, authority, human expectations, and possible escape routes. Requirements can conflict, while criteria that are too narrow encourage benchmark-specific compliance.

Initial research questions

  • What entities, events, constraints, and time horizons must a physical AI safety specification represent?
  • How should specifications express uncertainty, competing duties, and context-dependent harm?
  • Which requirements belong to the policy, runtime safety layer, operator, hardware, or deployment process?
  • How can an informal requirement be traced to a measurable criterion and an enforceable intervention?

What would count as progress

  • A shared vocabulary for hazards, harms, unsafe states, failure boundaries, interventions, and residual risk.
  • Formal and natural-language templates that independent teams apply consistently to the same scenarios.
  • Reference cases that expose ambiguity and conflict between plausible safety requirements.
  • Evidence that specifications generalize beyond the examples used to author them.

Who is needed

Robotics and AI researchersControl and formal-methods specialistsHuman-factors and HRI researchersDeployment safety engineersStandards, legal, and policy experts
02

Sufficiency

Whether a bounded representation of the present can preserve everything safety requires from an unbounded past.

The problem

What information, history, observations, forecasts, tests, and evidence are enough to support a particular safety judgment? A system may observe extensively yet omit the one fact that changes the risk assessment. A benchmark may contain many scenarios yet remain insufficient for the claim attached to it.

Why it remains unresolved

There is no general criterion for when a representation or body of evidence preserves everything relevant to safety. Relevance may depend on latent intent, earlier interactions, hidden physical state, or deployment conditions absent from the data. Unbounded history is also costly, privacy-sensitive, difficult to validate, and incompatible with real-time operation.

Initial research questions

  • Sufficient for which claim: detecting a hazard, forecasting harm, selecting an intervention, or certifying an operating envelope?
  • What past information can be discarded without changing a correct safety decision?
  • Can sufficiency be characterized causally rather than by predictive accuracy alone?
  • Can a monitor recognize insufficient evidence and defer, slow, stop, or request help?

What would count as progress

  • Claim-specific sufficiency criteria with explicit assumptions.
  • Datasets with controlled long-range dependencies, distractors, missing information, and hidden-state interventions.
  • Bounded representations that preserve safety-relevant decisions while removing irrelevant history.
  • Calibrated abstention or escalation when observations do not support a reliable judgment.

Who is needed

Representation-learning researchersCausal inference and statistics expertsState-estimation teamsInformation theoristsDeployment partners with long-horizon traces
03

Anticipation

Whether an unsafe consequence can be predicted, with calibrated confidence, before the last moment it can still be prevented.

The problem

Can a physical AI system identify unsafe consequences early enough for an effective response, while remaining calibrated about uncertainty and avoiding unusable false alarms? A continuously maintained estimate of risk is one candidate framework for studying this question, not a settled definition of safety.

Why it remains unresolved

Long-horizon forecasts compound model error and depend on how robots and people react to one another. Earlier warnings create more intervention time but also more uncertainty and false positives. A predictor can therefore appear accurate while still being operationally useless.

Initial research questions

  • Which future outcomes must be forecast, and at what spatial and temporal resolution?
  • How should uncertainty over plans, dynamics, human behavior, and external events be propagated?
  • How should detection lead time be measured when several interventions or failure boundaries exist?
  • When should uncertainty itself trigger slowing, deferral, or information gathering?

What would count as progress

  • Closed-loop evaluations in which actions affect subsequent observations and hazards.
  • Metrics for detection lead time, calibration, false-alarm burden, missed-hazard severity, and intervention feasibility.
  • Benchmarks with delayed consequences, branching futures, and locally safe actions that form unsafe sequences.
  • Prospective evidence that earlier warnings reduce the frequency or severity of failure.

Who is needed

World-modeling and planning researchersControl and uncertainty researchersHRI specialistsRuntime-monitor developersRobotics companies with real interaction traces
04

Falsification

Every test is finite. The world is not. Whether failures beyond the benchmark can be found, and what passing can ever prove.

The problem

How can we construct serious attempts to disprove a scoped safety claim? Benchmark success must be treated as evidence under stated assumptions, not confirmation that a system is universally safe.

Why it remains unresolved

Scenario spaces are combinatorial, failures can be rare, and test generators inherit assumptions from their designers. Simulator success can omit hardware and human factors, while aggregate scores conceal catastrophic tails, correlated failures, and untested regions.

Initial research questions

  • What safety claims are precise enough to be falsifiable?
  • How can search, simulation, formal analysis, and hardware testing complement one another?
  • How should tests target rare, severe, and multi-step failure modes?
  • What can a passing result legitimately support, given the tested domain and assumptions?

What would count as progress

  • Every result paired with its scope, assumptions, exclusions, and known failure modes.
  • Reproducible counterexample-generation protocols that outperform unguided testing.
  • Coverage measures tied to hazard structure rather than scenario counts alone.
  • Public failure records with minimal reproductions, severity, provenance, and remediation status.

Who is needed

Verification and validation researchersFormal-methods and testing specialistsScenario-generation teamsIndependent evaluatorsReliability and accident-analysis experts
05

Recoverability

Which states still admit a way back, and whether a long-horizon safety loop can keep the system inside them.

The problem

Does an unfolding unsafe situation still admit a feasible intervention that can avoid or materially reduce harm—and can the system recognize and execute it before the opportunity disappears? A maintained estimate of risk and intervention feasibility is one candidate framework for studying this question.

Why it remains unresolved

Recovery depends on dynamics, latency, actuator limits, contact state, surrounding people, and other agents. Stopping is not always safe. The boundary between recoverable and unrecoverable states is uncertain, action-dependent, and different for each kind of harm.

Initial research questions

  • How should recoverability be defined when harm can be reduced but not eliminated?
  • Which interventions—slow, stop, retreat, replan, hand over, or ask for help—remain feasible in each state?
  • How should sensing, computation, communication, and actuation latency enter the estimate?
  • When can a safety intervention create a different hazard?

What would count as progress

  • Estimators for intervention windows and points beyond which specified harms cannot be avoided.
  • Closed-loop tests of intervention success, residual harm, latency margin, and secondary hazards.
  • Recovery policies robust to dynamics error, delayed observations, and degraded actuators.
  • Hardware evidence across manipulation, locomotion, and mobile navigation.

Who is needed

Control theoristsMotion-planning and runtime-assurance researchersHardware and embedded engineersHuman-factors specialistsPartners able to run controlled recovery experiments
06

Invariance

Which guarantees, if any, survive the move from simulation to hardware, from one body to another, and from one machine to many.

The problem

Which safety claims, measurements, and mechanisms survive changes of body, task, environment, operator, scale, software stack, and deployment regime—and which must be revalidated?

Why it remains unresolved

Bodies differ in mass, reach, compliance, latency, sensing, and failure modes. Simulation omits contacts, wear, timing effects, and human responses. A metric may keep its name while changing operational meaning across embodiments; generalization claims are empty unless both the transformation and preserved property are named.

Initial research questions

  • Which properties should remain invariant, and under exactly which transformations?
  • When is recalibration enough, and when is a new specification or evaluation required?
  • How can sim-to-real gaps be decomposed into perception, dynamics, control, timing, and interaction?
  • How much hardware evidence is required before making a cross-embodiment claim?

What would count as progress

  • Matched evaluations across simulation and multiple physical embodiments.
  • Transfer matrices showing which claims hold, degrade, or fail across named changes.
  • Standardized reporting of body properties, sensing, latency, payload, environment, and operating envelope.
  • Explicit triggers for revalidation when a guarantee does not transfer.

Who is needed

Robotics companies across embodiment classesSimulation and digital-twin teamsDomain-adaptation researchersHardware and controls engineersBenchmark and metrology specialists
07

Adversary

Whether any safety guarantee survives an adaptive opponent inside the loop.

The problem

Does a scoped safety claim survive an opponent who can manipulate elements of the perception–reasoning–action loop? An adversary may target sensors, instructions, tools, communications, physical context, operators, or the runtime safety mechanism itself.

Why it remains unresolved

Physical attackers combine cyber, semantic, social, and mechanical strategies. Their actions alter the world, not only an input string. Threat models vary by access, knowledge, budget, and persistence, while defensive monitors can themselves be evaded, overloaded, spoofed, or induced to intervene dangerously.

Initial research questions

  • What attacker capabilities and objectives are realistic for each deployment?
  • How can prompt, sensor, tool, communication, and physical-scene attacks be evaluated together?
  • Can an attacker create an unsafe plan whose individual actions evade local checks?
  • Which defenses degrade gracefully when a sensor, model, or communication channel is compromised?

What would count as progress

  • Explicit threat models covering attacker access, knowledge, resources, constraints, and target claims.
  • Reproducible embodied attack suites evaluated against complete closed-loop systems.
  • Adaptive red-team protocols that test monitors, intervention logic, and fallback behavior.
  • Defenses evaluated against held-out attacks and independent teams.

Who is needed

AI and robotics red teamsCybersecurity and adversarial-ML researchersHardware-security specialistsRuntime-guardrail developersIncident-response and responsible-disclosure experts

Candidate framework · Problems 03 and 05

Anticipation and recoverability

A maintained estimate of risk is one candidate way to study whether danger can be recognized early enough for intervention. It is not PASC’s final safety definition. Its value must be judged by calibration, detection lead time, intervention success, and residual harm.

Problems → Research → Evidence → Benchmarks → Standards

PASC does not begin by declaring a standard. It begins by establishing the evidence a standard would require.

Contribute to the research agenda →