SurgCF-Bench: A Counterfactual Feasibility Benchmark for Long-Horizon Surgical Robot World Models
Abstract
World models for robot learning are often evaluated by next-state prediction, action likelihood, or short-horizon task success. For surgical robot autonomy, a complementary question is whether a model can reject a continuation that is locally plausible but incompatible with the current state, task phase, or trajectory context. We introduce SurgCF-Bench, an offline benchmark for counterfactual feasibility in surgical robot world models. The current release, SurgCF-Bench-ROSMA-v0, uses 207 public Da Vinci Research Kit kinematic episodes from ROSMA across Pea-on-a-Peg, Post-and-Sleeve, and Wire-Chaser-I. It defines single-step, scene-specific, long-horizon, and hard-negative tracks with counterfactuals generated by time shifts, cross-episode substitutions, cross-task substitutions, reversals, loops, skips, overscaling, Gaussian perturbations, phase-neighbor substitutions, and velocity-matched chunks. Lightweight baselines reveal a consistent failure mode: action-only likelihood is near chance on context-mismatched continuations, while state-action scoring works at short horizons but degrades as the horizon grows. On 0.5 s counterfactual segments, a state-action probe reaches 0.78-0.81 AUROC on context negatives; at 2 s and 5 s it drops to roughly 0.60-0.63 and 0.56-0.58. SurgCF-Bench therefore offers a reproducible public-data diagnostic for long-horizon contextual feasibility before hardware access or clinical validation is available.