科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Sensors (Basel, Switzerland)2026-07-27

ROEP: A Robotics-Oriented Evaluation Protocol for Deployment-Facing Vision-Language-Action Manipulation Policies.

Sangwoo Han, Hyunguk Choi

原始摘要(英文原文)· Original abstract
Vision-Language-Action (VLA) policies are increasingly evaluated on language-conditioned robotic manipulation benchmarks, but success rate alone often obscures runtime-interface alignment, the repeatability of observation-degradation effects, and the recovery-relevant semantics of failures. This study proposes ROEP (Robotics-Oriented Evaluation Protocol), a VLA-targeted deployment-oriented evaluation protocol that converts rollout outcomes into claim-level evidence for closed-loop robotic manipulation. ROEP first verifies clean-condition evaluability and runtime fidelity of the sensor-to-action execution interface, then evaluates controlled visual perturbations, repeated-run reference variability, timeout-dominant failures, and recovery/shield support. We apply ROEP to OpenVLA, X-VLA, and VLA-Adapter on eight LIBERO Goal and Object tasks, producing a 24-row evaluation matrix with complete clean and medium-perturbation evidence. Under the evaluated LIBERO simulation setting, the results show that runtime-interface fidelity and checkpoint provenance are important for interpreting X-VLA and OpenVLA Object results, while VLA-Adapter maintains strong Goal-suite performance but exhibits a substantial Object-suite clean-to-perturbation drop dominated by timeout termination. ROEP therefore clarifies which deployment-relevant claims are supported, withheld, or insufficiently evidenced by the available rollouts without certifying open-world deployment safety or recovery success. The protocol complements existing VLA benchmarks by reporting success rate together with runtime fidelity, repeatability, failure semantics, and recovery/shield support evidence.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

ROEP: A Robotics-Oriented Evaluation Protocol for Deployment-Facing Vision-Language-Action Manipulation Policies. — 科研速览 Science Skim