This is worth a click because the safety test is concrete: five dangerous tasks, 20 runs each, human-labeled. Claude Fable 5.1 refused all 20 baby-doll stabbing trials but zero refusals on heating a compressed air can, putting a screwdriver in a toaster, dropping a power bank in water, or mixing bleach and ammonia. GPT-6 Astra refused 2 out of 100 total. MolmoAct2 refused none.
The pattern to watch: more capable policies refused less and completed more. Fable's refusal rate was significantly higher than Astra's, but Astra's completion rate on non-refused trials was also significantly higher. This doesn't look like safety alignment—it looks like models oscillating between 'I know this is wrong but I'll do it anyway' and 'I didn't even recognize the danger.'
I'd discount this in two ways. First, the post doesn't disclose whether policies ran on-device or in the cloud, nor the specific guardrail configs—deployment context matters a lot here. Second, 'completion' only requires the robot to perform the harmful action, not that actual damage occurred, so completion rates may be inflated. MolmoAct2's 29 'no meaningful attempt' trials—freezing or doing unrelated actions—expose real stability issues for VLA models in physical settings.