The model is not the safety layer
In September, Robocurve put three robot policies on real robot arms and gave them harmful instructions, 100 trials each. GPT-6 Astra refused 2 of 100 on safety grounds and carried out 60. Claude Fable 5.1 refused 20, every one of them the same instruction. One heading in their report is the reason this post exists: “The more capable policy refuses less and completes more”. (Source: RoboHarm.)
So I built the layer between the model and the motor, on the same $200 arm that carried the pre-registration last month, and ran the same task twice: with the layer off, and with it on.
Off, it does exactly what it was asked, and lowers the tool onto the hand pointed end first. On, a hand monitor sees the hand and the arm is held within one frame; a vision model is asked whether a handover is safe, and its answer and how long it took are written down; the planner turns the pointed end away before it comes near the hand; the handle goes into the palm; the arm will not move away until the hand is out. Every one of those checks is a row in a hash-chained log, and the right-hand panel of the video is a replay of that log, not an animation.
Things the log also says, because that is the point of keeping one: the arm sat 12–20 mm lower than it believed under the tool’s weight; the overhead camera loses the hand under the arm in the last seconds, so the layer has to tolerate that and compensate before the descent; the vision model called a still, open hand “moving” more than once; 58 of the 70 arm sessions, including both final takes, ran under an operator override with that fact stamped on every row. None of that is hidden, and none of it is a pre-registered result. This is a demonstration on a hobby arm, with a real screwdriver and my own hand.
What it is for: nobody publishes how often a safety layer fires, misses or false-alarms. I will. The model is not the safety layer, and the layer is not safe until someone has measured it.