New Vision-Language-Action model beats larger baselines with 4x less data
Key claim: action chunking with a frozen vision backbone matters more than parameter count below 10M episodes. I reproduced the door subset on real hardware and got within 3 points.