Robot arm study finds nearly 4x success gap across input formats
What happened
Northwestern University and Stanford ran a controlled comparison on a real 7-DoF robot arm, changing only the input data format. Task success rates differed by nearly four times: stereo raw images with cross-view attention fusion hit 59%, channel concatenation 45%, a single color photo 43%, color plus hardware depth 41%, and 3D point clouds just 14%. Transparent glass cups caused the RGB-D camera to drop readings. The paper argues that when on-device depth sensing fails often on real data and downstream stages cannot correct it, end-to-end training through the gradient is the better trade. The paper is still under review, code is not public, and no third party has reproduced it.
Written by AI from the coverage · updated 47 minutes ago
Coverage
Follow the reports to see the story from different sides.
- Computing Life · Share · Yage玻璃杯从深度图里消失之后:机器学习拆任务的取舍
西北大学与斯坦福大学在真实七自由度机械臂上做受控对照,仅改变输入数据形式,成功率相差近四倍:双目原始画面加跨视角注意力融合达59%,通道直接拼接45%,单张彩色照片43%,彩色照片加硬件距离图41%,三维点云仅14%。透明玻璃杯会让RGB-D相机丢失读数,论文认为前端测距环节在真实数据上失效率高、下游又无法纠偏时,打通梯度端到端训练更划算。该论文仍在审稿阶段,代码未公开,也无第三方复现。
Heat over time
Not enough continuous observations to draw a trend yet.