Kai Ploeger, Jan Peters
Residual learning using directional task-error supervision enables stable three- to five-ball juggling on real robotic arms.
Existing residual learning methods suffer from low sample efficiency due to scalar rewards or random exploration. Five-ball juggling is extremely challenging, requiring years of human practice.
A residual learning framework that uses directional task-error as learning feedback and a task-error model for sample selection. A fixed-Jacobian Newton update combines analytic prior knowledge with directional information.
Stable three-, four-, and five-ball juggling on Barrett WAM arms with monotonic task-error decrease after the first attempt. Both directional feedback and informative prior are necessary; the simplest method (fixed-Jacobian Newton) is most reliable.