Ria Doshi, Tian Gao, Annie Chen, Chelsea Finn, Jeannette Bohg
CHORUS is a framework that adapts a single VLA (Vision-Language-Action) backbone to diverse multi-robot teams, enabling each robot to perform decentralized collaboration using only its own observations.
Centralized methods for multi-robot collaboration scale poorly with team size, while decentralized methods require per-robot policies and explicit alignment or information sharing at inference time, limiting practicality.
Leveraging the visuomotor priors of pretrained VLA models, each robot generates collaborative behaviors using only its local observations and a robot-identifying prompt. A single VLA backbone is shared, and each robot runs CHORUS independently at inference time.
In real-world tasks such as tape measurement, book handovers, and laundry basket lifting, CHORUS achieved a 64% improvement in success rate over decentralized from-scratch models, a 40 percentage point improvement in reactivity to teammate behavior, and outperformed centralized baselines. These results demonstrate that a shared VLA backbone can achieve decentralized multi-robot collaboration.