Michael Finkelson, Daniel Segal, Eitan Richardson, Shahar Armon, Nani Goldring, Poriya Panet, Nir Zabari, Benjamin Brazowski et al.
A method to generate realistic multi-speaker dialogue audio using natural language scene descriptions and multiple voice references.
Existing multi-speaker systems bind speakers via structured supervision, preventing the generation of natural, non-studio audio. They also suffer from a 'reference shortcut' where the model bypasses the text prompt by matching acoustic similarity.
A text-to-audio foundation model pretrained on wild data is conditioned on multiple reference voices and a free-form text prompt. Reference latents are concatenated into the model's token sequence with identity-aware positional encodings, and a high-noise-biased timestep distribution forces reliance on the text prompt for speaker assignment.
Outperforms existing systems on speaker-binding metrics in the CoVoMix2-Dialogue benchmark while generating rich conversational audio with overlapping speech and ambient sound. Demonstrates the advantage of using a general-purpose audio model conditioned on free-form scene descriptions.