Sara Dorfman, Maya Vishnevsky, Omer Dahary, Or Patashnik, Daniel Cohen-Or
To address the diversity problem in text-to-image models, this paper proposes a method that induces structured variation at the text level using a Vision Language Model, enabling users to navigate image galleries along meaningful axes.
State-of-the-art text-to-image models have high visual fidelity but suffer from a collapse in diversity. Existing diversity methods often produce outputs driven by incidental variations rather than meaningful design choices.
The method exploits the fact that recent models are trained on elaborate captions, decoupling semantic decisions from pixel generation. It uses a Vision Language Model (VLM) to operate on the full scene context and employs an agentic workflow to enforce structured variation attuned to the original prompt.
The approach demonstrates the creation of diverse and navigable design spaces where every variation corresponds to a specific, user-understandable semantic decision. It presents a paradigm shift from relying on stochastic variation within the model to inducing diversity directly at the text level.