Dian Zheng, Harry Lee, Manyuan Zhang, Kaituo Feng, Zoey Guo, Ray Zhang, Hongsheng Li
InterleaveThinker is a multi-agent pipeline that enables existing image generators to produce interleaved text-image sequences by leveraging planner and critic agents to plan, evaluate, and refine the generation process.
Recent image generators excel at single-image generation and editing but cannot perform interleaved generation (text-image sequences) due to architectural limitations. Existing Unified Multimodal Models (UMMs) also show limited performance on this task.
InterleaveThinker improves interleaved generation performance across various image generators (e.g., FLUX, SD3). It achieves performance comparable to Nano Banana and GPT-5 on interleaved generation benchmarks. Additionally, it significantly enhances the base model on reasoning-based benchmarks (WISE, RISE).