Haozhe Wang, Weijia Feng, Jinpeng Yu, Che Liu, Ping Nie, Fangzhen Lin, Jiaming Liu, Ruihua Huang et al.
This paper addresses the structural limitation of visual generation models, which hallucinate due to fixed training data, by proposing a co-training framework with search tools to effectively expand the knowledge boundary for agentic generation.
Visual generation models are trained on fixed corpora, causing them to confidently fabricate information for unbounded, evolving user requests (e.g., new characters, trending entities). Existing benchmarks fail to capture this 40-point performance collapse, and naive search integration injects noise, degrading performance.
The authors construct SearchGen-20K, a dataset of 20,839 prompts across 12 failure categories and 22 domains, and SearchGen-Bench. They propose a 'teach-then-search' co-training framework to discover the generator-specific, evolving knowledge boundary—the divide between what can be internalized through training and what must remain external.
Frontier open generators score only 21-28/100 on SearchGen-Bench. Even a minimal version of the proposed co-training recipe yields monotonic improvement, laying the foundation for recursive self-improvement in visual generation. The full dataset, co-training corpus, and search corpus are released as a replayable harness for tool-augmented, world-knowledge-grounded visual generation.