Xumin Yu, Zuyan Liu, Zhenyu Yang, Yuhao Dong, Shengsheng Qian, Jiwen Lu, Han Hu, Yongming Rao
We propose ViQ, a framework that unifies text and visual information into a single discrete representation, simplifying multimodal modeling and improving training efficiency.
Existing discrete visual representations struggle to balance semantics and fine details. Reconstruction-oriented representations often lack semantic information, while semantically stronger features suffer from severe detail loss. They are also often limited to fixed-resolution inputs.
ViQ structures quantization learning into two stages. First, it performs text-aligned pre-training of the visual encoder under supervision from a pretrained language model to handle native-resolution inputs. Second, during feature discretization, it proposes a proximal representation learning strategy and a position-aware head-wise quantization mechanism to progressively compact the feature space and enable flexible resolution processing.
ViQ achieves competitive performance compared to state-of-the-art multimodal vision encoders with continuous, high-dimensional features on various multimodal tasks, while maintaining high precision in low-level reconstruction. Training with visual quantized representations yields up to 20%-70% acceleration with different base LLMs and training recipes, significantly improving efficiency.