Shilong Xiang, Zirui Zhang, Chengzhi Mao
A self-supervised learning framework where VLLMs autonomously hypothesize, validate, and reinforce visual concepts without labels.
Self-supervised learning scales to large data but yields opaque features, while interpretable models rely on human annotations. A method to discover interpretable concepts without labels is needed.
We design a self-supervised preference optimization loop that uses VLLMs as active participants, not static feature extractors, to autonomously hypothesize, validate, and reinforce candidate visual attributes from raw images. Concept discovery is amortized into the VLLM backbone via a self-supervised preference objective.
We successfully extract domain-specific concepts in natural, medical, and physics domains where standard VLLMs fail. We achieve up to 24-point absolute improvement in downstream top-1 classification accuracy on unseen data. We show that interpretability can emerge through a model's autonomous interaction with incidental visual structures, without any human supervision.