TL;DR
This paper proposes a framework and benchmarks for quantitatively evaluating the interpretability of sparse autoencoders (SAEs) using human-annotated concepts.
Problem
Sparse autoencoders (SAEs) are widely used to extract interpretable concepts from vision and vision-language models, but existing evaluation methods largely rely on proxy metrics or qualitative inspection, failing to directly measure semantic correspondence.
Approach
Human-grounded evaluation framework: A framework is presented to quantify alignment between SAE latents and human-annotated concepts.
synCUB and synCOCO benchmarks: Synthetic benchmarks of paired images differing in exactly one attribute are constructed to enable intervention-style evaluation.
Fully-Binary Matching Pursuit (FBMP): A coalition-based matching procedure is introduced that supports many-to-one mappings between SAE latents and annotated concepts.
Targeted Attribute Perturbation Alignment Score (TAPAScore): A functional validation score is proposed to test whether matched concepts respond selectively and in the expected direction under targeted image-level attribute perturbations.
Results & Contribution
FBMP consistently outperforms one-to-one baselines.
Our matching procedure and TAPAScore are the only evaluated metrics that reliably distinguish trained SAEs from untrained ones under sanity checks.