Ananth Eswar, Pratinav Seth, Utsav Avaiya, Vinay Kumar Sankarapu
Proposes a causal audit framework to directly test the validity of attribution scores that identify important neuron rows in language models.
There is a lack of methods to directly test whether neuron rows identified as important by attribution scores are truly causally important or merely correlated with a model's capability.
Conducts two audits using one-shot neuron-row zeroing. First, evaluates if attribution methods better identify dispensable rows at the language modeling level. Second, tests if the identified rows are sufficient to install refusal behavior using a contrastive harmful-versus-benign signal.
Attribution methods substantially outperform baselines at identifying dispensable rows. The attributed rows are sufficient to install refusal on hate and crime while keeping benign over-refusal low, and layer-matched random controls fail. Crucially, highly rank-stable selectors can be among the least causally valid. Refusal lives in a redundant subspace, where different attribution methods install it through largely disjoint row sets. This demonstrates that rank-stability proxies miss the kinds of selector failures a direct causal audit can surface.