Amiri Hayes, Belinda Li, Jacob Andreas
We propose synthesizing Python programs to approximate attention heads in transformer language models for interpretability.
Deep neural networks are opaque; understanding attention heads in transformers is challenging.
1) Collect attention matrices for each head on random inputs. 2) Prompt a pretrained LM to generate Python programs that reproduce attention patterns. 3) Rerank programs by their fit on held-out data.
On GPT-2, TinyLlama-1.1B, and Llama-3B, fewer than 1,000 programs achieve >75% IoU similarity. Replacing 25% of heads with programs increases perplexity by only 16% and maintains QA performance.