Simon Schug
Sgatlin layers, where each expert is a single linear neuron with sparse gating, replace transformer feedforward layers to improve compute efficiency and interpretability.
Existing MoE models still use large dense experts, limiting sparsity benefits, and nonlinearities hinder interpretability.
Shrink each expert to a single linear neuron, apply sparse gating over many such neurons, and remove nonlinearities to form a network of sparsely gated linear neurons.
Under isoflop comparison, sgatlin improves perplexity over standard transformer feedforward layers. Its linearity enables interpretable feedforward circuits without auxiliary models, revealing semantically clustered structures causally involved in factual recall.