Pietro Barbiero, Giovanni De Felice, Mateo Espinosa Zarlenga, Francesco Giannini, Filippo Bonchi, Mateja Jamnik, Giuseppe Marra, Ruggero Noris
A framework based on Lagrangian mechanics that deductively designs interpretable machine learning methods using a general theory.
Interpretability research suffers from fragmented methodologies and inconsistent evaluation protocols due to the lack of a general theory.
From user-defined interpretability premises, derive symmetries and constraints, then find optimal interpretable models via minima of a Lagrangian. Identify limitations of existing methods and propose new research directions.
Identifies and addresses limitations of existing methods (traditional, concept-based, mechanistic interpretability), contributes to interpretability education and research community.