Korbinian Friedl, Francis Rhys Ward, Paul Yushin Rapoport, Tom Everitt, Jonathan Richens
A theoretical study proving that it is fundamentally impossible to train AI systems to honestly report their latent knowledge.
The problem of designing advanced AI systems to honestly report their beliefs about latent variables in the environment that are hidden from humans. This formalizes the problem of eliciting latent knowledge (ELK) and analyzes its limitations.
Uses Causal Influence Diagrams (CIDs) to model the relationship between an agent's training environment and its subjective representation of the world. Formalizes the distinction between observable and latent variables, defines honesty, and defines goal misgeneralization using CIDs. Presents an impossibility theorem proving that no behavior-based training strategy can guarantee an honest agent even under perfect feedback.
Proves an impossibility theorem that no feedback-based training strategy can guarantee an honest agent. This presents a fundamental theoretical limitation for AI alignment research and suggests the need for new methods beyond behavior-based approaches.