Kartik Narayan, Vishal M. Patel
This is the first work to introduce a Mixture of Experts (MoE) architecture for low-resolution face recognition, addressing domain gap and catastrophic forgetting.
Low-resolution face recognition is challenging due to poor feature extraction from degradations like blur, occlusion, and low contrast. The significant domain gap between high-resolution gallery and low-resolution probe images makes it difficult for a single feature encoder to generalize across both domains when fine-tuned, leading to catastrophic forgetting.
The paper proposes FaceMoE, adapting a MoE transformer architecture for this task. It uses multiple specialized feed-forward network (FFN) experts and a top-k router to dynamically assign tokens to appropriate experts. This design promotes emergent specialization across experts for different semantic regions of the face, enabling resolution-aware feature extraction. The top-k router facilitates sparse expert activation, allowing the model to preserve pretrained knowledge when fine-tuned on a low-resolution dataset while increasing model capacity without proportional computational overhead. The model is trained with a combined face recognition loss, router z-loss, and load balancing loss.
Extensive experiments across eleven datasets spanning high-resolution, mixed-quality, and low-resolution benchmarks show that FaceMoE significantly outperforms state-of-the-art methods. The primary contribution is the novel application of MoE to low-resolution face recognition, effectively tackling the core challenges of domain gap and catastrophic forgetting.