Komal Kumar, Tajamul Ashraf, Omkar Thawakar, R. Anwer, Hisham Cholakkal, Mubarak Shah, Ming-Hsuan Yang, P. Torr et al.
A systematic survey of LLM post-training techniques, analyzing key methodologies such as fine-tuning, reinforcement learning, and test-time scaling, and presenting future directions.
Pretraining alone is insufficient for LLMs' reasoning ability, factual accuracy, and alignment with user intent. Challenges such as catastrophic forgetting, reward hacking, and trade-offs between inference time and performance must be addressed during post-training.
Various post-training strategies including fine-tuning, reinforcement learning (especially RLHF), and test-time scaling are categorized and analyzed for their pros and cons. Key research directions such as model alignment, scalable adaptation, and inference-time reasoning are summarized.
Provides a comprehensive taxonomy of post-training methodologies, highlights core challenges and future research directions, and offers a publicly maintained repository for ongoing updates.