Juyong Jiang, Fan Wang, Jiasi Shen, Sungju Kim, Sunghun Kim
A comprehensive survey that systematically organizes Code LLM research for code generation, covering data, methodology, evaluation, ethics, and more, with performance comparisons on key benchmarks.
Although LLM research for code generation is actively conducted from both NLP and software engineering perspectives, there is a lack of a comprehensive literature review that summarizes the latest trends, making it difficult for researchers to grasp the overall picture.
A taxonomy is established with six categories: data curation, latest advances, performance evaluation, ethics, environmental impact, and real-world applications. Key studies are systematically reviewed within each category. Experimental results on HumanEval, MBPP, and BigCodeBench benchmarks are collected from original papers to ensure fair comparison.
The evolution of Code LLMs is viewed from a historical perspective, and performance comparisons across various difficulty levels and task types clearly show differences between models. The gap between academia and practical development is identified, future research directions are suggested, and a continuously updated GitHub resource is provided.