A. Torre, Barbara M. Flores, Diego Rodríguez
We propose an efficient method to compress large language models by combining knowledge distillation with chain-of-thought reinforcement learning, showing that small models retain most of the teacher's performance.
Large language models exhibit outstanding performance but require enormous computational resources and memory, making deployment difficult in resource-constrained environments. Existing knowledge distillation alone fails to fully transfer complex reasoning abilities.
Using Qwen 3B as teacher and Qwen 0.5B as student, we perform knowledge distillation on English Dolly-15k, Spanish Dolly-15k, and code BugNet and PyTorrent datasets. For code tasks, we integrate chain-of-thought prompting with Group Relative Policy Optimization using CoT-annotated Codeforces data to improve reasoning coherence and correctness. Post-training 4-bit weight quantization further reduces memory footprint and inference latency.
On English tasks, the distilled student retains 70-91% of teacher performance; up to 95% on Spanish; and up to 93.5% Rouge-L on code. Code tasks with chain-of-thought reinforcement learning show better reasoning coherence and correctness than knowledge distillation alone. The proposed framework provides practical small models for resource-constrained environments.