Scaling Behaviors of LLM Reinforcement Learning Post-Training: An Empirical Study in Mathematical Reasoning
Published in ACL 2026 Main Conference, 2025
Authors: Zelin Tan, Hejia Geng, Xiaohang Yu, Mulei Zhang, Guancheng Wan, Yifan Zhou, Qiang He, Xiangyuan Xue, Heng Zhou, Yutao Fan, Zhong-Zhi Li, Zaibin Zhang, Guibin Zhang, Chen Zhang, Zhenfei Yin, Philip Torr, and Lei Bai.
This empirical study examines how model scale, data volume, and compute interact during reinforcement-learning post-training for mathematical reasoning. It identifies predictive scaling behavior, a latent saturation trend in learning efficiency, and the value of repeated high-quality data in constrained regimes.
Links: ACL Anthology · arXiv · DOI
