Scaling Behaviors of LLM Reinforcement Learning Post-Training: An Empirical Study in Mathematical Reasoning

Published in ACL 2026 Main Conference, 2025

Authors: Zelin Tan, Hejia Geng, Xiaohang Yu, Mulei Zhang, Guancheng Wan, Yifan Zhou, Qiang He, Xiangyuan Xue, Heng Zhou, Yutao Fan, Zhong-Zhi Li, Zaibin Zhang, Guibin Zhang, Chen Zhang, Zhenfei Yin, Philip Torr, and Lei Bai.

This empirical study examines how model scale, data volume, and compute interact during reinforcement-learning post-training for mathematical reasoning. It identifies predictive scaling behavior, a latent saturation trend in learning efficiency, and the value of repeated high-quality data in constrained regimes.

Links: ACL Anthology · arXiv · DOI