PAPO: Stabilizing Rubric Integration Training via Decoupled Advantage Normalization
Published:
Authors: Zelin Tan, Zhouliang Yu, Bohan Lin, Zijie Geng, Hejia Geng, Yudong Zhang, Mulei Zhang, Yang Chen, Shuyue Hu, Zhenfei Yin, Chen Zhang, and Lei Bai.
PAPO integrates outcome- and process-level feedback through decoupled advantage normalization. The method preserves correctness as the primary training signal while using rubric-based process rewards to improve reasoning quality without encouraging reward-hacking behavior.
