Reward-Free Policy Optimization
Recent approaches to RL post-training for large language models increasingly remove the critic to reduce training instability and memory overhead. Even where a critic is trained, it is discarded once training ends, although it has learned to predict outcomes. We revisit this trend and show that a pretrained critic's ability to predict future outcomes can make it a valuable asset for efficient long-horizon reasoning.
The frozen critic's value at the last token
Share of rollouts per value bin, correct above the axis, incorrect below.
Instability in critic-based RL for long chain-of-thought reasoning is largely an optimization artifact. Keeping policy updates small and low in variance restores stable convergence, and the critic becomes what it was trained to be: a predictor of eventual success.
Contributions (i) and (ii) ↓A single calibrated, frozen critic serves as the reward, the GAE baseline and the forecaster. Training needs no completed rollouts, no step-level annotations and no external reward labels, and the binarized critic matches supervised PPO.
Contributions (ii) and (iii) ↓The critic scores unfinished rollouts with accuracy comparable to complete ones, so rollouts can be rewarded before they finish. This suits long-horizon reasoning, where outcomes arrive late and generation dominates cost.
Contribution (iv) ↓Unlock the critic · contribution (i)
Recent approaches to RL post-training increasingly remove the critic to reduce training instability and memory overhead. We find that this instability is largely an optimization artifact: it comes from the update recipe, not from the value network, and it recedes once each policy update is kept small and low in variance. Under this recipe, critic-based training converges stably, without length collapse or entropy collapse.
Unlock the critic, reward-free · contribution (ii)
A well-pretrained critic estimates the posterior probability of eventual success from later trajectory states and unfinished prefixes. Its predictions therefore provide outcome-derived, dense, per-prefix learning signals that need neither completed rollouts, step-level annotations, nor external reward labels. RFPO repurposes a single calibrated, frozen critic as the rollout-level reward, as the value baseline for generalized advantage estimation, and as a success forecaster for unfinished prefixes. The RL loop contains no verifier and trains no value network.
If a critic is good enough to be trusted as a baseline, it is also good enough to be the reward.
The value histogram at the top of this page is why this works: the frozen critic pushes incorrect rollouts toward 0 and correct ones toward 1 (AUC 0.96 on rollouts from the converged policy). It ranks attempts at the same problem just as well when the attempt never finished, so the reward is not simply detecting truncation. From a prefix alone, within-problem AUC reaches 0.84 at 4,096 tokens, and 0.84–0.92 on prefixes that have not produced an answer yet.
Ranking attempts at the same problem
Within-problem AUC of the frozen critic, averaged over ten policy checkpoints (16,384 rollouts each, 19–54% truncated).
Truncated attempts are ranked as well as the full set. Finished-only is lower because a problem that mixes finished and truncated attempts offers easy contrasts, and removing the truncated attempts removes them.
Ranking from a prefix
Within-problem AUC when the critic sees only the first n tokens.
Reward-free, safely · contribution (iii)
The critic's raw score carries a length bias, and a continuous reward lets the policy exploit it in either direction. Binarizing the debiased score closes that channel. Used as a continuous reward, the same critic first climbs above fully supervised PPO, which shows how much it knows, and then gives the gain back as response length drifts. Binarizing trades that ceiling for stable training.
Response length drifts under a continuous reward
Change in mean training response length, last step against first, by debiasing strength. The shaded band is the range the binarized runs stay within.
Don't wait for the end · contribution (iv)
Outcome rewards wait for the last token, so the longer the reasoning, the more training pays for waiting. A critic answers earlier.
Binarized, RFPO matches supervised PPO without a single label in the training loop, while substantially cutting compute and memory overhead. Because rollouts can be rewarded before they finish, training no longer has to pay for waiting on every trajectory to complete. Even when a tight generation cap leaves a large share of every batch unfinished, RFPO validates above supervised PPO trained under the same cap, at a lower cost.
Validation under a 4,096-token training cap
Macro-average of AIME 2026 and AMC 2023 at 12,288 tokens. Dots: single evaluations; lines: centered means of five.
Share of training rollouts truncated
PPO shortens its answers to fit the cap; RFPO keeps about half unfinished.
Median seconds per RL step
Same node, same cap. The gap is PPO's critic update.
At the standard 5,120-token cap
Separate 16-sample evaluation after 300 RL steps, pass@1 (%)
| Model | AIME25 | AIME26 | AMC23 | GPQA | Avg |
|---|---|---|---|---|---|
| Initial policy | 22.7 | 21.0 | 71.4 | 44.9 | 40.0 |
| PPO, 100% labels | 25.2 | 21.7 | 73.6 | 46.6 | 41.8 |
| RFPO, 50% labels | 22.7 | 21.2 | 72.8 | 48.5 | 41.3 |
| RFPO, 0% labels | 25.4 | 21.7 | 70.0 | 47.3 | 41.1 |
Level or ahead on AIME and GPQA, behind on AMC 2023 (40 problems, 2.5 points each). An RL step takes 421 s, against 582 s for PPO.
As reasoning traces lengthen and agentic episodes stretch, waiting for the outcome buys less and less for what it costs, and the value network is the one component of the standard recipe that speaks before the outcome arrives. So far our evidence comes from mathematical reasoning with one 4B model family and traces up to 12,288 tokens; larger models and other tasks are next.
The question is not whether a critic is affordable, but how much of what it already knows the current recipe throws away.
Our findings challenge the prevailing critic-free paradigm and establish critic-based, reward-free optimization as a scalable and computationally efficient path for LLM post-training.
Authors
Cite
@article{li2026unlocking,
title = {Unlocking the Critic: Reward-Free Policy
Optimization for LLM Post-Training},
author = {Li, Hongyang and Li, Xiao and Wu, Caesar and
Mammar, Said and Danoy, Gr{\'e}goire and Bouvry, Pascal},
journal = {arXiv preprint arXiv:2609.37119},
year = {2026},
url = {https://arxiv.org/abs/2609.37119}
}