LLM Post-Training Stages Interfere: Order Alone Shifts Math Accuracy by 50%
Combining SFT Warm-Up With Teacher Adaptation Raises On-Policy Distillation Accuracy 50% Relative

A new preprint from Salesforce AI Research, the University of Illinois Urbana-Champaign (UIUC), and Nanyang Technological University finds that the sequence in which post-training stages are applied to a language model is not interchangeable — and that getting the order wrong can cost more than 14 percentage points of math reasoning accuracy without adding a single step of training compute. Submitted to arXiv on September 25, 2026, "Understanding the Synergy between SFT, RLVR, and OPD in LLM Post-Training" is the first systematic controlled study of how Supervised Fine-Tuning, Reinforcement Learning with Verifiable Rewards, and On-Policy Distillation interact when composed in sequence.
Why Stage Order Changes What a Model Can Learn
The three stages in question are each well-studied individually. SFT trains a model on curated (prompt, response) pairs using cross-entropy loss on fixed outputs. RLVR uses binary or scalar rewards derived from verifiable outcomes — in this case, whether a math answer is correct — to optimize the model's policy using Group Relative Policy Optimization (GRPO). On-Policy Distillation (OPD) is a knowledge transfer method where a student model generates its own rollouts and receives token-level supervision from a teacher model's full probability distribution, rather than from a fixed set of sampled responses as in standard SFT.
The central problem the paper identifies is that each stage meaningfully alters the statistical relationship between student and teacher — what the authors call student-teacher compatibility. The key metric is reverse KL divergence: how far the student's output distribution is from the teacher's. When the student's policy falls in regions where the teacher assigns very low probability, the teacher's soft-label supervision during OPD becomes informationally weak. OPD requires low reverse KL divergence at the start of distillation to propagate a reliable learning signal. What SFT and RLVR each do to that gap is the paper's core finding.
SFT Warm-Up Closes the Gap That RLVR Widens
The researchers ran experiments across nine student-teacher model pairs, ranging from two to 53 times the parameter count difference, using Qwen3-4B-Base as the primary student model and Qwen3-14B as the primary teacher, with the DAPO-Math-17K dataset for training. Evaluation used a battery of established math and science reasoning benchmarks: AIME 2024, AIME 2025, AMC 2023, MATH-500, Minerva Math, OlympiadBench, and GPQA-Diamond.
The results are unambiguous on two fronts. First, applying a brief SFT warm-up before OPD reduces the reverse KL divergence from approximately 0.40 (for an unmodified base model) to approximately 0.10–0.14. This narrow gap enables the teacher's probability distribution to provide a consistently useful learning signal across the student's actual trajectories. Second, adapting the teacher model with RLVR raises its capability, and combining that teacher adaptation with the student SFT warm-up produces the headline result: benchmark accuracy rises from 29.2% with vanilla OPD to 43.8% — a 50% relative improvement achieved within the same distillation compute budget, according to the paper's authors.
The SFT warm-up alone is not responsible for the full gain; it is the combination of improving the teacher's capability via RLVR and improving the student's learning interface via SFT that drives the 29.2% to 43.8% trajectory. By itself, a brief SFT warm-up raises OPD accuracy from 33.6% to 35.4% in the paper's primary experimental setup — still a meaningful improvement, but the dramatic full gain requires both interventions.
The effect runs sharply in reverse when RLVR is applied to the student before OPD. RLVR pushes the policy toward high-reward solutions, collapsing output entropy and concentrating probability mass on the reward-favored answers. This moves the student's post-RLVR trajectories away from the teacher's distribution, raising reverse KL divergence back toward 0.40. The researchers find that RLVR-first pipelines leading into OPD regress rather than improve: the student unlearns some of what RLVR taught because OPD pulls it toward a teacher it has diverged from. SFT creates alignment between student and teacher distributions; RLVR on the student destroys it when the teacher has not been similarly updated.
OPD-Initialized Models Outperform SFT Starts in Downstream Reinforcement Learning
Read more: TrimSFT gains 26.9 points on MATH500 by targeting intermediate-confidence tokens
The paper's third major finding concerns what happens after distillation, when the model continues into additional RLVR runs. The researchers compared two initialization strategies: a model prepared for RLVR via SFT and one prepared via OPD, starting from matched accuracy baselines. After 260 steps of RLVR, OPD-initialized models reached 38.5–43.1% accuracy across the benchmark suite, versus 33.4–36.8% for SFT-initialized models — and the paper's authors note the separation grows over the observed horizon.
The mechanism is policy entropy. OPD trains the student against the teacher's full token-level probability distribution, conveying not just the correct answer but the teacher's uncertainty across alternative tokens. This preserves higher policy entropy than SFT, which trains on hard-labeled responses and flattens the policy toward the teacher's dominant modes. Higher entropy at the start of RLVR gives the optimizer more exploration space: RLVR reinforces correct-reward rollouts from a diverse distribution, and a narrow, low-entropy starting point reduces the ceiling of what RL can discover. The paper notes this gap widens with more RL compute — meaning OPD's advantage is not a one-time offset but one that compounds as labs scale their RL post-training budgets.
How the Findings Connect to Frontier Post-Training Pipelines
Every major frontier AI lab in 2026 runs multi-stage post-training pipelines, and the paper's findings sit in direct conversation with several active architectures.
Alibaba's Qwen3 uses a four-stage pipeline that incorporates SFT followed by RLVR and then distillation to consolidate RL-trained specialist models into a general model. That consolidation step is structurally consistent with the paper's recommendation: SFT warm-up before distillation.
DeepSeek-R1 uses cold-start SFT on chain-of-thought reasoning traces, followed by RLVR, then a secondary SFT-style distillation step using rejection-sampled data. The cold-start SFT is exactly the warm-up the paper advocates. Its subsequent distillation stage uses an SFT-style approach — training on fixed sampled outputs rather than on-policy rollouts with teacher soft labels — which means the consolidation stage may be a less optimal downstream RL initialization compared to a full OPD equivalent.
NVIDIA's Nemotron-Cascade 2 uses an SFT stage followed by cascade RL, with multi-domain OPD inserted throughout the RL process to recover benchmark regressions as the model passes through different training domains. This interleaved use of OPD within RL is structurally consistent with the paper's finding that OPD works best when student-teacher compatibility is maintained — the cascade architecture tries to ensure that compatibility is never fully lost.
The multi-teacher OPD (MOPD) pattern — multiple RL-specialized models distilling into a single student — is emerging as a preferred consolidation architecture. The paper's analysis implies MOPD effectiveness depends critically on whether the student has had adequate SFT warm-up before distillation begins, regardless of teacher count.
Read more: Meta FAIR study predicts byte-level distillation beats token models as compute scales
What Remains Untested and Where Caution Applies
The study is a preprint and has not yet undergone peer review. Its findings are based entirely on the Qwen3 model family evaluated on mathematics and science reasoning benchmarks. Whether the same stage-ordering effects appear with architectures from different model families, on coding tasks, on multimodal pipelines, or with distillation methods other than the specific OPD variant the paper tests, is not yet established.
The benchmark suite — AIME, MATH-500, and GPQA-Diamond among others — is credible and widely used in reasoning model evaluation, but is entirely concentrated in formal quantitative reasoning. Language tasks, instruction following, and general knowledge benchmarks are absent from the paper's experimental scope.
The paper also tests model pairs ranging from two to 53 times the parameter ratio. The 53× case is an extreme scenario unlikely to reflect typical deployment decisions, where student and teacher models are usually within a factor of 5–10 of each other. The paper does not distinguish which findings are robust across parameter ratio ranges and which are specific to particular ratio regimes.
All numerical results — including the 50% relative improvement figure — are reported by the paper's authors and have not been independently replicated by external researchers at the time of writing. Independent replication, particularly on non-Qwen model families and on tasks beyond math, will be essential before these stage-ordering recommendations can be treated as general engineering principles.
What Post-Training Engineers Should Watch Next
The paper's most practically significant contribution may be the framing itself: post-training stage composition is not a neutral engineering choice but a source of performance variation that can be studied, measured, and optimized. Previous work treated SFT, RLVR, and OPD as independent tools; this analysis establishes that they are co-dependent, with each stage altering the conditions under which the next stage operates.
If independent replication confirms the OPD-initialization advantage — particularly the finding that this advantage widens with more RL compute — the implication for frontier AI development is substantial. Labs currently using SFT-based consolidation in multi-stage pipelines would have a mechanistically grounded reason to evaluate OPD as a consolidation method, not because it is conceptually attractive but because it preserves the policy entropy that subsequent RL rounds require. The question of whether the same entropy-preservation advantage extends to preference optimization pipelines, or to systems that incorporate synthetic data distillation alongside RL, would be a natural extension of this work.