Gradient Conflict, Not Coverage Gap, Breaks Cross-Tokenizer Distillation
Alibaba Cloud Researchers Show Strict-Position Training With Top-16 Reverse KL Beats Complex Alignment

Researchers at Alibaba Cloud Computing have published findings that challenge a central assumption driving the past two years of cross-tokenizer on-policy distillation research: that recovering supervision from tokenizer-mismatched positions makes training better. In a preprint posted to arXiv on October 6, 2026, the team reports that the opposite is true — adding span-based supervision across mismatched token groups consistently lowers accuracy, and the reason is gradient conflict rather than any deficiency in the alignment strategy itself. The paper, "Rethinking Cross-Tokenizer On-Policy Distillation," is available on arXiv.
The result carries direct implications for practitioners and researchers building distillation pipelines across incompatible model families — a setting that has become operationally unavoidable as Qwen, Phi, Llama, Gemma, and IBM's Granite models each use proprietary tokenizers and cannot be paired for distillation using conventional vocabulary-matching methods. The practical stakes are significant: the ability to transfer reasoning capability from a large frontier teacher into a smaller, domain-adapted student underpins a substantial share of enterprise model deployment, and the teacher and student increasingly come from different organizational lineages with wholly incompatible vocabularies.
The Problem Cross-Tokenizer OPD Must Solve
On-policy distillation (OPD) trains a student model on its own sampled outputs while a larger teacher model provides token-by-token feedback, typically as a KL divergence signal. This approach, established by MiniLLM in 2023, avoids a core problem with standard knowledge distillation: when students train on teacher-generated text, they learn to recover from teacher errors but not their own, a compounding failure known as exposure bias that grows roughly with the square of sequence length. OPD eliminates that asymmetry by always grounding supervision on the student's own rollouts.
The difficulty arrives when teacher and student use different tokenizers. A tokenizer is more than a lookup table — it determines how any string of text is divided into discrete tokens, which means the same sentence can yield a completely different sequence of prediction steps depending on the model. When a Granite teacher processes the word "calculation," it might produce a single token; a Phi-4-mini student might split the same word across two or three. At that boundary, the teacher's next-token distribution and the student's next-token distribution are no longer describing the same prediction problem. The teacher's probability mass is spread across completions the student will never see in its vocabulary, and vice versa.
Quantitatively, the vocabulary intersection between major model families is often surprisingly small. The static Jaccard overlap — the fraction of tokens that appear in both vocabularies — sits between roughly 39 and 65 percent across the model pairs studied in the focal paper. In principle, that leaves a substantial portion of each generated sequence with no direct teacher supervision under naïve shared-vocabulary matching. The question the Alibaba researchers set out to answer was whether recovering that missing supervision is worth the added complexity — and whether the recovery methods published over the past two years were actually helping.
The field's dominant response since 2024 has been: find ways to recover that lost supervision. Align at the byte level. Build projection matrices. Use optimal transport. Find the finest granularity at which both tokenizers agree. Each successive method increased the fraction of teacher signal that reached the student.
What the Alibaba Study Actually Found
The Alibaba Cloud team ran experiments across three heterogeneous teacher-student pairs — including IBM Granite as teacher distilled into Microsoft Phi and Alibaba's own Qwen — on mathematical reasoning and code generation benchmarks. Their first finding cuts against the field's premise: strict 1:1 alignment groups, positions where a single teacher token and a single student token happen to cover the same text interval, already cover the vast majority of tokens in student-generated sequences, even though static vocabulary Jaccard overlap between model families ranges from only about 39 to 65 percent. The reason is that actual inference-time output skews heavily toward tokens both tokenizers handle the same way. Common numerals, punctuation, frequent English subwords, and structural tokens such as newlines and code delimiters appear in the shared intersection and dominate real generated text. Table 1 of the paper shows this strict-coverage finding across all tested model pairs. The theoretical mismatch rate is high; the practical mismatch rate during rollout is low.
The paper then tests two variants of the distillation loss. The strict loss applies reverse KL divergence over the shared vocabulary at strictly aligned positions only. The span loss adds a mean squared error term on the log-probabilities of observed token paths in mismatch groups. The combined objective weights them as L₁:₁ + λ · L_span, with λ controlling how much the mismatch supervision contributes.
The accuracy results, measured across math and code benchmarks, consistently show the same pattern: any positive value of λ reduces performance relative to strict-only distillation (λ = 0). The span loss does not just fail to help — it actively makes training worse.
Gradient Conflict as the Mechanism
The paper's most significant contribution is not the performance comparison but the gradient-level diagnosis of why the span loss hurts. The authors compute directional agreement between the gradients produced by the span loss and the gradients produced by the strict loss at checkpoints during training. Directional agreement here means whether the two gradient vectors point in roughly the same direction — whether adding the span supervision pushes model parameters in a way that is consistent with what the strict loss is already doing.
The span-loss gradients show weak or negative directional agreement with the strict-loss gradients throughout training. More troublingly, the relative magnitude of span gradients grows over the course of training, meaning the conflicting signal becomes proportionally louder as the student improves on the strict-position objective. This is not a case where the span supervision provides noisy but compatible signal; it is providing signal that structurally opposes the reliable strict-position learning.
The finding connects to a broader multi-objective training problem identified in the gradient surgery literature: when two loss components produce gradients that conflict, adding the second component can actively degrade the optimization trajectory even if that component has a well-motivated loss function in isolation. The span log-probability MSE is a reasonable thing to minimize in principle; in practice, the teacher's span-level probability estimates are computed over text intervals where the teacher's tokenization diverges from the student's, making the supervision noisy enough to conflict with the tighter signal available at strict positions.
A Simpler Approach That Matches the Best Current Methods
The paper proposes a practical simplification that outperforms four cross-tokenizer baselines: restrict the reverse KL at strict positions to the student's top-16 highest-probability tokens in the shared vocabulary, rather than computing it over all shared tokens. This top-k restriction has a natural motivation — the most information in the teacher signal comes from the distribution of likely next tokens, and very low-probability shared tokens contribute gradient noise out of proportion to their contribution to the learned distribution. At k = 16, performance matches strict full-vocabulary distillation and outperforms the baselines on math reasoning accuracy for all three model pairs.
The economics of this approach matter for deployment. Methods like SimCT, which builds minimal aligned units from the finest text granularity both tokenizers can express, and ESCD, which models event-set completions after partial entry into a teacher token group, both require additional data structures and inference-time computation per training step. The Alibaba approach requires only standard shared-vocabulary lookup plus a top-k sort — work that sits comfortably in the same computational budget as single-tokenizer OPD.
The Broader Competitive Context in 2026
The cross-tokenizer distillation problem has attracted at least six distinct research approaches in the past two years, each with a different intuition about what "alignment" should mean.
Optimal transport methods like MultiLevelOT (2024) and the Universal Logit Distillation approach frame the problem as matching probability distributions without requiring token-by-token correspondence, treating alignment as a global optimization problem. These methods showed strong results on question answering and summarization but were primarily tested outside the on-policy training regime. Approximate likelihood matching (ALM, March 2025) formalized the cross-tokenizer setting as likelihood alignment across any pair of tokenizers and demonstrated competitive performance on tokenizer transfer tasks including distillation from math-specialized models to general-purpose small models.
The byte-level distillation approach (BLD, April 2026) took a different path: rather than aligning at the subword level, it routes knowledge transfer through a shared byte-level interface by appending a lightweight byte-level decoder head to the student. The authors reported competitive results but acknowledged inconsistent gains across benchmarks. X-Token (May 2026) introduced a sparse projection matrix that maps each student token to weighted teacher tokens, choosing between two loss formulations — a full-projection mode and a hybrid near-equivalent-matching mode — depending on whether critical token types fall outside the common vocabulary.
Read more: Meta FAIR study predicts byte-level distillation beats token models as compute scales
SimCT and ESCD represent the most recent generation of methods before the focal paper. SimCT argued that strict shared-token matching discards a large fraction of the teacher signal at mismatched positions, and proposed supervision at the level of minimal aligned multi-token units — the finest shared text continuations both tokenizers can express. ESCD went further, arguing that the event-completion gap — what happens when a student token only partially realizes a teacher token group — requires explicit modeling of the valid completion set, weighted by teacher probability mass. ESCD reported substantial gains on code generation benchmarks using a Qwen3-32B teacher, including a jump on LiveCodeBench pass@2 from 18.1 to 42.9 percent.
The Alibaba paper's relationship to ESCD is particularly instructive. ESCD validated its approach using a gradient agreement metric (COUF) that measures how much completion supervision aligns with the base distillation signal — and reported that keeping the event-completion loss raises COUF from 0.7383 to 0.8495 for the Qwen pair. The focal paper uses the same logic from the other direction: its span-loss gradients show poor directional agreement with strict-loss gradients. Both papers agree on the underlying principle. Where they diverge is on whether any mismatch-group supervision can be made gradient-compatible. ESCD argues yes, by carefully modeling completion probabilities at the right level of granularity. The Alibaba paper argues that the MSE span supervision commonly used in the field does not meet that bar, and that the gradient conflict is severe enough that strict-only training outperforms the augmented approach.
Where the Method Falls Short
Code generation accuracy presents a more mixed picture than mathematical reasoning in the paper's reported results. Strict full-vocabulary OPD held at or near the best result on code tasks, but did not achieve the same consistent margin over baselines that appeared on math benchmarks. The authors do not attribute this to a specific mechanism. One plausible explanation is that code generation places higher demands on vocabulary coverage at mismatch positions — code token boundaries differ more systematically across model families for variable names, indentation, and symbolic operators than for natural language text. If mismatch-group tokens carry more semantic weight in code than in math, the strict-only approach may be trading a form of coverage that matters more in that domain.
The paper also does not provide a method for improving mismatch-group supervision — only evidence that the current MSE span loss is counterproductive. The three model pairs tested (all involving IBM Granite as teacher) leave open how the findings transfer to other architecture combinations. Granite-to-Phi and Granite-to-Qwen share specific structural properties, and a paper examining, say, Llama-to-Gemma or Mistral-to-DeepSeek would provide broader validation. No code repository was linked in the arXiv listing as of October 8, 2026, making independent reproduction dependent on the paper's methodological descriptions alone.
The Design Principle the Field Is Converging On
Across the six competing approaches, a pattern is emerging that the Alibaba paper makes explicit: gradient quality is the correct evaluation criterion for supervision components in distillation, not supervision breadth. ESCD's COUF metric and the focal paper's directional agreement analysis are measuring the same underlying property — whether an auxiliary loss component reinforces or conflicts with the primary learning signal. Methods that can demonstrate gradient alignment, as ESCD does for completion supervision, earn the right to add coverage. Methods that cannot demonstrate it, as the focal paper shows for MSE span supervision, should not add coverage regardless of how intuitive the motivation seems.
This principle has implications beyond cross-tokenizer OPD. Post-training pipelines increasingly combine multiple loss components — supervised fine-tuning, RL from verifiable rewards, OPD, and auxiliary regularization terms — often without examining gradient directional agreement between them. Recent work on post-training stage ordering has already shown that the sequence in which these components are applied changes performance dramatically. The gradient-quality lens suggests that the composition problem has a measurable diagnostic: if two loss components produce consistently misaligned gradients, one of them should either be redesigned or dropped.
For practitioners currently weighing cross-tokenizer distillation methods for production use, the immediate takeaway is specific: a strict-only top-16 reverse KL approach is simpler to implement, computationally equivalent to standard OPD, and reportedly competitive with or better than more elaborate alignment schemes on math reasoning tasks. The method requires no auxiliary data structures, no learned projection matrices, no byte-level decoder heads, and no multi-step alignment preprocessing — it is, in the authors' framing, the minimum viable cross-tokenizer OPD implementation, and their claim is that it has been beating more complex alternatives without researchers realizing it. Whether that holds on the code benchmarks that matter most to engineering teams, and whether it generalizes beyond the specific Granite teacher architecture, remains the open empirical question this paper leaves for the field to resolve.
Read more: LLM post-training stages interfere — order shifts math accuracy by 50%