Self-training data is usually built by sampling a model's answers independently and keeping the correct ones (RFT). We instead ask the model for several distinct approaches first, as a tree (GROOT) or a list (Verbalized Sampling), then have it solve the problem once per approach.
On frontier problems (solved ≤2/64 times), IID samples tend to be mode-collapsed and repeat one strategy; fine-tuning on this data just further reinforces this behaviour.
On hard competitive-programming problems (Figure 1), a model trained this way has almost the same pass@k as the base model at every k, even with 64 samples per training problem and a sampling temperature of 1.5.
Unlike prior work that focuses on the training objective, we focus on how the training data itself is sampled. The model first proposes n approaches to a problem and then writes one full solution for each. We try two ways of getting the approaches:
Each approach is given to the model as a hidden instruction, and the model writes its reasoning as if it had come up with the approach itself (we discard responses that mention the instruction). The training example contains only the original problem and the solution, so at test time the fine-tuned model sees the usual prompt. Both methods add one planning call per problem, which we don't count against the sampling budget.
Figure 3 shows samples from each method for one training problem (cobalt:14164): count the train routes [L, R] that lie inside a query range [p, q], with up to 200,000 trains and 100,000 queries. All text shown is the base model's output.
As GROOT and VS find correct solutions for more training problems than IID at the same budget (155 and 137 problems with 4 samples, against 42 for IID), their gains could come from simply having more correct training data. We test this in two ways: a high-budget IID variant that finds even more correct solutions, and a dataset of diverse samples filtered to keep only the incorrect ones.
First, IID with 64 samples per problem at T = 1.5 finds correct solutions for 379 problems, more than twice as many as GROOT-4, yet the model trained on that data reaches only 7.9 held-out pass@64, about half of GROOT-4's 15.5 (Figure 4).
Second, we train on only the incorrect strategic samples (ANTI). These models still outperform every IID model, including IID-64 at T = 1.5 (12.8 and 14.8 vs 7.9 held-out pass@64), and keeping the correct samples improves them further.
Sampling IID from Qwen3-235B-A22B instead of the student raises held-out pass@64 from 5.8 to 13.4. The student's own GROOT and VS samples reach 20.1 and 22.8. Running GROOT or VS with the teacher does worse than running them with the student, possibly because the teacher's traces are further from the student's distribution. With Nemotron3-Nano-4B and a Nemotron3-Super-120B teacher, the teacher improves every method, including GROOT and VS, but the student's own GROOT and VS data (19.9 and 20.8) still beat IID data from the teacher (13.8).
We also run RL from each fine-tuned model, using MaxRL (Tajwar et al., 2026). The GROOT-4 and VS-4 models begin at a validation pass@8 that RL from the base model does not reach, and RL from IID-4 needs about 64 steps to catch up.
The strategically trained models also end RL ahead. On the Cobalt frontier GROOT-4 ends at pass@64 of 39.0 and VS-4 at 37.8, against 36.6 for IID-4 and 29.3 for the base model. On the held-out benchmarks GROOT-4 keeps improving (15.5 → 18.0 pass@64) while VS-4 slips (17.5 → 16.0). So we recommend GROOT as an initialization for RL, and VS when the fine-tuned model is used directly.
We also run Recursive Self-Aggregation (RSA; Venkatraman et al., 2025), which refines a population of 16 solutions for 10 rounds by merging 4 at a time. The strategic models start higher and improve more. On the Cobalt frontier VS-4 goes from 4.2 to 10.6 pass@1, and IID-4 from 0.9 to 3.8.
Next-Chapter Prediction (Gurung & Lapata, 2025) asks the model to plan the next ~200 words of a novel. A plan's score is how much it lowers the perplexity of the real passage. This task has a much larger approach space and a continuous reward (i.e., no binary correctness), making it a useful test of whether diversity transfers beyond traditional tasks. To get a pass/fail signal, we say a plan clears the bar if it improves perplexity by at least a threshold (e.g., 15%). Figure 10 shows 32 plans from each sampler for one section.
The analogue of pass@k is coverage@k: the share of test sections where at least one of k sampled plans clears the bar. After training, the GROOT and VS models are nearly twice as likely as the base model to produce a plan that clears the 15% bar: coverage@32 of 3.6% and 4.1% vs 2.0%. The IID model reaches 2.1%, about the same as the base model. Nemotron3-Nano-4B shows the same pattern: at 16 plans, GROOT and VS reach 14.6% and 14.2% against 11.1% for the base model and 10.4% for IID.
Read the full paper on arXiv. To cite it, use the BibTeX below.
@misc{gurung2026strategicallydiverse,
title = {Strategically Diverse Sampling for Self-Training},
author = {Alexander Gurung and Esmeralda S. Whitammer and Mirella Lapata},
year = {2026},
eprint = {2609.31571},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2609.31571}
}