Strategically Diverse Sampling for Self-Training

Alexander Gurung1, Esmeralda S. Whitammer1,2, Mirella Lapata1
1 University of Edinburgh  2 CIFAR Fellow

Self-training data is usually built by sampling a model's answers independently and keeping the correct ones (RFT). We instead ask the model for several distinct approaches first, as a tree (GROOT) or a list (Verbalized Sampling), then have it solve the problem once per approach.

IID self-training barely improves pass@k on hard problems

On frontier problems (solved ≤2/64 times), IID samples tend to be mode-collapsed and repeat one strategy; fine-tuning on this data just further reinforces this behaviour.

On hard competitive-programming problems (Figure 1), a model trained this way has almost the same pass@k as the base model at every k, even with 64 samples per training problem and a sampling temperature of 1.5.

pass@k on frontier LiveCodeBench + OJBench problems after self-training
Figure 1. Qwen3-4B-Instruct after RFT on 4 or 64 samples per training problem, on held-out frontier problems (LiveCodeBench and OJBench, averaged): problems the base model solves at most 2 times in 64. pass@k is estimated from 128 samples per problem, on the same 289 problems for every model. Nemotron3-Nano-4B is evaluated on its own frontier, so its numbers aren't comparable to Qwen's.

Asking for distinct approaches first yields more diverse training data

Unlike prior work that focuses on the training objective, we focus on how the training data itself is sampled. The model first proposes n approaches to a problem and then writes one full solution for each. We try two ways of getting the approaches:

  • GROOT (new): the model writes a decision tree of approaches, and we take n root-to-leaf paths. Any two paths split at some decision in the model's own tree.
  • Verbalized Sampling (Zhang et al., 2026): the model lists n approaches with estimated probabilities. We apply it to approaches rather than full answers.

Each approach is given to the model as a hidden instruction, and the model writes its reasoning as if it had come up with the approach itself (we discard responses that mention the instruction). The training example contains only the original problem and the solution, so at test time the fine-tuned model sees the usual prompt. Both methods add one planning call per problem, which we don't count against the sampling budget.

Sampling training data at a fixed budget
Figure 2. IID draws n independent solutions and keeps the correct ones. GROOT builds a tree of strategies and follows n distinct paths. VS lists n approaches with probabilities. In both strategic methods each approach is passed back as a hidden instruction.

Figure 3 shows samples from each method for one training problem (cobalt:14164): count the train routes [L, R] that lie inside a query range [p, q], with up to 200,000 trains and 100,000 queries. All text shown is the base model's output.

Example outputs
Qwen3-4B-Instruct
Figure 3. GROOT: the model's tree. The four highlighted leaves are the paths it chose; click one to read the approach given to the solver. VS: four approaches with the model's own probability estimates (which need not sum to 1). IID: eight independent samples at T = 0.85, all incorrect; the first, a per-query scan over every train, is shown. Each approach is only a starting point: the solver can fix or abandon it, so under each approach we note what the solver actually wrote.

Gains are driven by diversity, not correctness

As GROOT and VS find correct solutions for more training problems than IID at the same budget (155 and 137 problems with 4 samples, against 42 for IID), their gains could come from simply having more correct training data. We test this in two ways: a high-budget IID variant that finds even more correct solutions, and a dataset of diverse samples filtered to keep only the incorrect ones.

First, IID with 64 samples per problem at T = 1.5 finds correct solutions for 379 problems, more than twice as many as GROOT-4, yet the model trained on that data reaches only 7.9 held-out pass@64, about half of GROOT-4's 15.5 (Figure 4).

Correct training data vs. the trained model
Figure 4. Left: how many Cobalt training-frontier problems (1,833 for Qwen) each sampler found at least one correct solution for. Right: held-out frontier pass@64 of the model fine-tuned (RFT) on those samples. The IID-64 run produces the most correct training data but a much weaker model than GROOT-4 or VS-4. Hover a row to see its numbers.

Second, we train on only the incorrect strategic samples (ANTI). These models still outperform every IID model, including IID-64 at T = 1.5 (12.8 and 14.8 vs 7.9 held-out pass@64), and keeping the correct samples improves them further.

Training on incorrect strategic samples
Qwen3-4B-Instruct
(,)
Figure 5. Frontier pass@k after training. Dotted lines are models trained only on strategic samples that failed the tests; they still sit well above every IID-trained model. Held-out is the mean of LiveCodeBench (LCB) and OJBench. Values come from the raw samples and can differ from the paper's table by 0.1 due to rounding.

Self-generated diversity beats IID sampling from a large teacher

Sampling IID from Qwen3-235B-A22B instead of the student raises held-out pass@64 from 5.8 to 13.4. The student's own GROOT and VS samples reach 20.1 and 22.8. Running GROOT or VS with the teacher does worse than running them with the student, possibly because the teacher's traces are further from the student's distribution. With Nemotron3-Nano-4B and a Nemotron3-Super-120B teacher, the teacher improves every method, including GROOT and VS, but the student's own GROOT and VS data (19.9 and 20.8) still beat IID data from the teacher (13.8).

Self-generated vs. teacher-generated data
Figure 6. Held-out frontier pass@64 of a 4B student trained on 4 samples per problem from itself or from a larger teacher: Qwen3-235B-A22B for Qwen3-4B (16k-token limit), Nemotron3-Super-120B-A12B for Nemotron3-Nano-4B (on its own frontier).

Strategically diverse models are a better starting point for RL and test-time scaling

We also run RL from each fine-tuned model, using MaxRL (Tajwar et al., 2026). The GROOT-4 and VS-4 models begin at a validation pass@8 that RL from the base model does not reach, and RL from IID-4 needs about 64 steps to catch up.

RL from each fine-tuned model

Figure 7. Best-so-far validation pass@8 during RL, mean of three seeds ± one standard error. The red arrow marks the largest head start of GROOT-4 or VS-4 over an IID run at the same pass@8 (circles show where each curve crosses that level).

The strategically trained models also end RL ahead. On the Cobalt frontier GROOT-4 ends at pass@64 of 39.0 and VS-4 at 37.8, against 36.6 for IID-4 and 29.3 for the base model. On the held-out benchmarks GROOT-4 keeps improving (15.5 → 18.0 pass@64) while VS-4 slips (17.5 → 16.0). So we recommend GROOT as an initialization for RL, and VS when the fine-tuned model is used directly.

Before and after RL
(,)
RL steps to reach a validation pass@8 of
Figure 8. Frontier pass@k before RL (dashed) and after RL (solid, mean of three RL seeds; the shaded band is ± one standard error across seeds). Below: RL steps until validation pass@8 first reaches the chosen value. Nemotron is scored on its own frontier, so its numbers aren't comparable to Qwen's.

We also run Recursive Self-Aggregation (RSA; Venkatraman et al., 2025), which refines a population of 16 solutions for 10 rounds by merging 4 at a time. The strategic models start higher and improve more. On the Cobalt frontier VS-4 goes from 4.2 to 10.6 pass@1, and IID-4 from 0.9 to 3.8.

Test-time scaling with RSA
Qwen3-4B-Instruct
(,)
Figure 9. Frontier pass@1 of the initial population (○) and after 10 aggregation rounds (●), mean of three runs. Held-out is the mean of LCB and OJBench.

Strategic sampling for diversity also helps story planning

Next-Chapter Prediction (Gurung & Lapata, 2025) asks the model to plan the next ~200 words of a novel. A plan's score is how much it lowers the perplexity of the real passage. This task has a much larger approach space and a continuous reward (i.e., no binary correctness), making it a useful test of whether diversity transfers beyond traditional tasks. To get a pass/fail signal, we say a plan clears the bar if it improves perplexity by at least a threshold (e.g., 15%). Figure 10 shows 32 plans from each sampler for one section.

Plans for one section
Qwen3-4B-Instruct
The real passage
Figure 10. Each dot is one plan from the base model, placed by its perplexity improvement on the real passage. Filled dots clear the bar. Section funny, chapter 3, part 5 of 7.

The analogue of pass@k is coverage@k: the share of test sections where at least one of k sampled plans clears the bar. After training, the GROOT and VS models are nearly twice as likely as the base model to produce a plan that clears the 15% bar: coverage@32 of 3.6% and 4.1% vs 2.0%. The IID model reaches 2.1%, about the same as the base model. Nemotron3-Nano-4B shows the same pattern: at 16 plans, GROOT and VS reach 14.6% and 14.2% against 11.1% for the base model and 10.4% for IID.

Next-chapter prediction coverage
Figure 11. Share of the 1,430 test sections where at least one of k sampled plans improves perplexity by at least 15%. Qwen bands are ± one standard error over three training runs; the Nemotron models are single runs.

Read the full paper on arXiv. To cite it, use the BibTeX below.

Citation

@misc{gurung2026strategicallydiverse,
  title  = {Strategically Diverse Sampling for Self-Training},
  author = {Alexander Gurung and Esmeralda S. Whitammer and Mirella Lapata},
  year   = {2026},
  eprint = {2609.31571},
  archivePrefix = {arXiv},
  primaryClass = {cs.CL},
  url    = {https://arxiv.org/abs/2609.31571}
}