Can online data selection improve training?
Training data are usually chosen before a run. This proposal asks whether a model can help decide what it should train on next. Start with more candidate data than the budget can consume, arrange it into reproducible shards, and measure the model after each shard.
If a shard produces an unusually large, unrecovered increase in a held-out domain loss, roll back the model. A first rejection defers that shard to a later rotation of the queue. A second rejection sends it to quarantine for inspection. Data quality becomes a traceable decision with a reason and a chance to be reconsidered.
Version 2 reports no results for this mechanism. Its “reviewed” status is the repository’s author-review status, not journal peer review. The figures below describe the proposed design.
Source: Abstract and §1 ↗.
From example selection to whole-shard decisions.
The proposal places its idea alongside four research directions: RHO-LOSS for held-out-loss-based example selection; DoReMi for proxy-model domain weighting; Online Data Mixing for treating domains as bandit arms; and DataDecide for testing whether small-scale data choices transfer to larger runs.
The proposed distinction is deliberately narrow: decide at whole-shard granularity, roll back the model, use pre-written thresholds, and inspect repeatedly rejected shards to inform source selection. The source cites these directions as related ideas; it does not report a head-to-head comparison or establish a novelty claim against all prior work.
Related-work framing: §6 ↗. The empirical motivation below comes from GoLLeM-v5, not from a completed queue experiment.
The motivation comes from GoLLeM-v5. Late mixture changes moved efficiency by amounts comparable to seed or checkpoint variation. Even with an unchanged recipe, the final 64M run lost 0.42 efficiency points between steps 120k and 140k; its paired interval included zero.
One from-scratch corpus comparison gave a more distinct signal in WikiText byte perplexity: the two seeds of the same arm differed by about 0.011, while the two corpora differed by 0.155–0.166. That is roughly fourteen times the seed spread. It is only one comparison of substantially different corpora, not evidence that every small shard can be classified reliably.
| Signal | Observed variation | Data-change effect |
|---|---|---|
| Board efficiency | About ±0.4 at a checkpoint | At most about 0.5 for mixture tweaks |
| WikiText byte perplexity | 0.011 between two Z1 seeds | 0.155–0.166 between corpora |
The proposal therefore measures bits per byte on four to six fixed control domains, such as educational English, encyclopaedic English, question–answer text and Polish text. The control sets must be separate from training data, benchmarks and the selection split.
A useful queue must beat more than random rejection.
Baseline A consumes the eligible pool in plain seeded order. Control C rejects random shards at the same rate as the proposed quarantine arm B. Both controls are needed: rejecting data changes diversity and order even without a useful quality signal.
Use the same token budget, two seeds per arm and the same final decay. Adopt the queue only if B beats both A and C on the independent selection axis by more than max(0.6, 2 × seed spread) efficiency points, or its per-axis equivalent. The Polish axis must be defined before use.
If B fails to clear the threshold against either control, retain plain-order training and use the quarantine list only for analysis. Lower control losses alone do not meet the success criterion.
Proposed rule: §4 (continued on page 4) ↗. No measured verdict is reported.
Separate the rejection rule from the success rule.
Let Ld,t be held-out control loss in bits per byte for domain d after shard t. Let σd be the calibrated spread of shard-to-shard loss changes on a homogeneous source. The basic rule can be summarised as:
reject(t) ⇐ ∃d: ΔLd,t > 3σd
and the rise does not recover over the next shard
This notation summarises the paper’s verbal rule; it does not add a new numerical definition of recovery. The recovery criterion and σ-level must be fixed before the experiment and calibrated jointly across domains to keep measured false rejections below 5%.
adopt ⇔ effB − effA > τ
and effB − effC > τ
The second expression is the proposed efficiency-axis adoption rule after matched final decay. It uses independent selection data, not the losses that trigger rejections. For the optional curve-guided variant, the frozen reference is L̂d(D) = ad + bdD−cd; its fit must first be identifiable.
Make rollback and retry part of the protocol.
- Build and freeze a rotation. Split the eligible pool into document-aligned shards. Record their hashes, source labels and seeded order in a versioned manifest. The proposal gives 1.5× the training budget as one example of a larger candidate pool.
- Train in the stable phase. A warmup–stable–decay schedule keeps the learning rate constant while the queue operates. Each checkpoint is a rollback point; the queue position must be saved with it.
- Check for recovery. The basic rule flags a domain-loss rise above a pre-set threshold, illustrated as 3σd, that does not recover over the next shard. Calibration must keep the measured false-rejection rate below 5% across the domains.
- Roll back and retry later. Restore the checkpoint before the rejected shard. Return the shard used to test recovery to the front of the queue. Defer the rejected shard until the next rotation; a second rejection leads to quarantine.
- Inspect the cause. Examine source mix, language, register, duplicates, benchmark overlap and boilerplate. Drop or filter identified problems; the proposal permits one return from quarantine when no cause is found.
The decay at the end of training uses accepted shards only. New sources enter at the next rotation after the input checks and their own calibration; they do not silently change an active manifest.
What about a shard that is harmless but unhelpful?
The optional curve-guided rule compares observed loss improvement with a frozen reference trajectory, L̂d(D) = ad + bdD−cd. It is used only after the fitted curve passes an identifiability check. The reference must not learn only from the shards already accepted by the queue. The paper also requires a pre-written stop condition if rejections become excessive or the pool becomes too small. These remain design choices to test. §2, curve-guided variant ↗
Compute and implementation
On the v5 trainer, a 20,000-step checkpoint interval contains 20,000 × 32 × 1,024 = 655,360,000 tokens, approximately 0.66B. The proposal estimates about 55 minutes per 64M shard and 78 minutes per 128M shard using its cited v5 throughput.
A rejection with the recovery check loses two shards of work. A branch check can add another shard on a second GPU. Equal token budgets in the proposed experiment therefore do not imply equal wall time or cost.
Screen cheaply first
Training on every shard in a 100B-token pool still costs 100B tokens of training, even if only 30B are retained. The suggested first tier uses forward-only scoring with a small reference model, then trains on selected candidates. Whether this predicts the training decision must also be measured.
Make rollback reproducible
The implementation needs a shard builder, queue-aware sampler, checkpointed position and a supervisor for rejection decisions. An injected rejection must demonstrate that the trainer resumes from the earlier state with the intended replacement shard.
At two bytes per token, a 25B-token queue needs roughly 50 GB before other files. The paper proposes a volume or shard streaming instead of assuming a small container disk can hold it.
Sources: §2, two-tier screening ↗; §5, implementation and cost ↗. All timings are estimates from the source hardware, not measurements of an implemented queue.
Calibrate, then compare three arms.
The experiment starts on the 64M proxy shape. Stage A uses ten shards from a homogeneous source to estimate domain noise, calibrate the rule and, where applicable, fit a reference curve. Stage B compares three arms at equal token budget, with two seeds each.
The proposed adoption rule requires B to beat both A and C after the same final decay, on the independent selection axis, by more than max(0.6, 2 × seed spread) efficiency points or its per-axis equivalent. The Polish axis must first be defined.
The queue is designed to optimise those losses. Its benefit must appear on a separate selection measure. If it fails that test, the quarantine record can still be useful for analysis without becoming a training rule.
What must the study actually discover?
- Do domain losses distinguish individual shards reliably, rather than only very different whole corpora?
- Does a cheap forward-only screening score predict the more expensive training-based decision?
- Can a delayed check and a second chance preserve beneficial new domains while rejecting harmful shards?
- Does a gain on the independent selection axis justify the extra rejected compute?
- Do source-level observations remain useful when the model size or training stage changes?
The proposed queue also produces a dataset of decisions: shard identity, source and register mix, screen score, measured loss change or residual, and rejection history. Those labels can support a source catalogue; transfer to a universal quality classifier remains an open hypothesis.
Try to break the mechanism and its interpretation.
The proposed implementation gate injects a rejected shard and checks exact restoration of the preceding model state and the intended replacement sequence. Queue position, immutable rotation manifests and shard hashes must make resume reproducible.
For early rejections, a branch from the same checkpoint uses a different shard from the same source. If it produces the same increase, the paper treats the training phase as a competing explanation and returns the shard. Log the rejection, recovery check, retry and every return so that the verdict can be reconstructed.
Proposed checks: §5 (continued on page 5) ↗. These are requirements for a future implementation, not checks completed by this web edition.
A new language or domain may make existing control losses worse before helping later. A queue can mistake that transition for poor data. Repeated decisions can also overfit the control sets, steadily narrowing what the model sees.
- Short-term effects: source-specific calibration, delayed judgement and a later retry are proposed safeguards, not guarantees.
- Goodhart’s law: independent selection sets check whether lower control losses translate into broader quality.
- Changing noise: recalibration follows a schedule written before training, rather than being adjusted after an inconvenient result.
- Reproducibility: every queue version, rejection, measured value and return needs a durable record.
- Limited labels: a shard’s accept/reject history describes quality relative to a particular model and training stage. It is not a universal dataset ranking.
The worthwhile output may be a better understanding of sources even if the queue fails to improve a model. That is why the proposal distinguishes a useful analysis tool from an adopted training method.
Source: §6, risks ↗, with stage-dependent labels discussed in §2 ↗.
No queue result has been measured yet.
The output is a specified mechanism, a falsifiable experiment and an implementation plan. There is no demonstrated queue quality gain, rejection rate, cost saving or successful rollout in this source edition.
After the study, report B−A and B−C on the independent selection axis, seed variation, the adoption thresholds and actual compute overhead. If both comparisons clear the rule, adoption is supported within that tested setting. Otherwise the proposal remains an analysis tool and the baseline stays in place.
Return to the evidence.
This web edition is an editorial adaptation of Arkadiusz Słota’s paper, not an additional experiment or a replacement for the original. Diagrams were redrawn for the web; quantitative values are taken from the cited sections.
- Original paper · version 2
- A data queue with quarantine for GoLLeM-v6 · PDF ↗
- Repository record
- Version, author, status, license and acknowledgements ↗
Source revision:cbb0ab90cf2005dcabb75dbf2e9b412b38d6b1a8
PDF SHA-256:b4babd69da52768951cf05c7ac44a017c6151e2f1557a0b3d9ace36eae972316 - Contributions
- Paper by Arkadiusz Słota with support from the Kolektyw AI agents: Latarnik, Hart, Monter and Wartownik. The PDF describes their individual roles. Web adaptation by Fabryka AI.
- License
- The source metadata states “not decided”. Publication here does not grant a new license to the underlying paper.