Fabryka.Research notes
Research notes / 02 · Training dataResearch proposal · author-reviewed · no results

Can a model help choose
its next training data?

An illustrated guide to a proposed training-data queue: measure each shard, retry uncertain data, and quarantine repeated failures.

Based on the paper’s September 26 snapshot. This edited web summary preserves the reported scope and status; it is not a live experiment tracker. The original PDF contains the full methods, tables and acknowledgements.
0.66BProposed tokens per shard, aligned with a checkpoint interval
2 chancesFirst rejection: defer. Second rejection: quarantine and inspect.
3 armsPlain order, quarantine and matched random rejection

The proposal in figures

A second chance before quarantine

Schematic · no results
Data flow from eligibility checks to training, control-loss measurement, rollback, retry and quarantine.
The proposed mechanism, not a recorded execution. First rejection defers the shard; a second rejection sends it to inspection. Read the method →

A test with two ways to fail

Study design · no results
Three arms: plain queue, quarantine queue and matched random rejection.
Two seeds per arm, matched token budgets and the same final decay. B must beat both A and C on an independent selection axis. See the adoption threshold →
01 / PROBLEM

Can online data selection improve training?

Training data are usually chosen before a run. This proposal asks whether a model can help decide what it should train on next. Start with more candidate data than the budget can consume, arrange it into reproducible shards, and measure the model after each shard.

If a shard produces an unusually large, unrecovered increase in a held-out domain loss, roll back the model. A first rejection defers that shard to a later rotation of the queue. A second rejection sends it to quarantine for inspection. Data quality becomes a traceable decision with a reason and a chance to be reconsidered.

Proposal, not a demonstrated improvement.

Version 2 reports no results for this mechanism. Its “reviewed” status is the repository’s author-review status, not journal peer review. The figures below describe the proposed design.

Source: Abstract and §1 ↗.

02 / PRIOR ART

From example selection to whole-shard decisions.

The proposal places its idea alongside four research directions: RHO-LOSS for held-out-loss-based example selection; DoReMi for proxy-model domain weighting; Online Data Mixing for treating domains as bandit arms; and DataDecide for testing whether small-scale data choices transfer to larger runs.

The proposed distinction is deliberately narrow: decide at whole-shard granularity, roll back the model, use pre-written thresholds, and inspect repeatedly rejected shards to inform source selection. The source cites these directions as related ideas; it does not report a head-to-head comparison or establish a novelty claim against all prior work.

Related-work framing: §6 ↗. The empirical motivation below comes from GoLLeM-v5, not from a completed queue experiment.

The motivation comes from GoLLeM-v5. Late mixture changes moved efficiency by amounts comparable to seed or checkpoint variation. Even with an unchanged recipe, the final 64M run lost 0.42 efficiency points between steps 120k and 140k; its paired interval included zero.

One from-scratch corpus comparison gave a more distinct signal in WikiText byte perplexity: the two seeds of the same arm differed by about 0.011, while the two corpora differed by 0.155–0.166. That is roughly fourteen times the seed spread. It is only one comparison of substantially different corpora, not evidence that every small shard can be classified reliably.

Motivating observations, not queue outcomes. §3, Table 1 ↗.
SignalObserved variationData-change effect
Board efficiencyAbout ±0.4 at a checkpointAt most about 0.5 for mixture tweaks
WikiText byte perplexity0.011 between two Z1 seeds0.155–0.166 between corpora

The proposal therefore measures bits per byte on four to six fixed control domains, such as educational English, encyclopaedic English, question–answer text and Polish text. The control sets must be separate from training data, benchmarks and the selection split.

03 / BASELINE & FALSIFICATION THRESHOLD

A useful queue must beat more than random rejection.

Baseline A consumes the eligible pool in plain seeded order. Control C rejects random shards at the same rate as the proposed quarantine arm B. Both controls are needed: rejecting data changes diversity and order even without a useful quality signal.

Use the same token budget, two seeds per arm and the same final decay. Adopt the queue only if B beats both A and C on the independent selection axis by more than max(0.6, 2 × seed spread) efficiency points, or its per-axis equivalent. The Polish axis must be defined before use.

What would reject adoption?

If B fails to clear the threshold against either control, retain plain-order training and use the quarantine list only for analysis. Lower control losses alone do not meet the success criterion.

Proposed rule: §4 (continued on page 4) ↗. No measured verdict is reported.

04 / MATHEMATICAL FORMULATION

Separate the rejection rule from the success rule.

Let Ld,t be held-out control loss in bits per byte for domain d after shard t. Let σd be the calibrated spread of shard-to-shard loss changes on a homogeneous source. The basic rule can be summarised as:

ΔLd,t = Ld,t − Ld,t−1
reject(t) ⇐ ∃d: ΔLd,t > 3σd
and the rise does not recover over the next shard

This notation summarises the paper’s verbal rule; it does not add a new numerical definition of recovery. The recovery criterion and σ-level must be fixed before the experiment and calibrated jointly across domains to keep measured false rejections below 5%.

τ = max(0.6, 2 × seed spread)
adopt ⇔ effB − effA > τ
and effB − effC > τ

The second expression is the proposed efficiency-axis adoption rule after matched final decay. It uses independent selection data, not the losses that trigger rejections. For the optional curve-guided variant, the frozen reference is L̂d(D) = ad + bdD−cd; its fit must first be identifiable.

§2, rejection and curve rules ↗; §4, adoption rule ↗.

05 / METHOD

Make rollback and retry part of the protocol.

FIG. 01 / PROPOSED DATA QUEUE
Screen, train, measure, roll back; first rejection defers a shard, second rejection quarantines it.
Conceptual diagram of §2 ↗. It is not an execution trace. The paper’s detailed second-chance policy is shown explicitly.
  1. Build and freeze a rotation. Split the eligible pool into document-aligned shards. Record their hashes, source labels and seeded order in a versioned manifest. The proposal gives 1.5× the training budget as one example of a larger candidate pool.
  2. Train in the stable phase. A warmup–stable–decay schedule keeps the learning rate constant while the queue operates. Each checkpoint is a rollback point; the queue position must be saved with it.
  3. Check for recovery. The basic rule flags a domain-loss rise above a pre-set threshold, illustrated as 3σd, that does not recover over the next shard. Calibration must keep the measured false-rejection rate below 5% across the domains.
  4. Roll back and retry later. Restore the checkpoint before the rejected shard. Return the shard used to test recovery to the front of the queue. Defer the rejected shard until the next rotation; a second rejection leads to quarantine.
  5. Inspect the cause. Examine source mix, language, register, duplicates, benchmark overlap and boilerplate. Drop or filter identified problems; the proposal permits one return from quarantine when no cause is found.

The decay at the end of training uses accepted shards only. New sources enter at the next rotation after the input checks and their own calibration; they do not silently change an active manifest.

What about a shard that is harmless but unhelpful?

The optional curve-guided rule compares observed loss improvement with a frozen reference trajectory, L̂d(D) = ad + bdD−cd. It is used only after the fitted curve passes an identifiability check. The reference must not learn only from the shards already accepted by the queue. The paper also requires a pre-written stop condition if rejections become excessive or the pool becomes too small. These remain design choices to test. §2, curve-guided variant ↗

Compute and implementation

On the v5 trainer, a 20,000-step checkpoint interval contains 20,000 × 32 × 1,024 = 655,360,000 tokens, approximately 0.66B. The proposal estimates about 55 minutes per 64M shard and 78 minutes per 128M shard using its cited v5 throughput.

A rejection with the recovery check loses two shards of work. A branch check can add another shard on a second GPU. Equal token budgets in the proposed experiment therefore do not imply equal wall time or cost.

Screen cheaply first

Training on every shard in a 100B-token pool still costs 100B tokens of training, even if only 30B are retained. The suggested first tier uses forward-only scoring with a small reference model, then trains on selected candidates. Whether this predicts the training decision must also be measured.

Make rollback reproducible

The implementation needs a shard builder, queue-aware sampler, checkpointed position and a supervisor for rejection decisions. An injected rejection must demonstrate that the trainer resumes from the earlier state with the intended replacement shard.

At two bytes per token, a 25B-token queue needs roughly 50 GB before other files. The paper proposes a volume or shard streaming instead of assuming a small container disk can hold it.

Sources: §2, two-tier screening ↗; §5, implementation and cost ↗. All timings are estimates from the source hardware, not measurements of an implemented queue.

06 / EXPERIMENT

Calibrate, then compare three arms.

The experiment starts on the 64M proxy shape. Stage A uses ten shards from a homogeneous source to estimate domain noise, calibrate the rule and, where applicable, fit a reference curve. Stage B compares three arms at equal token budget, with two seeds each.

FIG. 02 / PROPOSED EXPERIMENT
A: plain queue. B: quarantine queue. C: random rejection at B’s rate. Two seeds each, matched token budget.
Study design from §4 ↗. Random rejection controls for changing the order and diversity simply by rejecting data.

The proposed adoption rule requires B to beat both A and C after the same final decay, on the independent selection axis, by more than max(0.6, 2 × seed spread) efficiency points or its per-axis equivalent. The Polish axis must first be defined.

Better control losses are not the success criterion.

The queue is designed to optimise those losses. Its benefit must appear on a separate selection measure. If it fails that test, the quarantine record can still be useful for analysis without becoming a training rule.

07 / RESEARCH

What must the study actually discover?

  • Do domain losses distinguish individual shards reliably, rather than only very different whole corpora?
  • Does a cheap forward-only screening score predict the more expensive training-based decision?
  • Can a delayed check and a second chance preserve beneficial new domains while rejecting harmful shards?
  • Does a gain on the independent selection axis justify the extra rejected compute?
  • Do source-level observations remain useful when the model size or training stage changes?

The proposed queue also produces a dataset of decisions: shard identity, source and register mix, screen score, measured loss change or residual, and rejection history. Those labels can support a source catalogue; transfer to a universal quality classifier remains an open hypothesis.

Questions and possible by-products from §2 ↗ and §2–4 ↗.

08 / VERIFICATION

Try to break the mechanism and its interpretation.

The proposed implementation gate injects a rejected shard and checks exact restoration of the preceding model state and the intended replacement sequence. Queue position, immutable rotation manifests and shard hashes must make resume reproducible.

For early rejections, a branch from the same checkpoint uses a different shard from the same source. If it produces the same increase, the paper treats the training phase as a competing explanation and returns the shard. Log the rejection, recovery check, retry and every return so that the verdict can be reconstructed.

Proposed checks: §5 (continued on page 5) ↗. These are requirements for a future implementation, not checks completed by this web edition.

A new language or domain may make existing control losses worse before helping later. A queue can mistake that transition for poor data. Repeated decisions can also overfit the control sets, steadily narrowing what the model sees.

  • Short-term effects: source-specific calibration, delayed judgement and a later retry are proposed safeguards, not guarantees.
  • Goodhart’s law: independent selection sets check whether lower control losses translate into broader quality.
  • Changing noise: recalibration follows a schedule written before training, rather than being adjusted after an inconvenient result.
  • Reproducibility: every queue version, rejection, measured value and return needs a durable record.
  • Limited labels: a shard’s accept/reject history describes quality relative to a particular model and training stage. It is not a universal dataset ranking.

The worthwhile output may be a better understanding of sources even if the queue fails to improve a model. That is why the proposal distinguishes a useful analysis tool from an adopted training method.

Source: §6, risks ↗, with stage-dependent labels discussed in §2 ↗.

09 / RESULT

No queue result has been measured yet.

Version 2: proposal; results = none.

The output is a specified mechanism, a falsifiable experiment and an implementation plan. There is no demonstrated queue quality gain, rejection rate, cost saving or successful rollout in this source edition.

After the study, report B−A and B−C on the independent selection axis, seed variation, the adoption thresholds and actual compute overhead. If both comparisons clear the rule, adoption is supported within that tested setting. Otherwise the proposal remains an analysis tool and the baseline stays in place.

SOURCE & ATTRIBUTION

Return to the evidence.

This web edition is an editorial adaptation of Arkadiusz Słota’s paper, not an additional experiment or a replacement for the original. Diagrams were redrawn for the web; quantitative values are taken from the cited sections.

Original paper · version 2
A data queue with quarantine for GoLLeM-v6 · PDF ↗
Repository record
Version, author, status, license and acknowledgements ↗
Source revision: cbb0ab90cf2005dcabb75dbf2e9b412b38d6b1a8
PDF SHA-256: b4babd69da52768951cf05c7ac44a017c6151e2f1557a0b3d9ace36eae972316
Contributions
Paper by Arkadiusz Słota with support from the Kolektyw AI agents: Latarnik, Hart, Monter and Wartownik. The PDF describes their individual roles. Web adaptation by Fabryka AI.
License
The source metadata states “not decided”. Publication here does not grant a new license to the underlying paper.