Problem
Can small models with different mechanisms complement one another on symbolic music? The pilot studies next-character prediction in ABC-encoded Bach melodies, and separately checks whether generated sequences complete with valid internal bar lengths.
Prior art
The repository’s earlier music-expert work treats an ensemble as the baseline for representation stitching. This later Bach pilot compares a count-based n-gram, an NPLM with a fixed context window and attention-based GPT models. Historical Irish-music results use another corpus and are not pooled with these scores.
Baseline & falsification threshold
N-gram order 6 is the count-based reference. For the hybrid, the stronger single predictor—NPLM—is the baseline. The report evaluates paired family-level differences, but does not state a pre-registered minimum practical improvement for every comparison. We therefore report the measured effect and interval without inventing an adoption threshold.
Mathematical formulation
phybrid(x) ∝ √(pNPLM(x) · pGPT(x))
The second line describes the equal-weight logit ensemble: a normalised geometric mean. Averaging probabilities instead produces an arithmetic mixture and a different result.
Method
All listed predictors use the same ABC encoding and are scored on its musical part; EOS is excluded from the denominator. Bootstrap samples whole tune families. The hybrid variants and 50/50 weights were recorded before the first test evaluation.
Experiment
The test covers 35 melodies in 17 families. NPLM has about 79k parameters; each GPT about 820k. They differ in context, optimisation and regularisation, so this is a pilot comparison of configured systems, not an isolated architecture or equal-compute study.
Research
NPLM and GPT B disagree on their most likely character at 18.09% of positions, versus 8.28% for the two GPTs. Their character-loss correlation is lower too. These observations motivate complementary-error research, but useful ensemble improvement must still be measured directly.
Verification
NPLM beats GPT B in 13 of 17 families. The paired PPL-ratio interval is 0.929–0.988. For the hybrid versus NPLM, the exploratory interval is 0.939–0.968, without multiple-comparison correction. The test has now been used; further tuning requires a new evaluation pool or a pre-planned grouped protocol.
Result
NPLM test PPL is 1.9501; GPT B is 2.0375; their logit hybrid is 1.8595, 4.65% below NPLM. That result is limited to this pilot. An ensemble executes both models; it is not a merged single model.
Inspect the original work.
Credit: Slayer Micro-Models contributors. Editorial web adaptation: Fabryka AI.
Original report and surrounding artifacts ↗
Repository: slayerlabs/micro-models
Revision: df1c984a9d3be6e5bb2092dbc89d2a220b2825f7. Repository and source licenses apply; this summary grants no additional rights to the source material.