Problem
Can a cheap continuous probe reveal differences between nearby small models before a more expensive evaluation? This run evaluates Polish subject–verb agreement while preserving missing coverage and the provisional status of the input pairs.
Prior art
The pairs come from a pinned Polish MultiBLiMP source. Koliber is primarily Polish, while Pollock’s model card lists English training. This language-domain mismatch prevents an overall-capability ranking, even when the evaluation items and numerical settings match.
Baseline & falsification threshold
The paired models form a descriptive comparison, not a trained multi-seed baseline. The repository proposes a seed-calibrated discrimination factor and thresholds for later ladder admission. Training-seed variance was not measured here, so this run cannot clear those decision gates.
Mathematical formulation
DF = |mean paired difference| / σseed
The continuous pair score is averaged over items. It is not accuracy. DF remains unavailable without the seed standard deviation; dividing by evaluation-item noise would answer a different question.
Method
Both models retain their own tokenizers and document-start tokens. They run in fp32 with math SDPA, batch size 8, context 512, on an M4 Max GPU through MPS. The source checks suite hashes, ordered pair IDs and critical-region coverage before publishing.
Experiment
The models score the same 2,000 full-sentence pairs. A common subset of 1,171 pairs supports critical-region scoring. The three paradigms are subject–verb number, gender and person. No corpus BPB was measured because the required frozen corpus was unavailable.
Research
The next research step is a reviewed complete tier with held-out corpus loss, a matched English probe and training-seed calibration. Cheap runtime is useful only if the resulting signal separates the intended models reliably under a frozen protocol.
Verification
The repository stores per-item scores, model revisions, protocol metadata, paired aggregates and hashes. Native linguistic review and training-overlap assessment are still missing. The article reproduces recorded results; it does not complete those outstanding validations.
Result
Mean sentence-pair probability is 0.9083 for Koliber and 0.5914 for Pollock. Critical-region values are 0.8913 and 0.5718. Measured evaluation times were 111.3 s and 100.4 s, excluding loading and output serialization. These are provisional Polish-probe results, not a full-tier or matched-language ranking.
Inspect the original work.
Credit: SlayerLab ladder contributors. Editorial web adaptation: Fabryka AI.
Original report and surrounding artifacts ↗
Repository: slayerlabs/ladder
Revision: 7b7c7a6d8d6a542ef74490a1ae841d012042188b. Repository and source licenses apply; this summary grants no additional rights to the source material.