Fabryka.Research notes
Research / ladderProvisional diagnostic · incomplete tier

A Polish probe, with its limits visible

Koliber and Pollock score the same 2,000 Polish agreement pairs. A useful diagnostic does not automatically become a complete benchmark.

Repository-reported evidence, presented with its original scope. This web edition does not rerun the experiments or change the source’s result status.

Agreement by number, gender and person

Provisional measurement
Agreement by number, gender and person
Shared pair IDs and protocol. Number: 667 pairs; gender: 667; person: 666. Native review and training-overlap checks remain pending. Source ↗
01 / PROBLEM

Problem

Can a cheap continuous probe reveal differences between nearby small models before a more expensive evaluation? This run evaluates Polish subject–verb agreement while preserving missing coverage and the provisional status of the input pairs.

Source note at the inspected revision ↗

02 / PRIOR ART

Prior art

The pairs come from a pinned Polish MultiBLiMP source. Koliber is primarily Polish, while Pollock’s model card lists English training. This language-domain mismatch prevents an overall-capability ranking, even when the evaluation items and numerical settings match.

Source note at the inspected revision ↗

03 / BASELINE & FALSIFICATION THRESHOLD

Baseline & falsification threshold

The paired models form a descriptive comparison, not a trained multi-seed baseline. The repository proposes a seed-calibrated discrimination factor and thresholds for later ladder admission. Training-seed variance was not measured here, so this run cannot clear those decision gates.

Source note at the inspected revision ↗

04 / MATHEMATICAL FORMULATION

Mathematical formulation

pair probability = σ(log p(good) − log p(bad))
DF = |mean paired difference| / σseed

The continuous pair score is averaged over items. It is not accuracy. DF remains unavailable without the seed standard deviation; dividing by evaluation-item noise would answer a different question.

Source note at the inspected revision ↗

05 / METHOD

Method

Both models retain their own tokenizers and document-start tokens. They run in fp32 with math SDPA, batch size 8, context 512, on an M4 Max GPU through MPS. The source checks suite hashes, ordered pair IDs and critical-region coverage before publishing.

Source note at the inspected revision ↗

06 / EXPERIMENT

Experiment

The models score the same 2,000 full-sentence pairs. A common subset of 1,171 pairs supports critical-region scoring. The three paradigms are subject–verb number, gender and person. No corpus BPB was measured because the required frozen corpus was unavailable.

Source note at the inspected revision ↗

07 / RESEARCH

Research

The next research step is a reviewed complete tier with held-out corpus loss, a matched English probe and training-seed calibration. Cheap runtime is useful only if the resulting signal separates the intended models reliably under a frozen protocol.

Source note at the inspected revision ↗

08 / VERIFICATION

Verification

The repository stores per-item scores, model revisions, protocol metadata, paired aggregates and hashes. Native linguistic review and training-overlap assessment are still missing. The article reproduces recorded results; it does not complete those outstanding validations.

Source note at the inspected revision ↗

09 / RESULT

Result

Mean sentence-pair probability is 0.9083 for Koliber and 0.5914 for Pollock. Critical-region values are 0.8913 and 0.5718. Measured evaluation times were 111.3 s and 100.4 s, excluding loading and output serialization. These are provisional Polish-probe results, not a full-tier or matched-language ranking.

Source note at the inspected revision ↗

PROVENANCE

Inspect the original work.

Credit: SlayerLab ladder contributors. Editorial web adaptation: Fabryka AI.

Original report and surrounding artifacts ↗

Repository: slayerlabs/ladder
Revision: 7b7c7a6d8d6a542ef74490a1ae841d012042188b. Repository and source licenses apply; this summary grants no additional rights to the source material.