ci_leaderboard recipe: judge-free, 30 rows/axis, small-n directional scores.
| # | model | Elo (BT-MLE) | MS (public) | MS (experimental ⚠) | axes |
|---|---|---|---|---|---|
| 1 | qwen2.5-0.5b-instruct | 1000.0 | 0.3667 | 0.2759 ⚠ | 3 |
| 2 | qwen2.5-1.5b-instruct | 1000.0 | 0.3667 | 0.2759 ⚠ | 3 |
⚠ Experimental = self-built axes below the measurement-grade item floor (cultural): authored by the benchmark team, not yet independently human-validated, small-n (wide CI). DIRECTIONAL ONLY — a rank-order hint, never a citable score. Kept visible (not hidden) so progress on these axes is trackable while the native-annotation program brings them to headline grade (n≥200, κ≥0.7).
41 verifiable pairwise matches fed this board. Elo ranking is
BT-MLE (order-invariant); a model is credibly above another only when their bootstrap CIs
don't overlap — see glossobench elo for full CIs. This page is generated
by CI from a small, judge-free, tiny-row-cap recipe (directional / demo only,
not a real GlossoBench-5/7 claim) — see
glossobench for the real harness.