GlossoBench — ms leaderboard v1.0

ci_leaderboard recipe: judge-free, 30 rows/axis, small-n directional scores.

#modelElo (BT-MLE)MS (public)MS (experimental ⚠)axes
1qwen2.5-0.5b-instruct1000.00.36670.2759 ⚠3
2qwen2.5-1.5b-instruct1000.00.36670.2759 ⚠3

Experimental = self-built axes below the measurement-grade item floor (cultural): authored by the benchmark team, not yet independently human-validated, small-n (wide CI). DIRECTIONAL ONLY — a rank-order hint, never a citable score. Kept visible (not hidden) so progress on these axes is trackable while the native-annotation program brings them to headline grade (n≥200, κ≥0.7).

41 verifiable pairwise matches fed this board. Elo ranking is BT-MLE (order-invariant); a model is credibly above another only when their bootstrap CIs don't overlap — see glossobench elo for full CIs. This page is generated by CI from a small, judge-free, tiny-row-cap recipe (directional / demo only, not a real GlossoBench-5/7 claim) — see glossobench for the real harness.