Lëtzebuergesch · LLM Evaluation

LuxBench

Which language model actually speaks Luxembourgish — not German in disguise?

A native-language benchmark of 21 frontier models across 12 tasks: free writing, translation from German, French & English, grammar, idiom, native vocabulary and reading comprehension. Scored by a deterministic linguistic anchor plus a blind three-model judge panel.

21
Models tested
12
Tasks each
GPT-5
Best overall
Qwen3.7 Max
Best Chinese model

The leaderboard

Final = 0.55 · judge + 0.45 · objective
# Model Final score Judge Objective LB-rate Drift Lexical
0 German-drift markers 1–3 markers 4+ markers CN Chinese lab US US lab EU European lab † judged single-model (indicative)

Tiebreaker — the hard round

Top 8 · drift-exposing grammar

The leaders clustered at 85–95, so a second round pushed harder: the Eifeler-Regel (Luxembourgish's n-deletion rule — the classic native-vs-German tell), preposition contractions, German-drift correction, and longer free composition. This is where a Chinese model came out on top.

#ModelHard scoren-ruleDrift-fixContractions

The €0 correctness layer

74%

A deterministic normalizer — no model, running on a 4-core CPU — removed 74% of German-drift markers (19 → 5) across all outputs and applied 46 Eifeler-Regel corrections, catching errors even the top models made.

den Mann → de Mann  ·  Ech ginn muer → Ech gi muer  ·  Ich → Ech

Because Luxembourgish orthography is codified (the 2019 reform, the Eifeler-Regel), correctness can be enforced, not merely learned — a moat that holds no matter which base model runs underneath.

How it's scored

Objective anchor

Deterministic, no LLM: German-drift ratio (Luxembourgish-only vs German-only function words), orthography density (ë / é / ä), and exact-match native vocabulary. Immune to judge unreliability.

Blind judge panel

Three frontier models — Claude Opus 4.8, GPT-5, Gemini 3.1 Pro — grade anonymized responses for authenticity, correctness and task success, so no model can favour its own output.

Two rounds

A 12-task standard battery ranks all 21; a harder tiebreaker round separates the top 8 on the grammar that actually distinguishes native Luxembourgish from German.