
Guides & Tutorials
12 min
LLM-as-a-Judge with an Open Model: 990 Planted Errors, 99.6% Caught, and What the Rubric Changes
We used Qwen 3.8 27B over an OpenAI-compatible API as a judge of dataset labels, and scored it against errors we planted ourselves so every number is exact. With a written rubric the judge caught 491 of 493 wrong labels and also rejected 58 labels the dataset called correct, 43 of them the same systematic mistake in the generator. Without the rubric it agreed with the dataset more and missed 20 planted errors. The rubric is not a tuning knob; it decides what the judge measures.
LLM-as-a-judgeevaluationdata quality