
Benchmarks & Comparisons
8 min
OpenJev vs an LLM Judge on 990 Planted Errors: Recall, Calibration, Cost
The same 990 support tickets with 495 deliberately wrong labels, judged two ways: OpenJev over /v1/classify and Qwen 3.8 27B as an LLM judge. Asked one pair at a time OpenJev catches 82.6% of the planted errors; asked to pick among all twelve intents it catches 98.2%, rejects the same 33 mislabelled 'complaint' rows the 27B judge rejected, and its probabilities are calibrated to within 3 points. Every number is exact because we planted the errors.
benchmarksOpenJevLLM-as-a-judge