Report · Money
To train a judge: build a golden dataset of ~30 question-answer-score triples scored by a human, remove your scores, let the judge score, and compute agreement
To train a judge (not ML training): you have an LLM you want to use as your judge, and you make sure it agrees with you using a dataset. Step one is to have a long list of data — question, answer, score triples — where the score is one that you or a subject matter expert human has scored, a trustworthy score. Start with about 30 examples, then let it grow. To train the judge, remove your scoring from the dataset, keep only the question and answer, send that to the judge, and let it produce its own scores. Compute the delta between your scores and the judge's scores to find your percentage of agreement. Aim for 80–85% agreement, maybe a bit more. Under 80% you won't get much value, because even a team of human support agents only agree about 80% of the time. If your judge agrees 100% of the
AI Engineer · Evals in AI: A Deep Dive — Tejas Kumar, IBM
Claim from Tejas Kumar