Writing rubrics a reviewer model can actually apply
Most quality rubrics are written for people who already know what good looks like. Here is how to write one for a model.

Ines Moreau
Head of Research
Share

Hand a typical QA scorecard to a reviewer model and it will return confident, consistent and mostly meaningless scores. The problem is rarely the model. It is that the scorecard was written for people who share years of unstated context about what “empathetic” means at this company.
A rubric for a model has to carry that context itself.
Replace adjectives with observations
“Was the agent empathetic?” is not a criterion; it is an invitation to guess. Rewrite it as things a reader could point to in the transcript:
The agent acknowledged the customer’s problem before asking for information.
The agent did not repeat an apology more than once.
The agent did not use the phrases on the brand guide’s “avoid” list.
Each of those can be checked, and each can be wrong in a way you can see.
Anchor every score
A five-point scale with no anchors produces a lot of fours. Give the reviewer an example of each level, taken from real conversations, and tell it which observation separates one level from the next.
The evidence: required line matters more than it looks. It makes the reviewer quote the part of the transcript that justifies the score, which makes disagreements quick to resolve.
One question per criterion
If a criterion contains the word “and”, split it. “Accurate and complete” is two questions, and a conversation can easily be one without the other. When they share a score, you can no longer tell which one moved.
Calibrate with people, not against them
Before a rubric goes live, three people score the same forty conversations independently. Where they disagree with each other, the rubric is ambiguous, and no model will fix that. Where they agree with each other and disagree with the reviewer, look at the reviewer’s quoted evidence. Usually the rubric has a gap the humans filled from experience.
Fix the rubric, rerun, and repeat until the reviewer lands within a point of the human consensus on at least nine in ten conversations. Then, and only then, let it score production traffic.
Share this post
evaluation
rubrics

Written by
Ines Moreau
Head of Research




