AI Evaluation & Quality
Decide whether a model's output is good enough to ship: write the rubric, sample the failures nobody reported, catch the average that hides a group being failed, and say no to a launch.
The role in 2026
The person who decides whether a model's output is good enough to put in front of a customer.
Gone to the AI: Reading every output by hand. Sampling is cheap now; deciding what counts as good enough is not.
- Write down what good means, so two reviewers score the same
- Sample the failures nobody reported, not the traffic
- Say no to a launch that is better on average and worse for a group
10 briefs · about 4 weeks at your own pace
See the briefs
- 1Six answers from the support bot. Which ship? · memo · ~20 min
- 2Turn “the answers should be good” into something two people would score the same · memo · ~20 min
- 3Accuracy went up. Complaints went up too. · memo · ~20 min
- 4Better on average, worse for one group. Ship or not? · memo · ~20 min
- 5Where the errors actually cost us · worksheet · ~25 min
- 6Claude drafted the evaluation report. You sign it. · an AI draft you edit and sign · ~25 min
- 7Build the test set for the failures nobody reports · memo · ~20 min
- 8Five slides: is it good enough to ship? · deck · ~30 min
- 9Build an eval set, run it, and tell me what broke · built with AI tools, handed in as a link or PDF · ~30 min
- 10Capstone: the evaluation report someone could act on · built with AI tools, handed in as a link or PDF · ~30 min