A golden set is a collection of 20β50 real inputβideal-output pairs, harvested from the manual era of the client's process β before any AI was involved.
Real inputs beat synthetic ones. They surface the edge cases that matter: the weird formatting, the off-topic question, the ambiguous brief that a real user actually sent.
Don't treat the golden set as permanent. Update it quarterly with new failure modes discovered through production sampling. Your eval should get harder over time, not easier.