WHY IT WORKS
A golden dataset turns "does this look better" (a vibe, subject to whoever is looking) into "does this score better against 50 fixed, agreed-on examples" (a number, comparable run over run). This is the entire reason software testing exists, applied to a system whose outputs are less deterministic than a function's return value.
LLM-as-judge works because grading against a specific written rubric ("is the category correct: yes/no. Is the drafted response polite and on-topic: 1-5") is a much narrower, more reliable task for a model than the original open-ended generation task - judging is closer to classification than to creative generation, and narrower tasks are where LLMs are most consistent.
The regression gate is what actually prevents the failure in the problem statement: a change that improves the aggregate score on the golden set can ship; a change that drops it cannot, automatically, before a human ever has to notice in production.