Back to blogEngineering

EvaluatingLLMsforProduction:BuildinganEvalPipelineYourTeamTrusts

Renata Osei· Head of AI Engineering· November 18, 2025· 8 min read

Start with a golden dataset, not a vibe check

The most common failure in LLM development is deciding a change is good because a handful of manual test prompts look better. That approach does not scale past the first few iterations, and it hides regressions in cases nobody happened to try. We start every serious LLM feature with a golden dataset: 50 to 300 labeled examples pulled from real usage where possible, each with an expected output or a rubric for what a correct answer looks like.

The dataset needs to include the ugly cases, not just the clean ones: ambiguous questions, edge-case inputs, and known failure modes from earlier iterations. A golden dataset that only contains easy examples will tell you your system is better than it is, right up until it meets a real user.

Human review loops that actually scale

Automated scoring (exact match, semantic similarity, or a second LLM acting as judge) is useful for catching regressions quickly, but it cannot fully replace human judgment, especially for subjective quality like tone or reasoning soundness. We run a lightweight weekly review where a domain expert scores a random sample of 20 to 40 production outputs on a simple rubric, and we track that score over time alongside the automated metrics.

The key to making this sustainable is keeping the review interface fast: one screen, one output, three to five rating buttons, under 15 seconds per item. Reviews that take longer than that get skipped under deadline pressure, and a review process nobody actually does is worse than not having one, because it creates false confidence.

Regression testing across prompts and models

Every prompt change, every model version bump, and every retrieval tweak gets run against the full golden dataset before it ships, the same way a code change runs against a test suite. We track the score over time in a simple dashboard, and any change that drops the aggregate score by more than a small threshold (we use 3 percentage points as a default) requires an explicit sign-off rather than an automatic merge.

This matters most when swapping underlying models. A newer, cheaper, or faster model can look like a clear upgrade in general benchmarks while quietly regressing on your specific use case, particularly on domain-specific formatting or tone requirements. We have seen model upgrades that improved general reasoning benchmarks while dropping our client-specific golden set score by double digits.

Choosing metrics that map to business outcomes

Accuracy on a golden set is a proxy, not the goal. We push every eval pipeline to also track at least one metric tied to the actual business outcome: ticket deflection rate, time to first response, or downstream conversion. When the proxy metric and the business metric disagree, the business metric wins, and that disagreement is usually the signal that the golden dataset needs to be updated.

Wiring evals into the deployment pipeline

The eval suite only earns its keep if it runs automatically, not manually before a big demo. We wire it into the same CI pipeline that runs unit tests, gate deployments on the aggregate score, and store historical results so a regression introduced three weeks ago is traceable to the exact commit that caused it, rather than discovered by a confused user.

LLM EvaluationTestingAI EngineeringQuality

Wanthelpshippingsomethinglikethis?