AI Automation

AI Agents in Production: Evals, Golden Sets, SLAs

AI Agents in Production: Evals, Golden Sets, SLAs

Opening answer (BLUF)

Once AI agents are in production, quality is a measurement problem, not a shipping problem. Controls that stop a bad action do not tell you whether the agent is getting better or quietly slipping. NIST’s AI Risk Management Framework states that AI systems should be tested before deployment and regularly while in operation, with performance checked under conditions similar to the live setting and behavior monitored in production [1]. The loop that works is a golden set from real cases, offline evals paired with online sampling, and a quality SLA operations can actually staff.

Measurement is not the same as control

Logging, approvals, and kill switches are controls. They bound what the agent is allowed to do. They do not produce a trend line for task success, helpfulness, or policy violations. If the only signal you have is “nothing catastrophic happened this week,” you cannot tell a prompt tweak from a model swap from luck. The OECD’s robustness principle says AI systems should function appropriately throughout their lifecycle, and that risks should be continually assessed and managed, not checked once at launch [3]. Accountability is the companion: apply systematic risk management at each lifecycle phase on an ongoing basis, with traceability of datasets, processes, and decisions so outputs can be analyzed later [4].

Build a golden set from real cases

A golden set is a versioned collection of tasks with defined inputs, expected outcomes, and pass or fail criteria. It is not a brainstorm of hypothetical prompts. Anthropic’s applied guidance is to start with 20 to 50 simple tasks drawn from real failures, converting what you already test by hand, plus items from the bug tracker and support queue once the agent is live [6]. OpenAI’s eval design guidance says the same: mix production data, expert-curated cases, and historical logs, and mine those logs for new eval cases as you go [7].

Write each task so two domain experts would independently reach the same verdict. Anthropic’s test is practical: could a competent person pass the task themselves, and is every grader check visible in the task description [6]. Attach a reference solution, a known-good outcome that already passes the graders, so you know the task is solvable and the grader is not broken.

Label more than the prompt: the business outcome that counts as done, policy constraints, difficulty, and source (ticket ID or sampled trace). Balance the set. If you only test cases where a tool should be called, you will over-call it. Anthropic’s web-search evals had to cover both “search when the fact is live” and “answer from knowledge when the fact is stable,” because one-sided suites produce one-sided agents [6]. The same pattern applies to refunds, escalations, and “I don’t know.”

Keep the golden set under version control next to the prompt and tool schema. When a live failure is confirmed, add it. When a case saturates, graduate it into the regression suite.

Offline evals versus online sampling

Offline evals run the current agent (model, prompt, tools, and harness together) against the frozen golden set. They are repeatable, they do not touch users, and they can run on every prompt or model change. Anthropic treats automated evals as the pre-launch and CI layer [6]. OpenAI’s continuous-evaluation pattern is to run evals on every change and grow the set as new forms of nondeterminism appear [7].

Online sampling scores a slice of live traffic. NIST’s MEASURE 2.4 subcategory is the standards language for this: the functionality and behavior of the AI system and its components should be monitored when in production [1]. MEASURE 2.3 is the offline counterpart: performance or assurance criteria demonstrated for conditions similar to the deployment setting [1]. Google’s agent-evaluation metrics are designed to be registered once and applied to both offline assessments and continuous online monitors [8].

Use offline runs to answer, “Did this change regress known work?” Use online sampling to answer, “Has the live mix drifted away from the golden set?” Anthropic is blunt about the tradeoff: production monitoring reveals real user behavior and catches issues synthetic evals miss, but problems can reach users first, and live traces rarely come with ground truth [6]. Sampling therefore needs a grader (code check, LLM-as-judge, or human) and a sampling plan, not a dashboard of token counts. An agent that is perfect on last quarter’s golden set can still fail this week’s new phrasing. An agent that “feels fine” in production can still have dropped on a policy it used to handle. You need both clocks.

Score task success, helpfulness, and policy violations separately

A single “quality score” hides the failure mode. Google’s agent metrics make the first cut cleanly: multi-turn task success asks whether the goal was achieved, not how, while trajectory quality asks whether the path was logical and efficient [8]. Anthropic’s advice is to grade the outcome first (did the reservation exist in the database, not merely did the agent say “booked”), because path-matching graders punish valid solutions the eval designer did not anticipate [6]. Keep path and tool-use scores as diagnostics unless the path itself is the policy (identity must be verified before a refund tool may fire).

We recommend three rates that operations can staff:

  1. Task-success rate. Did the work get done against the defined outcome? For agents that call tools, this is state in the system of record, not fluency of the final message. Google scores hallucination as a separate claim-level check against tool results, which is the right place to put “confidently wrong” rather than burying it inside success [8]. NIST’s generative AI profile names that failure confabulation: confidently stated but erroneous content that can mislead the user [2].
  1. Helpfulness (or interaction quality). Was the conversation usable? An agent can close a ticket and still fail on tone, turn count, or unexplained decisions. Anthropic’s support-agent pattern scores state, transcript constraints, and a rubric for grounded explanations as separate graders [6]. OpenAI similarly splits instruction following and functional correctness from tool selection and argument precision [7].
  1. Policy-violation rate. Did the agent break a rule you care about, even if the user was happy? Safety, PII, spend limits, and “never invent a policy” belong here. Google’s safety metric is a binary pass or fail against stated harm categories [8]. A drop in helpfulness is a product issue. A rise in policy violations is an incident.

Do not stop at pass@1. Agent runs vary. Anthropic distinguishes pass@k (at least one success in k trials) from pass^k (every one of k trials succeeds) [6]. Their worked example: a 75% per-trial success rate yields about 42% probability of succeeding on all of three independent trials [6]. Customer-facing AI agents that users expect to behave the same way every time should report pass^k, or at least a reliability band, not only a single lucky pass.

Stanford’s 2022 HELM project made the same point at research scale. For each of 16 core scenarios, HELM measured seven metrics (accuracy, calibration, robustness, fairness, bias, toxicity, and efficiency) rather than treating accuracy as the only number that counted [5]. At the time, models had been evaluated on only 17.9% of those core scenarios on average, which HELM raised to 96.0% under standardized conditions [5]. The lesson that transfers is the refusal to let one metric stand in for the rest.

Regression tests when you change a prompt or model

Capability evals ask what the agent can do that it currently cannot. They should start with a low pass rate. Regression evals ask whether the agent still handles the work it already owned. Anthropic says those should sit near a 100% pass rate, because a drop is a break, not a research result [6]. As capability cases saturate, graduate them into the regression suite [6].

Run the regression suite on every prompt edit, tool-schema change, retrieval change, and model swap. OpenAI lists continuous evaluation on every change as a first-class step, not a later optimization [7]. If you skip it, you will learn about the regression from a user.

When a new model lands, do not try it in production and hope. Anthropic’s observation from customer work is that teams with an eval suite can judge a new model, tune prompts, and upgrade in days, while teams without one spend weeks on manual testing [6]. Run both suites, keep cost and latency as separate metrics (not substitutes for quality), and only then decide.

Read transcripts, not just scores. A fail can be a genuine agent mistake or a grader that rejected a valid solution [6]. If the suite is saturated at 100%, it is a regression net, not an improvement signal [6]. Add harder or newer live cases rather than celebrating a ceiling.

A quality SLA operations can actually staff

A quality SLA is not an uptime clone. Uptime asks whether the endpoint responded. A quality SLA asks whether a sampled, graded slice of work stayed inside agreed bounds, who grades it, and what happens when it does not.

NIST’s MEASURE 3.1 subcategory calls for personnel and documentation to track existing, unanticipated, and emergent risks based on intended and actual performance in deployed contexts [1]. MEASURE 3.3 adds user-report channels that feed evaluation metrics [1]. The OECD accountability principle is the governance overlay: proper functioning has to be demonstrable through documentation and analysis of outputs, not asserted [4].

Write the SLA so a real roster can keep it. A pattern we use with operations teams (including leaders we work with from Charlotte and Philadelphia) looks like this:

  • Before every change: the frozen golden set must not regress on task success or policy violations. Capability scores may move. Policy and core task success may not, unless the change is an explicit, reviewed tradeoff.
  • Every week: sample a fixed number of live traces. Score task success, helpfulness, and policy violations. File misses into the review queue with the trace attached.
  • Every month: a human calibration set to check that the LLM-as-judge still agrees with a domain expert. OpenAI lists ignoring human feedback, and failing to calibrate automated scoring, as anti-patterns [7]. Anthropic’s production pattern is LLM graders with product-owned criteria and periodic human calibration [6].
  • On miss: a named owner, a time box to reproduce, and a rule for whether the agent stays up or a skill is disabled. Evals feed that review queue. They do not replace a human stop or approve step.

Staff the work in hours, not slogans. Someone has to own the golden set, read a sample of transcripts, and decide when a live miss becomes a new test. If that work has no hours on the roster, you do not have a quality SLA.

Practical takeaways

  • Treat evals as the scoreboard for AI agents. Controls stop some bad actions. They do not tell you if quality is rising.
  • Build the first golden set from real failures and live logs (20 to 50 unambiguous tasks), then grow it from confirmed misses [6].
  • Run offline evals on every change and sample live traffic on a cadence. NIST wants both pre-deployment testing and in-production monitoring [1].
  • Report task-success, helpfulness, and policy-violation rates separately, plus a reliability view (pass^k) where consistency matters [6].
  • Keep a near-perfect regression suite for work the agent already owns, and graduate saturated capability tests into it [6].
  • Write a quality SLA with sample sizes, owners, hours, thresholds, and a miss procedure. Calibrate automated graders against humans on a schedule [7].

How we can help

Our team designs the measurement layer around production AI agents: golden sets from your real cases, offline and online eval loops, and quality SLAs that operations can staff. That work sits alongside the AI tools and agent workflows we already run for clients. Have more questions or want to get in touch? Reach us through our contact page, call (980) 322-4500, or email [email protected].

Citations

  1. NIST, "Artificial Intelligence Risk Management Framework (AI RMF 1.0)" (2023-01)
  2. NIST, "Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1)" (2024-07)
  3. OECD.AI, "Robustness, security and safety (Principle 1.4)" (2024-05)
  4. OECD.AI, "Accountability (Principle 1.5)" (2024-05)
  5. Stanford CRFM, "Language Models are Changing AI: The Need for Holistic Evaluation" (2022-11-17)
  6. Anthropic, "Demystifying evals for AI agents" (2026-01-09)
  7. OpenAI, "Evaluation best practices" (2026)
  8. Google Cloud, "Manage evaluation metrics" (2026-09-03)
Our Strongest Offering

Forge Your Next Website

Forged Sites are custom-built, static-first websites with a full AI content engine on board — no CMS to log into, no plugins to break, no builder to fight.

  • Working target: WCAG 2.2 AA
  • During work hours, an account manager still reviews material changes
  • DraftDash auto-drafted blogs keep your content engine running
  • Ethel AI-powered forms filter spam and capture genuine leads