# Agent Evals overview

> Connect outcome scores to the runs that produced them so you can compare changes using production results.

A support answer can pass every code check and still fail to help the customer. Agent Evals lets you score the result that matters, link it to the run that produced the answer, and compare changes using those outcomes.

Inngest already records each durable run, its steps, and its [trace](/docs-markdown/platform-and-operations/traces). Agent Evals adds outcome signals to those runs, including signals that arrive days after the work finishes. When a score is bad, you open the execution behind it instead of guessing from a final output.

## Score your first run

Add `scoreMiddleware()` to your client, then record a score from inside a function:

```typescript {{ title: "TypeScript" }}
import { Inngest } from "inngest";
import { scoreMiddleware } from "inngest/experimental";

export const inngest = new Inngest({
  id: "support-agent",
  middleware: [scoreMiddleware()],
});

export default inngest.createFunction(
  { id: "answer-support-ticket", triggers: { event: "support/ticket.created" } },
  async ({ event, step }) => {
    const answer = await step.run("generate-answer", () =>
      generateAnswer(event.data.message)
    );

    await step.score("score-answer-valid", {
      name: "answer-valid",
      value: validateAnswer(answer),
    });

    return answer;
  }
);
```

```python {{ title: "Python" }}
# Requires the inngest release after 0.5.19.
import inngest

from .stubs import generate_answer, validate_answer

inngest_client = inngest.Inngest(app_id="support-agent")

@inngest_client.create_function(
    fn_id="answer-support-ticket",
    trigger=inngest.TriggerEvent(event="support/ticket.created"),
)
async def answer_support_ticket(ctx: inngest.Context) -> str:
    message = str(ctx.event.data["message"])

    async def generate() -> str:
        return await generate_answer(message)

    answer = await ctx.step.run("generate-answer", generate)

    # Wrap the score write in a step so a retry or replay doesn't record
    # it twice. Pass the run ID explicitly.
    async def score_answer() -> None:
        await inngest_client.score(
            name="answer-valid",
            value=validate_answer(answer),
            run_id=ctx.run_id,
        )

    await ctx.step.run("score-answer-valid", score_answer)

    return answer
```

`step.score()` is a durable step, so a retry or replay doesn't record the score twice. The score appears on the run's trace and in the function's [score trends](/docs-markdown/agent-evals/scores#where-scores-appear).

## Grow from one score to a full eval

Most teams start with a single score and add the other parts as their questions get harder. Each part answers a new question:

1. **[Scores](/docs-markdown/agent-evals/scores): is this result good?** `step.score()` records a named number or boolean on a run or step. Use it for checks you can make right away, such as a guardrail, a validation, or an LLM judge.
2. **[Deferred scoring](/docs-markdown/agent-evals/deferred-scoring): did it help later?** `defer()` starts a scorer that waits for feedback, a review, or a conversion, then scores the run that started it. The original run doesn't stay open while it waits.
3. **[Experiments](/docs-markdown/agent-evals/experiments): is the change better?** `group.experiment()` picks one variant per run and keeps it on retries. Pass the returned `experimentRef` with a score so the outcome counts toward the variant that served it.
4. **[Sessions](/docs-markdown/agent-evals/sessions): what else happened?** A session ID on your events groups the runs behind one conversation, ticket, or job, so you can inspect everything that led to an outcome.

### Example: a support agent

A support agent answers tickets, and you're testing a new answer strategy against the current one:

- An experiment sends each ticket through one of the two strategies.
- The agent checks its answer and records `answer-valid` with `step.score()`.
- The agent defers a scorer that waits for the customer's rating. When the rating arrives, the scorer records `customer-helped` on the original run and credits the variant that served it.
- A `ticket_id` session collects every run for the ticket, including follow-ups.

In the experiment view, you compare both scores for each variant: **did the agent produce a valid answer**, and **did the answer help the customer?**

## Choose the right tool

| Goal                                                                            | Use                                                                                            |
| ------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------- |
| Score a guardrail, validation check, model confidence, or inline LLM judge      | [`step.score()`](/docs-markdown/agent-evals/scores)                                            |
| Wait for user feedback, ticket resolution, conversion, or review before scoring | [Deferred scoring](/docs-markdown/agent-evals/deferred-scoring)                                |
| Score a past run from another service that already knows the outcome            | [`inngest.score()` with `runId`](/docs-markdown/agent-evals/scores#score-from-outside-the-run) |
| Compare prompts, models, providers, tools, or workflow rewrites                 | [Experiments](/docs-markdown/agent-evals/experiments)                                          |
| Keep a user, account, or tenant on the same variant while testing               | [`experiment.bucket()`](/docs-markdown/agent-evals/experiments#choose-an-assignment-strategy)  |
| Find all runs for one conversation, ticket, import, or agent task               | [Sessions](/docs-markdown/agent-evals/sessions)                                                |
| Grade answers with a model                                                      | [Score with an LLM judge](/docs-markdown/agent-evals/guides/llm-judge)                         |

## Where results appear

| Surface                                                                       | Use it to                                                                                          |
| ----------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------- |
| [Traces](/docs-markdown/platform-and-operations/traces)                       | Inspect a run's steps, model calls, tool calls, errors, the selected variant, and its scores.      |
| [Function dashboard](/docs-markdown/agent-evals/scores#where-scores-appear)   | Watch score trends over time. A drop after a prompt change tells you something broke.              |
| [Experiment view](/docs-markdown/agent-evals/experiments#score-a-variant)     | Compare run counts and score aggregates for each variant.                                          |
| [**AI → Sessions**](/docs-markdown/agent-evals/sessions#inspect-related-runs) | Find every run related to a conversation, ticket, or job.                                          |
| [Insights](/docs-markdown/platform-and-operations/insights)                   | Query historical events, runs, and steps with SQL.                                                 |
| [Extended Traces](/docs-markdown/reference/typescript/v4/extended-traces)     | Capture spans from your model SDKs, databases, and HTTP calls inside each step with OpenTelemetry. |
| [Inngest MCP](/docs-markdown/ai-dev-tools/mcp)                                | Read experiment run counts and score aggregates from a coding agent.                               |

## Watch it in action

## Next steps

- [Quick start](/docs-markdown/agent-evals/quick-start) builds a scored experiment with later feedback.
- [Guides](/docs-markdown/agent-evals/guides) cover LLM judges, user feedback, model comparisons, and rollouts.
- [Reference](/docs-markdown/agent-evals/reference) lists every SDK call and attribution rule.