Platform

Agent Evals overview

Connect outcome scores to the runs that produced them so you can compare changes using production results.

SDK support
TypeScript
PythonNext release

A support answer can pass every code check and still fail to help the customer. Agent Evals lets you score the result that matters, link it to the run that produced the answer, and compare changes using those outcomes.

Inngest already records each durable run, its steps, and its trace. Agent Evals adds outcome signals to those runs, including signals that arrive days after the work finishes. When a score is bad, you open the execution behind it instead of guessing from a final output.

Score your first run

Add scoreMiddleware() to your client, then record a score from inside a function:

import { Inngest } from "inngest";
import { scoreMiddleware } from "inngest/experimental";

export const inngest = new Inngest({
  id: "support-agent",
  middleware: [scoreMiddleware()],
});

export default inngest.createFunction(
  { id: "answer-support-ticket", triggers: { event: "support/ticket.created" } },
  async ({ event, step }) => {
    const answer = await step.run("generate-answer", () =>
      generateAnswer(event.data.message)
    );

    await step.score("score-answer-valid", {
      name: "answer-valid",
      value: validateAnswer(answer),
    });

    return answer;
  }
);

step.score() is a durable step, so a retry or replay doesn't record the score twice. The score appears on the run's trace and in the function's score trends.

Grow from one score to a full eval

Most teams start with a single score and add the other parts as their questions get harder. Each part answers a new question:

  1. Scores: is this result good? step.score() records a named number or boolean on a run or step. Use it for checks you can make right away, such as a guardrail, a validation, or an LLM judge.
  2. Deferred scoring: did it help later? defer() starts a scorer that waits for feedback, a review, or a conversion, then scores the run that started it. The original run doesn't stay open while it waits.
  3. Experiments: is the change better? group.experiment() picks one variant per run and keeps it on retries. Pass the returned experimentRef with a score so the outcome counts toward the variant that served it.
  4. Sessions: what else happened? A session ID on your events groups the runs behind one conversation, ticket, or job, so you can inspect everything that led to an outcome.

Example: a support agent

A support agent answers tickets, and you're testing a new answer strategy against the current one:

  • An experiment sends each ticket through one of the two strategies.
  • The agent checks its answer and records answer-valid with step.score().
  • The agent defers a scorer that waits for the customer's rating. When the rating arrives, the scorer records customer-helped on the original run and credits the variant that served it.
  • A ticket_id session collects every run for the ticket, including follow-ups.

In the experiment view, you compare both scores for each variant: did the agent produce a valid answer, and did the answer help the customer?

Choose the right tool

GoalUse
Score a guardrail, validation check, model confidence, or inline LLM judgestep.score()
Wait for user feedback, ticket resolution, conversion, or review before scoringDeferred scoring
Score a past run from another service that already knows the outcomeinngest.score() with runId
Compare prompts, models, providers, tools, or workflow rewritesExperiments
Keep a user, account, or tenant on the same variant while testingexperiment.bucket()
Find all runs for one conversation, ticket, import, or agent taskSessions
Grade answers with a modelScore with an LLM judge

Where results appear

SurfaceUse it to
TracesInspect a run's steps, model calls, tool calls, errors, the selected variant, and its scores.
Function dashboardWatch score trends over time. A drop after a prompt change tells you something broke.
Experiment viewCompare run counts and score aggregates for each variant.
AI → SessionsFind every run related to a conversation, ticket, or job.
InsightsQuery historical events, runs, and steps with SQL.
Extended TracesCapture spans from your model SDKs, databases, and HTTP calls inside each step with OpenTelemetry.
Inngest MCPRead experiment run counts and score aggregates from a coding agent.

Watch it in action

Next steps

  • Quick start builds a scored experiment with later feedback.
  • Guides cover LLM judges, user feedback, model comparisons, and rollouts.
  • Reference lists every SDK call and attribution rule.