Agent Evals overview
Connect outcome scores to the runs that produced them so you can compare changes using production results.
- Sessions on eventsv4.7+(supported)
- Experiments and variantsv4.8+(supported)
- Run and step scoresv4.8+ Beta(supported)
- Deferred scoringv4.8+ Beta(supported)
Scoring and deferred scoring APIs may change before general availability.
- Sessions on eventsNext release(supported)
- Experiments and variantsNext release(supported)
- Run and step scoresNext release(supported)
- Deferred scoringNext release(supported)
Available in the Python SDK release after 0.5.19. No step.score(), createScorer(), or weighted and custom selectors yet.
A support answer can pass every code check and still fail to help the customer. Agent Evals lets you score the result that matters, link it to the run that produced the answer, and compare changes using those outcomes.
Inngest already records each durable run, its steps, and its trace. Agent Evals adds outcome signals to those runs, including signals that arrive days after the work finishes. When a score is bad, you open the execution behind it instead of guessing from a final output.
Score your first run
Add scoreMiddleware() to your client, then record a score from inside a function:
import { Inngest } from "inngest";
import { scoreMiddleware } from "inngest/experimental";
export const inngest = new Inngest({
id: "support-agent",
middleware: [scoreMiddleware()],
});
export default inngest.createFunction(
{ id: "answer-support-ticket", triggers: { event: "support/ticket.created" } },
async ({ event, step }) => {
const answer = await step.run("generate-answer", () =>
generateAnswer(event.data.message)
);
await step.score("score-answer-valid", {
name: "answer-valid",
value: validateAnswer(answer),
});
return answer;
}
);
step.score() is a durable step, so a retry or replay doesn't record the score twice. The score appears on the run's trace and in the function's score trends.
Grow from one score to a full eval
Most teams start with a single score and add the other parts as their questions get harder. Each part answers a new question:
- Scores: is this result good?
step.score()records a named number or boolean on a run or step. Use it for checks you can make right away, such as a guardrail, a validation, or an LLM judge. - Deferred scoring: did it help later?
defer()starts a scorer that waits for feedback, a review, or a conversion, then scores the run that started it. The original run doesn't stay open while it waits. - Experiments: is the change better?
group.experiment()picks one variant per run and keeps it on retries. Pass the returnedexperimentRefwith a score so the outcome counts toward the variant that served it. - Sessions: what else happened? A session ID on your events groups the runs behind one conversation, ticket, or job, so you can inspect everything that led to an outcome.
Example: a support agent
A support agent answers tickets, and you're testing a new answer strategy against the current one:
- An experiment sends each ticket through one of the two strategies.
- The agent checks its answer and records
answer-validwithstep.score(). - The agent defers a scorer that waits for the customer's rating. When the rating arrives, the scorer records
customer-helpedon the original run and credits the variant that served it. - A
ticket_idsession collects every run for the ticket, including follow-ups.
In the experiment view, you compare both scores for each variant: did the agent produce a valid answer, and did the answer help the customer?
Choose the right tool
| Goal | Use |
|---|---|
| Score a guardrail, validation check, model confidence, or inline LLM judge | step.score() |
| Wait for user feedback, ticket resolution, conversion, or review before scoring | Deferred scoring |
| Score a past run from another service that already knows the outcome | inngest.score() with runId |
| Compare prompts, models, providers, tools, or workflow rewrites | Experiments |
| Keep a user, account, or tenant on the same variant while testing | experiment.bucket() |
| Find all runs for one conversation, ticket, import, or agent task | Sessions |
| Grade answers with a model | Score with an LLM judge |
Where results appear
| Surface | Use it to |
|---|---|
| Traces | Inspect a run's steps, model calls, tool calls, errors, the selected variant, and its scores. |
| Function dashboard | Watch score trends over time. A drop after a prompt change tells you something broke. |
| Experiment view | Compare run counts and score aggregates for each variant. |
| AI → Sessions | Find every run related to a conversation, ticket, or job. |
| Insights | Query historical events, runs, and steps with SQL. |
| Extended Traces | Capture spans from your model SDKs, databases, and HTTP calls inside each step with OpenTelemetry. |
| Inngest MCP | Read experiment run counts and score aggregates from a coding agent. |
Watch it in action
Next steps
- Quick start builds a scored experiment with later feedback.
- Guides cover LLM judges, user feedback, model comparisons, and rollouts.
- Reference lists every SDK call and attribution rule.