• Documentation
  • Examples
  • Patterns
Contact salesSign Up
  • Documentation
  • Examples
  • Patterns
  • Overview
    • All patterns
  • AI Evals
    • Run experiments in production
    • Score agents on real outcomes
    • Score agent accuracy
  • Durable Workflows
    • Reliably run critical workflows
  • Flow Control
    • Flash sales and bursty workflows
  • Event Coordination
    • Building flows for lost customers
    • Reliable scheduling systems
    • Running functions in parallel
  • Scheduling
    • Running code at specific times
  • Background Jobs
    • Build reliable webhooks
    • Keeping your API fast
  • Contact salesSign Up
00 · Experiment, score, pick winners

AI Evals

Run experiments on live traffic, keep cohorts stable, and compare variants against real outcome signals. The substrate for evaluating models, prompts, and rewrites in production.

EVENTsend.emailupdate.crmscore.leadnotify.slack
eventuser.signed_up · 4 listeners

Patterns

01Run experiments in productionUse group.experiment() to split traffic, keep cohorts stable, compare variants, and roll changes forward safely.02Score agents on real outcomesAttach scores to runs and steps, and defer scoring until the real-world signal arrives, so you measure agents on what actually happened.03Score agent accuracyGrade an agent against a known answer (exact match, set overlap, or numeric tolerance) and record the score on the run that produced it.
Next primitive →Durable Workflows

Was this page helpful?

© 2026 Inngest Inc. All rights reserved.
We're hiring!
Star our open source repositoryJoin our Discord communityFollow us on XFollow us on Bluesky