PlatformAgent Evals

Agent Evals limits

Plan Agent Evals usage around score volume, scoring executions, and trace retention.

Before scoring production traffic, estimate how many scores you will write and how many function executions your scorers need. Check trace retention so you know how long you can inspect the runs behind a result. Use the current pricing page for Cloud plan allowances.

Cloud plan allowances

PlanScores ingestedExecutionsTrace and log history
Free10,000 included per month50,000 included per month24 hours
Pro50,000 included per month, then $1.50 per 1,0001 million included per month, with add-on capacity7 days
Business250,000 included per month10 million included per month, with add-on capacity14 days
EnterpriseCustomCustom90 days

A score is one ingested outcome value. If you write several scores for one run, each write counts toward monthly score usage. The included score quantities are allowances, not hard caps. Customers can keep writing scores under usage-based billing. Pricing defines an execution as one function run or one step execution. A deferred scorer runs as its own function, so include its run and steps when you estimate executions. If it writes a score, include that score too. Check pricing before making a purchasing decision because allowances and overage terms can change.

Application code chooses which runs to score. Agent Evals doesn't sample scores or cap score spend for you. Manage eval costs shows how to choose.

Session metadata limits

Sessions make related runs easy to find. Add no more than 5 entries to an event's meta.sessions. A non-batched run can belong to up to 5 sessions. A batched run can combine up to 25 unique session key and ID pairs from its events. A session key can be up to 128 bytes and a session ID up to 512 bytes. Keys and IDs cannot be empty. IDs must be strings or finite numbers; Inngest converts numbers to strings. See Supported values and Session propagation.

Sessions use an asynchronous search index. A recently sent event or started run may take time to appear in session search. Session search is an investigation view, not a realtime status signal.

Field limits

FieldLimit
Score name1 to 128 bytes (UTF-8); no control characters or single quotes
Score valueA finite number or a boolean
Session key128 bytes
Session ID512 bytes
Sessions per event5
Sessions per batched run25 unique key and ID pairs

Keep a score's name and meaning consistent across runs you compare. Scoring is in beta.

Score and experiment constraints

A deferred scorer is an independent function run with its own retries, concurrency, and steps. It can't use onFailure or event batching. Each defer() call uses an ID unique within its parent run. Deferred functions are experimental. There is no separate Agent Evals quota for deferred scorers. Normal function and plan limits still apply to their runs.

group.experiment() selects one variant per run and memoizes the choice across retries. Each variant callback must call at least one step.* tool. This rule applies whether a variant wraps one step or a larger workflow path. Experiments shows the variant rules in use.