PlatformAgent Evals
Agent Evals guides
Grade answers with a model, score user feedback, compare models, roll out rewrites, and control eval costs.
These guides build on the Quick start. Each one covers a single task from start to finish.
Score results
- Score with an LLM judge grades answers against a rubric with a model, and adds a cheap deterministic check.
- Score user feedback records a thumbs-up, a resolved ticket, or ground truth that arrives after the run.
Compare changes
- Compare models and prompts splits traffic between two answer strategies and scores each by later feedback.
- Roll out a workflow rewrite canaries a multi-step path and ramps it up as results come in.
- Read results and roll out explains how to compare variants and decide when to ship.
Operate evals
- Manage eval costs shows how to count runs, steps, scores, and model calls, and how to choose which runs to score.
Next steps
- Best practices lists the measurement rules that keep comparisons meaningful.
- Troubleshooting helps find missing or misattributed scores.