PlatformAgent Evals

Agent Evals guides

Grade answers with a model, score user feedback, compare models, roll out rewrites, and control eval costs.

These guides build on the Quick start. Each one covers a single task from start to finish.

Score results

  • Score with an LLM judge grades answers against a rubric with a model, and adds a cheap deterministic check.
  • Score user feedback records a thumbs-up, a resolved ticket, or ground truth that arrives after the run.

Compare changes

Operate evals

  • Manage eval costs shows how to count runs, steps, scores, and model calls, and how to choose which runs to score.

Next steps