# Compare models and prompts

> Split production traffic between two answer strategies and compare them by the feedback each one earns.

You have a new prompt or model and want to know whether it helps customers more than the current one. This guide sends each account to one of two strategies, records an immediate check, and credits later customer feedback to the variant that served the answer.

## 1. Define the variants

Put each strategy in a variant callback. `experiment.bucket()` keeps an account on the same strategy across tickets, so a customer doesn't bounce between styles:

```typescript {{ title: "TypeScript", filename: "inngest/answer.ts" }}
import { experiment } from "inngest";
import { inngest } from "./client";
import { feedbackScorer } from "./scorers";

export const answerTicket = inngest.createFunction(
  { id: "answer-ticket", triggers: { event: "support/ticket.created" } },
  async ({ event, step, group, defer }) => {
    const { result: answer, experimentRef } = await group.experiment(
      "answer-style",
      {
        variants: {
          concise: () =>
            step.run("answer-concise", () =>
              answer({ prompt: CONCISE_PROMPT, ticket: event.data })
            ),
          detailed: () =>
            step.run("answer-detailed", () =>
              answer({ prompt: DETAILED_PROMPT, ticket: event.data })
            ),
        },
        select: experiment.bucket(event.data.accountId, {
          weights: { concise: 50, detailed: 50 },
        }),
      }
    );

    await step.run("send-answer", () => sendReply(event.data.ticketId, answer));

    // Immediate check, credited to the variant.
    await inngest.score.experiment({
      name: "answer-valid",
      value: isValidAnswer(answer),
      experiment: experimentRef,
    });

    // Later outcome, credited to the same variant.
    defer("score-feedback", {
      function: feedbackScorer,
      data: { ticketId: event.data.ticketId },
      experiment: experimentRef,
    });

    return answer;
  }
);
```

```python {{ title: "Python", filename: "answer.py" }}
# Requires the inngest release after 0.5.19.
import inngest
from inngest.experimental import experiment

from .client import inngest_client
from .scorers import feedback_scorer
from .stubs import (
    CONCISE_PROMPT,
    DETAILED_PROMPT,
    answer,
    is_valid_answer,
    send_reply,
)

@inngest_client.create_function(
    fn_id="answer-ticket",
    trigger=inngest.TriggerEvent(event="support/ticket.created"),
)
async def answer_ticket(ctx: inngest.Context) -> str:
    ticket = ctx.event.data
    ticket_id = str(ticket["ticketId"])

    async def answer_concise() -> str:
        return await answer(prompt=CONCISE_PROMPT, ticket=ticket)

    async def answer_detailed() -> str:
        return await answer(prompt=DETAILED_PROMPT, ticket=ticket)

    res = await ctx.group.experiment(
        "answer-style",
        variants={
            "concise": lambda: ctx.step.run("answer-concise", answer_concise),
            "detailed": lambda: ctx.step.run(
                "answer-detailed", answer_detailed
            ),
        },
        select=experiment.bucket(
            str(ticket["accountId"]),
            weights={"concise": 50, "detailed": 50},
        ),
    )
    reply = res.result

    async def send() -> None:
        await send_reply(ticket_id, reply)

    await ctx.step.run("send-answer", send)

    # Immediate check, credited to the variant.
    async def score_valid() -> None:
        await inngest_client.score_experiment(
            name="answer-valid",
            value=is_valid_answer(reply),
            experiment=res.experiment_ref,
            run_id=ctx.run_id,
        )

    await ctx.step.run("score-answer-valid", score_valid)

    # Later outcome, credited to the same variant.
    ctx.defer(
        "score-feedback",
        function=feedback_scorer,
        data={"ticketId": ticket_id},
        experiment=res.experiment_ref,
    )

    return reply
```

To compare models instead of prompts, change the model in each callback. To compare providers, call a different SDK. Change one thing at a time so the result tells you which change mattered.

## 2. Wait for feedback

The scorer waits for the customer's rating and returns a score. Because `defer()` received the `experimentRef`, Inngest credits the score to the variant that served the answer:

```typescript {{ title: "TypeScript", filename: "inngest/scorers.ts" }}
import { createScorer } from "inngest/experimental";
import { z } from "zod";
import { inngest } from "./client";

export const feedbackScorer = createScorer(
  inngest,
  { id: "feedback-scorer", schema: z.object({ ticketId: z.string() }) },
  async ({ event, step }) => {
    const feedback = await step.waitForEvent("wait-for-feedback", {
      event: "support/feedback.received",
      timeout: "7d",
      if: `async.data.ticketId == '${event.data.ticketId}'`,
    });
    if (!feedback) return null;
    return { name: "customer-helpful", value: feedback.data.helpful };
  }
);
```

```python {{ title: "Python", filename: "scorers.py" }}
# Requires the inngest release after 0.5.19.
import datetime

import inngest
from inngest.experimental import create_defer

from .client import inngest_client

@create_defer(inngest_client, fn_id="feedback-scorer")
async def feedback_scorer(ctx: inngest.Context) -> None:
    ticket_id = str(ctx.event.data["ticketId"])
    feedback = await ctx.step.wait_for_event(
        "wait-for-feedback",
        event="support/feedback.received",
        timeout=datetime.timedelta(days=7),
        if_exp=f"async.data.ticketId == '{ticket_id}'",
    )
    if feedback is None:
        return

    # Credit the score to the variant that ctx.defer() received.
    parent = ctx.parents[0]
    helpful = feedback.data.get("helpful") is True

    async def write_score() -> None:
        if parent.experiment is not None:
            await inngest_client.score_experiment(
                name="customer-helpful",
                value=helpful,
                experiment=parent.experiment,
                run_id=parent.run_id,
            )
        else:
            await inngest_client.score(
                name="customer-helpful", value=helpful, run_id=parent.run_id
            )

    await ctx.step.run("score", write_score)
```

## 3. Compare the variants

Open the experiment in the dashboard to see run counts and score aggregates for each variant, or read them with the `get_experiment` tool in the [Inngest MCP server](/docs-markdown/ai-dev-tools/mcp). Compare `customer-helpful` only after both variants have had the same seven days to collect feedback, and keep the count of runs still waiting beside every average. Open a few traces from each variant to see why answers scored the way they did.

[Read results and roll out](/docs-markdown/agent-evals/guides/interpreting-results) walks through the comparison and the rollout decision.

## 4. Ramp or roll back

When the candidate meets your criteria, raise its weight in stages, for example `{ concise: 20, detailed: 80 }`, then `{ concise: 0, detailed: 100 }`. Changing bucket weights can move an account to the other variant. If an account must never switch back, store assignments in your database and select them with `experiment.custom()`. To stop the test quickly, pin the safe variant with `experiment.fixed("concise")`.

## Next steps

- [Experiments](/docs-markdown/agent-evals/experiments) explains every assignment strategy.
- [Score with an LLM judge](/docs-markdown/agent-evals/guides/llm-judge) adds a model-graded score while you wait for feedback.