# `group.experiment`

`group.experiment()` selects one of several named variants for a run. Inngest keeps that selection through retries, so you can compare changes on real traffic without switching a run between variants.

They are useful when:

- You want to roll out a new step or workflow path to a small share of runs.
- You need to compare models, prompts, or providers.
- You want the same user or account to receive a consistent variant across runs.
- You need to score outcomes and compare the variants that produced them.

`group.experiment()` is in beta. It requires TypeScript SDK v4.8.0 or later. In Python, `ctx.group.experiment()` requires the release after 0.5.19 and supports the `fixed` and `bucket` selectors. You don't need a separate package or feature-flag service.

## Run an experiment

Import the `experiment` helper from `inngest`. Then call `group.experiment()` in your function handler. The handler receives `group` alongside `event` and `step`.

```typescript {{ title: "TypeScript" }}
import { experiment } from "inngest";

const { result, variant, experimentRef } = await group.experiment(
  "answer-model",
  {
    variants: {
      current: () => step.run("answer-current", () => answerWithCurrentModel()),
      new: () => step.run("answer-new", () => answerWithNewModel()),
    },
    select: experiment.weighted({ current: 90, new: 10 }),
  }
);
```

```python {{ title: "Python" }}
# Requires the inngest release after 0.5.19.
from inngest.experimental import experiment

# Python has no weighted() selector. Bucketing on the run ID gives each
# run its own weighted assignment.
res = await ctx.group.experiment(
    "answer-model",
    variants={
        "current": lambda: ctx.step.run(
            "answer-current", answer_with_current_model
        ),
        "new": lambda: ctx.step.run("answer-new", answer_with_new_model),
    },
    select=experiment.bucket(
        ctx.run_id, weights={"current": 90, "new": 10}
    ),
)
result, variant, experiment_ref = (
    res.result,
    res.variant,
    res.experiment_ref,
)
```

Only the selected callback runs. `result` contains its return value, `variant` names the selected path, and `experimentRef` lets you credit a score to that path. Each callback must call at least one `step.*` method. Put side effects inside steps because code outside steps can run again on replay.

## How selection works

The selection is a durable, memoized step. The first time a run reaches `group.experiment()`, Inngest evaluates the `select` strategy and saves the chosen variant. When the run retries or replays, it reuses that variant.

A variant callback can use any step tool, including `step.run()`, `step.invoke()`, `step.waitForEvent()`, and `step.sendEvent()`. Each of those steps retries and memoizes like any other step. If a callback finishes without calling a step tool, the SDK throws a `NonRetriableError`.

Each `group.experiment()` call counts as one step in your run and billing metrics. Steps inside the selected callback count separately.

## Test a larger workflow path

Use the same API when each variant has several steps. Only the selected path runs:

```typescript {{ title: "TypeScript" }}
const { result, experimentRef } = await group.experiment("invoice-engine", {
  variants: {
    current: async () => {
      const invoice = await step.run("generate-current", () =>
        generateInvoiceV1(event.data)
      );
      await step.run("send-current", () => sendInvoice(invoice));
      return invoice;
    },
    rewrite: async () => {
      const invoice = await step.run("generate-rewrite", () =>
        generateInvoiceV2(event.data)
      );
      await step.run("send-rewrite", () => sendInvoice(invoice));
      return invoice;
    },
  },
  select: experiment.weighted({ current: 99, rewrite: 1 }),
});
```

```python {{ title: "Python" }}
# Requires the inngest release after 0.5.19.
async def current() -> dict[str, str]:
    invoice = await ctx.step.run(
        "generate-current", generate_invoice_v1, ctx.event.data
    )
    await ctx.step.run("send-current", send_invoice, invoice)
    return invoice

async def rewrite() -> dict[str, str]:
    invoice = await ctx.step.run(
        "generate-rewrite", generate_invoice_v2, ctx.event.data
    )
    await ctx.step.run("send-rewrite", send_invoice, invoice)
    return invoice

res = await ctx.group.experiment(
    "invoice-engine",
    variants={"current": current, "rewrite": rewrite},
    select=experiment.bucket(
        ctx.run_id, weights={"current": 99, "rewrite": 1}
    ),
)
```

`group.experiment()` records one selection step. It also meters each step the selected path runs. Give each experiment a unique, stable ID within a function, and keep variant names stable so traces and outcome comparisons remain clear.

## Choose a variant

- **Split new runs:** `experiment.weighted({ current: 90, new: 10 })` assigns variants by relative weights. Change the weights to ramp up a new path; existing runs keep their choice.
- **Keep a cohort together:** `experiment.bucket(userId, { weights: { current: 90, new: 10 } })` maps a stable user or account ID to a variant across runs.
- **Use your own assignment:** `experiment.custom(fn)` reads a feature flag or database record. Use it when assignment must change without a deploy or must stay fixed through a migration.
- **Pin one variant:** `experiment.fixed("current")` always selects the same path for an override or test.

### Split new runs with `weighted`

Each new run gets its own assignment. Weights are relative, so `{ control: 9, candidate: 1 }` and `{ control: 90, candidate: 10 }` give the same split.

```typescript {{ title: "TypeScript" }}
select: experiment.weighted({ control: 90, candidate: 10 })
```

```python {{ title: "Python" }}
# Requires the inngest release after 0.5.19.
# Python has no weighted(). Bucket on the run ID for a per-run split.
select=experiment.bucket(
    ctx.run_id, weights={"control": 90, "candidate": 10}
),
```

Selection is seeded by the run ID, so a run keeps its variant on retries. New weights in a later deploy only affect new selections. Because the seed is the run, not the user, one user can get different variants across runs.

### Keep a user on one variant with `bucket`

Pass a stable value, such as a user, account, or tenant ID. Inngest hashes it to a variant, so a user doesn't see a different experience each time they trigger the function.

```typescript {{ title: "TypeScript" }}
select: experiment.bucket(event.data.userId, {
  weights: { control: 80, candidate: 20 },
})
```

```python {{ title: "Python" }}
# Requires the inngest release after 0.5.19.
select=experiment.bucket(
    str(ctx.event.data["userId"]),
    weights={"control": 80, "candidate": 20},
),
```

Pass a present, stable key to `experiment.bucket()`; a missing key sends those runs to the same bucket. Changing bucket weights can change a user's future assignment. If a user must never switch back, store the assignment and use `experiment.custom()`.

### Read the assignment from your system with `custom`

Use a feature flag, rollout table, entitlement, or stored migration assignment. The selector can be sync or async, and its result is memoized for the run.

```typescript {{ title: "TypeScript" }}
select: experiment.custom(async () => {
  const assignment = await rolloutTable.get(event.data.accountId);
  return assignment ?? "control";
})
```

```python {{ title: "Python" }}
# Requires the inngest release after 0.5.19.
from inngest.experimental import experiment

# Python has no custom() selector. Read the assignment in a step, then
# pin it with fixed(). The step memoizes the assignment for the run.
async def read_assignment() -> str:
    assignment = await rollout_table.get(ctx.event.data["accountId"])
    return assignment or "control"

assignment = await ctx.step.run("read-assignment", read_assignment)

res = await ctx.group.experiment(
    "copy-style",
    variants=variants,
    select=experiment.fixed(assignment),
)
```

Its selector must return a name in `variants`. An unknown name fails the run.

### Pin one variant with `fixed`

Always select the same variant. Use it for a manual override, to test one path, or to pin a winner while you remove the experiment.

```typescript {{ title: "TypeScript" }}
select: experiment.fixed("candidate")
```

```python {{ title: "Python" }}
# Requires the inngest release after 0.5.19.
select=experiment.fixed("candidate"),
```

## Use the result and variant name

`group.experiment()` always returns `{ result, variant, experimentRef }`:

- **`result`**: the selected callback's return value.
- **`variant`**: the selected variant's name. Use it for logging, analytics, or later decisions.
- **`experimentRef`**: a replay-stable handle, `{ experimentName, variant }`. Pass it with a score to credit the variant that served this run.

If your team already uses an analytics warehouse, record the variant from a step:

```typescript {{ title: "TypeScript" }}
const { result, variant, experimentRef } = await group.experiment("copy-style", {
  variants: {
    short: () => step.run("short-copy", () => generateShortCopy(event.data)),
    detailed: () =>
      step.run("detailed-copy", () => generateDetailedCopy(event.data)),
  },
  select: experiment.bucket(event.data.userId, {
    weights: { short: 50, detailed: 50 },
  }),
});

await step.run("track-experiment", () =>
  analytics.track("experiment.variant_selected", {
    experiment: "copy-style",
    variant,
    userId: event.data.userId,
  })
);

return result;
```

```python {{ title: "Python" }}
# Requires the inngest release after 0.5.19.
from inngest.experimental import experiment

res = await ctx.group.experiment(
    "copy-style",
    variants={
        "short": lambda: ctx.step.run(
            "short-copy", generate_short_copy, ctx.event.data
        ),
        "detailed": lambda: ctx.step.run(
            "detailed-copy", generate_detailed_copy, ctx.event.data
        ),
    },
    select=experiment.bucket(
        str(ctx.event.data["userId"]),
        weights={"short": 50, "detailed": 50},
    ),
)

async def track() -> None:
    await analytics.track(
        "experiment.variant_selected",
        {
            "experiment": "copy-style",
            "variant": res.variant,
            "userId": ctx.event.data["userId"],
        },
    )

await ctx.step.run("track-experiment", track)

return res.result
```

## Score the served variant

An experiment tells you which variant ran. It doesn't decide which one is best. Attach your own outcome signal with a score.

The returned `experimentRef` identifies the served variant. If the outcome is available during this run, pass the ref to `inngest.score.experiment()`:

```typescript {{ title: "TypeScript" }}
await inngest.score.experiment({
  name: "invoice-valid",
  value: validateInvoice(result),
  experiment: experimentRef,
});
```

```python {{ title: "Python" }}
# Requires the inngest release after 0.5.19.
await inngest_client.score_experiment(
    name="invoice-valid",
    value=validate_invoice(res.result),
    experiment=res.experiment_ref,
    run_id=ctx.run_id,
)
```

When another run scores a later outcome, save the original `experimentRef` and run ID with the result. Pass both to `inngest.score.experiment()` so the score attaches to the run and variant that produced it. You can also pass the ref to a deferred scorer through `defer()`. See [Experiments](/docs-markdown/agent-evals/experiments) for the full path from selection to scoring.

For score types and attribution rules, see [Scores](/docs-markdown/agent-evals/scores), [Deferred scoring](/docs-markdown/agent-evals/deferred-scoring), and the [scoring reference](/docs-markdown/reference/typescript/v4/functions/scoring).

## Best practices

- Give every experiment in a function a unique ID.
- Keep variant names stable. They appear in traces and analytics.
- Use `weighted()` when each run can get a fresh assignment.
- Use `bucket()` when a stable user or account experience matters.
- Use `custom()` when assignment must be controlled outside code or changed without a deploy.
- Keep selectors simple. Put the behavior you're comparing inside the variant callbacks.
- Prefer distinct experiment names across functions. Two functions can share a name, and the dashboard scopes detail views by function, but distinct names are easier to compare in lists.

## Troubleshooting

| Issue                                              | Solution                                                                                                                                                         |
| -------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Experiments aren't available.                      | Upgrade to TypeScript SDK v4.8.0 or later.                                                                                                                       |
| The same user sees different variants across runs. | `experiment.weighted()` is seeded by run ID, not user ID. Use `experiment.bucket(userId, { weights })` to keep a user on one variant.                            |
| The run fails with an unknown variant.             | Make sure the `experiment.fixed()` or `experiment.custom()` value exactly matches a key in `variants`. A typo fails the run.                                     |
| The run fails with a `NonRetriableError`.          | Make sure every variant callback calls at least one `step.*` tool.                                                                                               |
| `experiment.bucket()` gives surprising skew.       | Check that the bucket value isn't `null` or `undefined`. Missing values hash as an empty string, so they all land in the same bucket. The trace shows a warning. |
| A 50/50 split looks uneven.                        | Small samples are noisy. Send more events before you assume selection is wrong.                                                                                  |

For scoring problems, such as a missing `step.score()` or scores that don't attach to a variant, see [Agent Evals troubleshooting](/docs-markdown/agent-evals/troubleshooting).

## Related docs

- [`group.experiment()` reference](/docs-markdown/reference/typescript/v4/functions/group-experiment)
- [Experiments](/docs-markdown/agent-evals/experiments)
- [Compare models and prompts](/docs-markdown/agent-evals/guides/compare-models-and-prompts)
- [Roll out a workflow rewrite](/docs-markdown/agent-evals/guides/workflow-rewrite)
- [Versioning](/docs-markdown/durable-execution/guides-and-advanced/versioning)