Scenario 2 of 3
Crash. Retry. Same sandbox.
Completed steps keep their results. When a step fails, the retry reuses the same sandbox and filesystem. If the run fails permanently, onFailure cleans up.
| Status | Run ID | Trigger | App | Queued at | Started at | Ended at | Duration |
|---|---|---|---|---|---|---|---|
| queued | - | playground/durability.run | sandbox-playground | - | - | - | - |
- Run ID
- -
- App
- sandbox-playground
- Function
- durability-demo
- Duration
- -
- Queued at
- -
- Started at
- -
- Ended at
- -
Trace
Press Run to watch the sample trace unfold.
- ○create-sandboxCreate the sandbox
- ○wait-runningWait until RUNNING
- ○write-markerWrite /tmp/marker
- ○flaky-after-createInjected failure
- ○run-helloRun: echo hello; node --version
- ○read-markerRead /tmp/marker
- ○destroy-sandboxDestroy the sandbox
Run an example to watch recovery.
Command outputstdout / stderr
Run an example to see its output.The real function this demo illustrates · 42 lines
export const buildAndTest = inngest.createFunction(
{
id: "durability-demo",
retries: 2,
triggers: [{ event: "playground/durability.run" }],
// Runs only after the last retry fails. Finds the sandbox by its
// deterministic name and destroys it, so nothing leaks.
onFailure: async ({ event, step }) => {
const { visitorId, sessionId } = event.data.event.data;
const page = await step.sandbox.list("list-sandboxes", { limit: 100 });
for (const sb of page.items) {
if (sb.name === sandboxName(visitorId, sessionId)) await sb.destroy(`destroy-${sb.id}`);
}
},
},
async ({ event, step, attempt }) => {
// Memoized: on a retry this returns the SAME sandbox, no second VM.
const created = await step.sandbox.create("create-sandbox", {
name: sandboxName(event.data.visitorId, event.data.sessionId),
vcpu: 2,
memoryMb: 2048,
runningTimeout: false,
});
const sandbox = await created.waitUntilRunning("wait-running", { timeout: "60s" });
await sandbox.commands.run("write-marker", `echo "written on attempt ${attempt + 1}" > /tmp/marker`);
// A step that fails once. Inngest retries the whole function; every step
// above replays from its stored result, so the sandbox is not recreated.
await step.run("flaky-after-create", () => {
if (attempt === 0) throw new Error("Injected failure after creating the sandbox");
});
const hello = await sandbox.commands.run("run-hello", "echo hello from $(hostname); node --version");
// The marker from attempt 1 is still here: same VM, same filesystem.
const marker = await sandbox.commands.run("read-marker", "cat /tmp/marker");
await sandbox.destroy("destroy-sandbox");
return { sandboxId: sandbox.id, attempts: attempt + 1, hello: hello.stdout, marker: marker.stdout };
},
);Choose a retryable crash to compare saved steps with a retry loop that starts over.
The same job without durable steps: every call, retry, poll and cleanup written by hand · 43 lines
async function buildAndTest(job: Job, attempt = 1): Promise<Result> {
let sandboxId: string | undefined;
try {
// 1. Create. If this times out, was a VM created? Unknown.
const res = await fetchWithTimeout(`${API}/v2/sandboxes`, { method: "POST", body: JSON.stringify(job.spec) });
sandboxId = (await res.json()).data.id;
// 2. Poll until RUNNING, by hand.
for (let i = 0; i < 60; i++) {
const s = (await (await fetch(`${API}/v2/sandboxes/${sandboxId}`)).json()).data;
if (s.status === "RUNNING") break;
if (s.status === "FAILED") throw new Error("sandbox failed to start");
await sleep(1000);
}
// 3. Each call gets its own retry loop and error classification.
await retry(3, () => exec(sandboxId!, "echo written > /tmp/marker"), { retryIf: isTransient });
await postProcess(); // crashes here on attempt 1
const hello = await retry(3, () => exec(sandboxId!, "echo hello; node --version"), { retryIf: isTransient });
const marker = await retry(3, () => exec(sandboxId!, "cat /tmp/marker"), { retryIf: isTransient });
return { hello, marker };
} catch (err) {
if (err instanceof PermanentError || attempt >= 3) throw err;
await sleep(backoff(attempt));
// Naive retry: starts from the top and creates a SECOND sandbox.
// The first is still RUNNING (and billing) unless the finally below ran,
// and if the process itself died (OOM, deploy, SIGKILL) it did not.
return buildAndTest(job, attempt + 1);
} finally {
if (sandboxId) {
try { await fetch(`${API}/v2/sandboxes/${sandboxId}`, { method: "DELETE" }); }
catch { /* leaked; hope the nightly cleanup cron catches it */ }
}
}
}
// Elsewhere: a cron that lists sandboxes older than 1h and deletes them.
// Elsewhere: tracing for each call, by hand.
// Not handled: a crash between create and finally (sandbox leaked).
// Not handled: "request timed out but the command ran" (ambiguous ops).
// Not handled: which errors are safe to repeat and which are not.