Skip to main content
Live. This area is documented as current, user-reliable behavior.

Goal

Understand when to move backend work into Durable Jobs and what StackShift guarantees after a job starts.

Prerequisites

  • A StackShift API key
  • Node.js or another backend runtime that can expose an HTTP handler

Workflow

1
Define a job handler in code.
2
Start the work once with an idempotency key.
3
Break longer work into steps so completed work is remembered.
4
Pause for external events when the flow needs user, webhook, or payment confirmation.
5
Use the run timeline, logs, attempts, and status to inspect what happened.

SDK 0.2.0 callback authentication

JavaScript and Go SDK 0.2.0 provide built-in callback verification. JavaScript requires Node.js 20 or newer. Signed callbacks also require the coordinated backend reliability rollout described below. The adapters authenticate raw-body Ed25519 signatures, timestamp freshness, and trusted account/run binding before dispatch, with a 4 MiB body limit. Configure the public key and account ID obtained from the authenticated verification-key endpoint; never trust callback-provided keys or URLs. Saved steps stop on failed or malformed state reads; only HTTP 404 means a missing step. Continuation checkpoints remain separate from events. Go preserves zero counter increments. Execution remains at least once and requires stable downstream idempotency keys. For JavaScript, use handleRequest for public routes. For Go, configure Options.Verification before mounting HTTPHandler; missing configuration or unsigned callbacks return HTTP 401. Low-level invocation methods remain for already authenticated input only.
JavaScript 0.2.0 handler
Go 0.2.0 verification configuration

Reliability update awaiting rollout

The reliability changes below are implemented but not deployed. They require migration 000524 and coordinated API and worker replacement. Existing deployments may still exhibit the earlier behavior. The prepared release guard requires coordinated replacement, refuses partial upgrades or pending legacy work, and verifies a database backup before migration; signing keys persist across deployments. Execution is at least once. Use a stable downstream idempotency key for each side effect, such as runId plus an operation name. A crash after an external effect but before its acknowledgment can repeat it, even when enqueue deduplication and saved steps are enabled.
  • Customer jobs use a public HTTPS handler URL, supplied on first registration or inherited from an existing customer definition. Accepted runs pin that URL and version; later submissions do not retarget queued work.
  • Platform job, queue, and event names are reserved. Use the owning platform feature to manage those runs. Customer callbacks have separate execution slots from platform jobs.
  • A matching event resumes one oldest active waiter. Use different correlation keys or separate events for multiple recipients. Successful waits preserve the retry budget and expose result.event separately from result.checkpoint.
  • Retry policy: explicit request override, then queue default, then three attempts. HTTP 408, 429 and 5xx are retryable with jitter and bounded Retry-After. Other non-2xx replies, redirects, malformed successful replies, and responses over 4 MiB fail permanently. Empty 2xx replies mean an empty result.
  • HTTP callbacks have a five-minute limit. Internal handlers have a thirty-minute execution deadline. Three expired-lease recoveries are allowed before a run fails for operator inspection.
  • Expired idempotency keys, state, and locks are removed in bounded batches. Events older than 90 days are removed only when no wait or unexpired idempotency key references them. Run history and referenced events are retained; no automatic terminal-run deletion is enabled.
  • Asset workflow creation, durable enqueue, linking, and inbound webhook delivery reservations commit atomically. Failed submissions do not consume the delivery ID.
  • List failed runs and inspect their logs before manually retrying. Retry and cancellation are recorded in the run log. Platform runs must be managed through their owning feature.

Verify callbacks before executing jobs

After the coordinated rollout, fetch GET /api/v1/jobs/verification-key using your StackShift API key. Configure its data.public_key and data.user_id on your handler. Do not fetch a key from a callback-supplied URL. Keep the private signing seed on StackShift servers only. Older Node SDK 0.1.1 does not authenticate callbacks. For that version, verify the raw body and expected account before calling handleInvocation using the wrapper below. SDK 0.2.0 includes handleRequest and verifyInvocation, so no separate wrapper is required. Deduplicate external side effects using stable business keys; signature verification alone does not prevent retries.
verify-invocation.mjs
Authenticated handler route

What Durable Jobs is

StackShift Durable Jobs is a reliable execution engine for backend work. It runs jobs outside the request path and keeps enough durable state to retry, resume, and finish safely. Use it for work that should not disappear when a request times out, a process restarts, or an external system is slow. Execution is at least once. Make external side effects idempotent.

Short example

Start a job from an API route, webhook, or backend service. StackShift stores the run and invokes your handler.

Mental model

  • Jobs execute work.
  • Steps break work into durable units.
  • State remembers progress across retries and resumes.
  • Events let a job pause without running compute, then resume when the right signal arrives.
  • Idempotency keys make duplicate starts safe.

Expected result

You can explain Durable Jobs as one product concept: jobs execute work, steps divide it, state remembers progress, and events resume paused runs.

Quick Start

Install the SDK, initialize a client, enqueue a job, add a worker handler, and inspect the run.

Idempotency

Use idempotency keys so duplicate requests do not create duplicate work.

Event Waiting + Correlation

Pause a workflow until the right external event arrives, then resume the correct run.