Goal
Understand when to move backend work into Durable Jobs and what StackShift guarantees after a job starts.Prerequisites
- A StackShift API key
- Node.js or another backend runtime that can expose an HTTP handler
Workflow
1
Define a job handler in code.
2
Start the work once with an idempotency key.
3
Break longer work into steps so completed work is remembered.
4
Pause for external events when the flow needs user, webhook, or payment confirmation.
5
Use the run timeline, logs, attempts, and status to inspect what happened.
SDK 0.2.0 callback authentication
JavaScript and Go SDK 0.2.0 provide built-in callback verification. JavaScript requires Node.js 20 or newer. Signed callbacks also require the coordinated backend reliability rollout described below. The adapters authenticate raw-body Ed25519 signatures, timestamp freshness, and trusted account/run binding before dispatch, with a 4 MiB body limit. Configure the public key and account ID obtained from the authenticated verification-key endpoint; never trust callback-provided keys or URLs. Saved steps stop on failed or malformed state reads; only HTTP 404 means a missing step. Continuation checkpoints remain separate from events. Go preserves zero counter increments. Execution remains at least once and requires stable downstream idempotency keys. For JavaScript, use handleRequest for public routes. For Go, configure Options.Verification before mounting HTTPHandler; missing configuration or unsigned callbacks return HTTP 401. Low-level invocation methods remain for already authenticated input only.JavaScript 0.2.0 handler
Go 0.2.0 verification configuration
Reliability update awaiting rollout
The reliability changes below are implemented but not deployed. They require migration 000524 and coordinated API and worker replacement. Existing deployments may still exhibit the earlier behavior. The prepared release guard requires coordinated replacement, refuses partial upgrades or pending legacy work, and verifies a database backup before migration; signing keys persist across deployments. Execution is at least once. Use a stable downstream idempotency key for each side effect, such as runId plus an operation name. A crash after an external effect but before its acknowledgment can repeat it, even when enqueue deduplication and saved steps are enabled.- Customer jobs use a public HTTPS handler URL, supplied on first registration or inherited from an existing customer definition. Accepted runs pin that URL and version; later submissions do not retarget queued work.
- Platform job, queue, and event names are reserved. Use the owning platform feature to manage those runs. Customer callbacks have separate execution slots from platform jobs.
- A matching event resumes one oldest active waiter. Use different correlation keys or separate events for multiple recipients. Successful waits preserve the retry budget and expose result.event separately from result.checkpoint.
- Retry policy: explicit request override, then queue default, then three attempts. HTTP 408, 429 and 5xx are retryable with jitter and bounded Retry-After. Other non-2xx replies, redirects, malformed successful replies, and responses over 4 MiB fail permanently. Empty 2xx replies mean an empty result.
- HTTP callbacks have a five-minute limit. Internal handlers have a thirty-minute execution deadline. Three expired-lease recoveries are allowed before a run fails for operator inspection.
- Expired idempotency keys, state, and locks are removed in bounded batches. Events older than 90 days are removed only when no wait or unexpired idempotency key references them. Run history and referenced events are retained; no automatic terminal-run deletion is enabled.
- Asset workflow creation, durable enqueue, linking, and inbound webhook delivery reservations commit atomically. Failed submissions do not consume the delivery ID.
- List failed runs and inspect their logs before manually retrying. Retry and cancellation are recorded in the run log. Platform runs must be managed through their owning feature.
Verify callbacks before executing jobs
After the coordinated rollout, fetch GET /api/v1/jobs/verification-key using your StackShift API key. Configure its data.public_key and data.user_id on your handler. Do not fetch a key from a callback-supplied URL. Keep the private signing seed on StackShift servers only. Older Node SDK 0.1.1 does not authenticate callbacks. For that version, verify the raw body and expected account before calling handleInvocation using the wrapper below. SDK 0.2.0 includes handleRequest and verifyInvocation, so no separate wrapper is required. Deduplicate external side effects using stable business keys; signature verification alone does not prevent retries.verify-invocation.mjs
Authenticated handler route
What Durable Jobs is
StackShift Durable Jobs is a reliable execution engine for backend work. It runs jobs outside the request path and keeps enough durable state to retry, resume, and finish safely. Use it for work that should not disappear when a request times out, a process restarts, or an external system is slow. Execution is at least once. Make external side effects idempotent.Short example
Start a job from an API route, webhook, or backend service. StackShift stores the run and invokes your handler.Mental model
- Jobs execute work.
- Steps break work into durable units.
- State remembers progress across retries and resumes.
- Events let a job pause without running compute, then resume when the right signal arrives.
- Idempotency keys make duplicate starts safe.
Expected result
You can explain Durable Jobs as one product concept: jobs execute work, steps divide it, state remembers progress, and events resume paused runs.
Related guides
Quick Start
Install the SDK, initialize a client, enqueue a job, add a worker handler, and inspect the run.
Idempotency
Use idempotency keys so duplicate requests do not create duplicate work.
Event Waiting + Correlation
Pause a workflow until the right external event arrives, then resume the correct run.