Skip to main content
Not yet released. This guide describes implemented changes awaiting rollout. Availability requires the corresponding backend and dashboard release.

Goal

Request copies and inspect verified applied allocations through a durable operation.

Prerequisites

  • The workload scaling migrations, updated server and worker, and a ready agent advertising manual_application_copies_v1.
  • A stateless service without a local persistent volume for horizontal scaling.

Workflow

1
As the project owner or team owner/admin, open project settings and inspect requested copies, applied copies, observations, and activity. Team members can inspect scaling but cannot change it.
2
Choose copy and resource bounds. Confirm singleton schedules and, for workers, concurrent consumption and acknowledgement/retry safety.
3
Apply scaling and follow the operation until each exact copy is ready. A request acceptance is not completion.
4
Enable automation only after reviewing its limits. Turn it off to return to manual control.

Opt-in sleep and request-triggered wake

Sleep when idle is implemented locally and unreleased. It requires matching API, worker, node agent and standalone gateway versions. Enable it only for a stateless web process with no local persistent volume, in-process jobs, schedules or background work. Databases and workers stay running. One awake copy is required; automatic vertical/horizontal scaling must be off. In project settings, enable Sleep when idle, confirm statelessness and no background work, and apply the policy. policy.idle accepts enabled, idle_seconds (60–86400), wake_seconds (1–120), max_requests (1–1024), and no_background_work. The CLI JSON request and governed project.scaling.request tool accept the same policy. Omission leaves sleep disabled. Public HTTP and authorized private HTTP, gRPC and raw TCP traffic use the gateway. The gateway holds the request or connection while the exact retained workload starts and passes its configured readiness probe. Use TCP readiness for a non-HTTP service. Requests have a bounded wake wait and queue limit. HTTP wake failure returns 503 with Retry-After; private TCP failure closes the connection before forwarding bytes. The gateway does not replay a forwarded write. Open HTTP requests and private TCP connections prevent sleep. Long-lived gRPC streams or pooled connections may keep a service awake. Existing direct private connections must drain before the first sleep. The private gateway keeps the existing application network authorization and checks access again after wake. Sleeping production applications retain their admission reservation and remain compute-billable. Previews release CPU/RAM capacity and must reacquire it on wake; insufficient capacity leaves wake pending until its timeout. Hosted preview compute uses measured awake intervals at the existing rate, with fractional-cent carry and idempotent hourly settlement. The initial partial hour uses the observed runtime start time, and prior hourly usage is credited without changing measured duration. BYOS application deployments remain free. Server rental, databases and retained storage are unchanged. Disabling sleep wakes the service and restores its direct routes. Redeploy preserves the approved policy on the new release. Retirement stops the exact old workload, revokes wake access and finalizes its meter. Explicit Stop is not undone by incoming traffic. Gateway registrations and deployment journals survive restarts; unavailable agent control or failed readiness causes a bounded failure rather than routing to an unverified workload. Routine platform health polling does not wake a deliberately sleeping service. Explicit HTTP monitors and customer probes are traffic and may prevent sleep. Process health while sleep is enabled does not establish application-level HTTP health. Native containerd/Kata deployment qualification and rollout remain separate from local protocol and lifecycle tests.

Autonomous and external agents

Objective actions project.scaling.request, project.scaling.resize, and database.scaling.request use committed approval receipts and the owning scaling controller. A resize binds current CPU, memory and policy revision, exact requested resources, disruption acknowledgement, and an idempotency key. Deployments, manual changes, repairs and automation share resource admission. External MCP connections require explicit grants for these new actions. Existing resource or deployment grants do not acquire scaling authority. Dispatch rechecks target ownership, approval revision, current grant and owner policy. External controller reconciliation does not spend or require an AI inference budget; infrastructure limits and database cost acknowledgements remain separate. Activity retains the agent action and controller operation. Accepted requests remain pending until the exact operation and applied resources or healthy copies verify. Replica completion snapshots are checked against the exact controller completion time; live metrics freshness is reported separately. Saving policy before a running release retains intent and does not establish runtime success. Cancellation revokes further agent dispatch and waits for an already admitted operation to settle; it does not reverse capacity or silently disable automation. Agent resource context includes desired and applied resources, copy readiness, policy bounds, blocked reasons and stored metrics freshness. Writer or deployment samples are labelled with their scope; missing samples are unavailable, never zero. A failed build or rolled-back candidate reporting OOM no longer doubles saved workload memory and starts another release. Build failures require build-worker diagnosis. Runtime failures require fresh evidence for the current runtime and an approved controller resize within the launch resource caps. Unknown controller outcomes require reconciliation before any replacement action.

Placement and lifecycle

Hosted Free allows one copy, Ignite two, Pro four, and Enterprise eight per service. Connected workloads use their owner’s selected server and measured available capacity. Copy limits do not use the number of projects. Release commands run once. Platform scheduled commands target the primary workload. Additional copies receive STACKSHIFT_REPLICA_INDEX and STACKSHIFT_SCHEDULES_ENABLED=false; application schedulers must honor the declared singleton contract. Redeploy prepares the requested copies from one immutable image before completion, publishes the new route pool, and drains the prior release. Stop and restart operate on the recorded physical set. Updating an agent preserves running allocations.

Automatic decisions

  • Automation requires a fresh readiness check of every exact applied copy against its committed release policy and resources. A running process alone does not prove HTTP readiness. Agents without the current observation endpoint block automation until updated. Readiness qualifies across polls: copies are probed concurrently, and counted successes remain at least the sealed period apart. Pending qualification blocks automatic decisions. Individual HTTP/TCP probes retain their configured timeout, up to 900 seconds.
  • Each successful observation needs readiness evidence at most 90 seconds old for every copy. Failed or cancelled rounds, changed sealed journals, lost monitoring continuity and agent restarts reset qualification. The node retains progress for at most 256 release identities; idle entries expire after one hour. Exhausted observation capacity blocks automation until progress expires.
  • Application and database automation share a separate worker pool with at most four observations running and four resource identities queued. One resource has at most one queued or running task, and execution reloads its current policy. Slow probes leave resource operation reconciliation available. When the pool is busy, later policies wait for capacity; delayed or stale samples reset their sustained windows instead of authorizing a change. Worker shutdown cancels and joins observation tasks.
  • Memory above 85% for five minutes has priority and requests bounded vertical growth.
  • With horizontal scaling off, quota-normalized CPU above 80% for three minutes requests vertical growth.
  • Horizontal scale-out requires CPU above 70% for three minutes; scale-in requires CPU below 30% for twenty minutes, safe projected utilization on remaining copies, and healthy memory.
  • Application cooldown is five minutes. At most two automatic vertical increases are allowed per hour. Active workload mutations, invalid samples, owner bounds, hosted limits, and admission suppress competing decisions.

API, CLI, manifests, and agent tools

GET /api/v1/projects/{id}/scaling returns policy, desired/applied counts, exact replicas, blocked reason, and operation status. PUT requires expected_revision, a unique idempotency_key, copies, policy, and safety. Use stackshift project scaling PROJECT and stackshift project scale PROJECT —file scaling.json. The CLI uses the same permission and intent checks as the dashboard. Application manifests accept memoryMB, cpuMillicores, and scaling with copies, enabled, vertical, horizontal, minCopies, maxCopies, maxMemoryMB, maxCpuMillicores, stateless, singletonSchedules, concurrentWorkers, and acknowledgedJobs. New services retain their policy for their first immutable release. Existing resource growth must complete before a combined policy update is reapplied. Governed agent inspection exposes current state. Applying scaling requires approval for the concrete revision and request and calls the owning service.

Expected result

Copies use the committed immutable image and allocation. New copies enter routing after readiness; removed copies leave routing before bounded drain and exact cleanup.

Common failures

  • An older or unready agent blocks scaling with an update or qualification reason.
  • Missing, stale, unhealthy, or incomplete observations suppress automation.
  • Insufficient measured host capacity retains the intent for reconciliation.
  • Cron and scheduler processes remain singleton; local persistent storage blocks horizontal copies.