Skip to main content
Not yet released. This guide describes implemented changes awaiting rollout. Availability requires the corresponding backend and dashboard release.

Goal

Distinguish configured resources, measured runtime limits, and unavailable observations.

Prerequisites

  • An environment with the workload signals migration and updated server and node agent.

Workflow

1
Inspect the configured memory and CPU allocation for the workload.
2
Check measured memory limits and CPU utilization relative to the assigned CPU quota.
3
Inspect the original sample timestamp and pressure signals before making a resource decision.
4
Follow sustained warning, critical, and recovery transitions in alerts.

Resource precedence and connected servers

An explicit workload memory or CPU allocation takes precedence over a named preset. Presets supply defaults; hosted plan allowances govern admission. Project create and update requests accept memory_mb and cpu_millicores. Memory must be positive; CPU must be at least 10 millicores. When a request includes both a preset and explicit allocations, those explicit allocations are checked against the hosted allowance after the preset defaults are applied. A connected customer server uses owner limits and the selected server capacity. Changing an account plan or updating a node agent does not apply a new workload allocation. This implementation is unreleased. Manual app memory and CPU updates use durable operations. Replicas, autoscaling, database purchase authorization, and dashboard controls remain later milestones.

Pending resource changes and host admission

On hosted servers, an agent reserves full memory limits and Kata memory overhead before creating, starting, or increasing a workload. Platform shared container hosts use bounded 2:1 CPU-reservation admission; per-container CPU ceilings and scheduling weights do not change. Kata CPU reservations remain exclusive. CPU contention can increase latency when workloads are busy together; shared CPU is not a dedicated-core guarantee. Admission shares one host lock across containerd and Kata. Pending claims remain durable across restarts. On customer-owned connected servers, container CPU limits share the host scheduler rather than consuming exclusive core reservations. Running containers with known memory usage count their observed usage; new starts, restarts, pending requests, and workloads with unknown usage reserve their full memory limits. Concurrent starts remain serialized. Per-container hard limits, storage reservations, host headroom, and Kata guarantees still apply. Simultaneous growth of running containers can exhaust RAM, so monitor memory pressure and OOM signals. Resource increases check only the dimensions that grow. An existing CPU reservation overage cannot block a memory-only increase, while a CPU increase still needs CPU capacity on hosted servers. Shared owner admission requires the updated API, worker, and enrolled node agent. The server verifies customer ownership for owner admission and platform ownership for shared CPU admission. Old agents retain strict reservations. The isolated capacity correction is deployed on the primary shared host and UBB; other nodes require their matching agent update. It introduces no environment variable or database migration. This limited rollout does not qualify every scaling or migration capability described here. The autoscaling verifier recognizes both supported cgroup-v2 CPU weight conversions. Legacy CPU fields are compared by their effective ceiling, and a committed scaling request may reconcile only its CPU reservation when all other resource limits match. The agent preserves the underlying verification error for diagnosis. This correction is deployed on UBB; other nodes need the matching agent release. Database disk reservations include the data-volume quota plus a 2 GiB writable-layer allowance. The agent reconciles a data-only reservation from authenticated volume ownership and the actual filesystem quota, checks capacity, verifies the live allocation, and durably seals the result. Retries do not add the allowance twice. This accounting repair does not expand the volume, restart the database, or change CPU and memory limits. The database accounting correction is implemented but unreleased and requires a node-agent update; it introduces no environment variable or SQL migration. Startup recovery defers unverified repairs without interrupting healthy database proxies. Scaling remains blocked when quota evidence or capacity is unavailable. Pending mutations and prepared replication require their existing recovery path. Database autoscaling still requires its own explicitly configured policy and bounds; application autoscaling does not enable it. The initial host budget keeps at least 1 GiB or 10% of RAM, whichever is larger, and 10% of CPU for host work. An unavailable or incomplete owned workload inventory blocks new admission. A failed live ClickHouse memory increase retains its durable intent and any already applied larger limit. Retry the same operation and target after measurements recover; a changed request cannot silently shrink an applied increase. Project update requests submit manual memory and CPU changes through a durable operation. resource_state on project detail reads reports the operation and the last verified applied allocation. Downsizes require apply_resource_downsize=true; pending changes block competing configuration, deployment, checkpoint restore, and cleanup mutations. The agent protocol binds each change to an immutable target and revision. Interrupted changes retain intent and already applied larger limits. Fresh actual cgroup memory, CPU quota, and CPU weight must match before acknowledgement. Historical replay cannot acknowledge another target or replace newer applied state. An isolated ARM64 Linux containerd test verified actual limits, safe retry after a metadata seal failure, restart preservation, crashed-process claim recovery, and concurrent memory admission. The targeted UBB agent rollout additionally verifies live application reservations; the broader production runtime matrix and Kata VM resizing remain unqualified. The server serializes admission and pending reservations per node before app preparation, manual resize, database creation and start, one-off jobs, pooler creation, and qualified stack service creation. Stack service bundles reserve atomically. Agent verification remains the final admission backstop. Capacity includes pending and surge workloads, stopped and retained storage, memory and CPU charges under the selected host policy, persistent volumes, writable layers, VM overhead, and host headroom. Filesystem availability covers agent state, content, and the configured snapshotter; devmapper also requires actual thin-pool data and metadata availability. Reservations have no expiry. Observed reservations remain accounted for when inventory snapshots arrive out of order, using exact physical or operation identity. Release requires exact container, task, and shim absence, or completed one-off cleanup. Retained volumes keep their storage promises until verified cleanup. Separate attempt identities prevent delayed old cleanup from releasing a newer lifecycle. Existing stack execution and WordPress import commands and package uploads use the shared one-off admission and exact cleanup receipt protocol, preserving service security, commands, files, and import authority. A readonly agent preflight identifies which services a stack apply will actually create. ImportPending skips web creation reservations; publication after verified import completion admits those services. A changed import phase requires a new preflight. One-off storage includes known persistent volumes and its writable layer. Unknown cleanup retains reservations. Unsupported live Kata resize, unknown cgroup mappings, unavailable snapshotter or thin-pool inventory, and incomplete resource allocations are blocked. Legacy persistent storage without exact ownership and unqualified block-volume cleanup remain conservatively charged. Updating an agent does not resize workloads. Beyond the explicitly described UBB correction, actual thin-pool hardware, Kata VM behavior, the broader production runtime matrix, upgrade packaging, and frontend controls remain unqualified. The limited rollout does not establish general availability of every scaling capability described here.

Moving a workload between servers

For new deployment-v2 moves between known nodes, the deployment plan saves the current managed DNS records before preparation. Preparation checks destination capacity and direct application readiness without moving public DNS away from the previous server. At cutover, the worker publishes the destination DNS record. The destination agent then requires repeated public probes bound to that candidate before promoting its route. A successful response from the previous server cannot satisfy this check. Rollback restores the saved DNS records before contacting the destination agent. Failed DNS restoration keeps rollback pending for retry; records changed outside the saved plan require operator reconciliation. DNS propagation can still interrupt traffic, so this is not a zero-downtime migration guarantee. This correction is implemented locally and unreleased. Update the destination node agent and deploy matching API and worker versions before starting a move. Drain old deployment workers during rollout; do not downgrade with new migration plans still active. Previously sealed plans are not retrofitted with DNS snapshots. No new environment variable or database migration is required for this correction. Automated protocol and provider tests cover preparation, cutover, restart recovery, and rollback. Live multi-host migration with production DNS has not yet been qualified.

Measured resource signals

  • The existing cpu_percent field measures CPU time relative to one core. signals.cpu_limit_percent reports utilization relative to the measured assigned quota.
  • signals.sampled_at preserves the original observation time through storage. Receiving an older sample does not make it fresh.
  • Repeated CPU reads inside the five-second sampling interval reuse the last complete sample with its original timestamp. A first sample or reset CPU counter remains unknown until a complete interval is available. This prevents readiness checks from erasing valid autoscaling measurements.
  • Memory signals include the actual hard limit, anonymous and file memory, inactive file cache, a separate working-set estimate, swap, OOM events, and pressure.
  • CPU throttling and memory events are cumulative counters. Compare consecutive valid samples to measure new events.
  • Missing signals are omitted. Historical samples without availability information remain unknown.

Sustained alert transitions

  • Memory warnings require usage at or above 80% for five minutes. CPU warnings require quota-normalized usage at or above 80% for three minutes.
  • Usage at or above 95% for one minute escalates a resource alert to critical. An existing critical alert remains critical through the warning band.
  • Recovery requires five minutes below 75%. Gaps longer than two minutes restart stabilization; stale, missing, future, or duplicate samples cannot resolve an alert.
  • Stabilization windows and active alerts persist across controller restarts. Concurrent observers serialize the transition and reuse the same alert identity.

Expected result

Explicit allocations survive deployment. Missing telemetry stays unknown, and stale observations cannot clear resource alerts.

Common failures

  • An older node agent does not report quota-normalized CPU or pressure signals. Update the agent before relying on those observations.
  • A missing runtime memory limit is unknown. A configured preset is not proof of the actual cgroup limit.