Goal
Give applications and operators a concrete S2 reliability runbook based on the public service contract and observable behavior.Prerequisites
- An S2-backed production workload
- An application recovery plan that identifies critical objects
Workflow
1
Monitor bucket bytes, object counts, request errors, bandwidth, and quota headroom.
2
Alert on sustained error rates, repeated throttling, and unexpected changes in object inventory.
3
Keep application retry behavior bounded and safe for idempotent object operations.
4
Use version restore for accidental overwrites and the StackShift recovery process for service-level incidents.
5
Verify critical keys, sizes, and checksums before declaring application recovery complete.
What to monitor
- Current bytes and object count compared with the bucket and plan quota.
- Request volume, error count, bandwidth, and top prefixes from bucket analytics.
- Recent access-log status codes and request identifiers when an application reports a failure.
- Lifecycle-rule errors, event-delivery retries, and inventory-export failures.
- Unexpected inventory differences for application-critical prefixes.
Application behavior during an incident
- Use exponential backoff with jitter for temporary
503 ServiceUnavailableand429 SlowDownresponses. - Honor
Retry-Afterwhen it is present and stop after a bounded number of attempts or an application deadline. - Retry reads freely when they have no side effects. Retry writes only when the key, body, checksum, and precondition make the operation idempotent.
- Do not automatically retry authentication, authorization, invalid-request, retention, or quota errors without correcting their cause.
- Surface a degraded state to callers instead of silently dropping uploads or inventing a successful object record.
Recovery validation
A bucket returning traffic again is only the start of recovery validation. Compare the expected inventory for critical prefixes, then verify object sizes and SHA-256 checksums for records the application cannot recreate. Keep object identifiers and expected checksums in the application database when business workflows must reconcile stored files. S2 object metadata is useful for storage operations, but it should not replace the application record that explains why an object exists. Service recovery must restore objects, retained versions, PostgreSQL and Garage metadata, policies, and required encryption dependencies together into an isolated destination. Record backup age, restoration duration, checksum results and access checks. A configured backup destination alone does not establish recoverability. The release candidate serializes cold-tier removal and restoration across processes and checks the stored revision before changing its class. Tiering compares primary and backup bytes before removal. Retained versions sharing the same stored blob follow its physical placement. Failed or uncertain primary deletion leaves cold metadata so the backup remains the recovery source. Cold restoration verifies the recorded size and SHA-256 before marking the object hot. Encrypted restoration also decrypts and authenticates the restored stream before committing that transition; customer-key objects require the original key. Corrupt copies are discarded and remain cold. Actual provider/KMS recovery qualification is recorded separately. Gateway reads bind response metadata and body to the same fetched object revision.Single-server capacity expansion
S2 keeps one storage server. The release-candidate capacity probe supports larger volumes and multiple declared data volumes on that host. It validates exact filesystem UUIDs, mount paths and data placement before reporting capacity. Deploy the helper and collector together using deployments/s2/CAPACITY.md. Assignable capacity excludes filesystem-reserved blocks, a recovery reserve of the larger of 8 GiB or 20 percent per writable volume, and a 1 GiB operational margin. Root retains a separate 10 GiB reserve. These physical measurements are distinct from customer billable usage and do not change prices. Missing mounts omit aggregate byte measurements. Inspect per-volume readiness and collector observation time instead of treating unavailable measurements as zero. The probe is startup and operational evidence; it is not a runtime reservation system for concurrent writes. Before expansion, verify an independent backup and an isolated restore. After expansion, reconcile filesystem size and Garage placement, then verify representative object reads, ranges and writes. Multi-volume hardware and recovery qualification remain release gates until executed by the operator.Recovery controls
- Restore a selected object version after an accidental overwrite or delete when versioning was enabled before the event.
- Use inventory exports to compare large keyspaces without loading every object body.
- Use retention rules for deletion protection and lifecycle rules for intentional aging; neither replaces recovery testing.
- Rotate any access key that may have been exposed during investigation and review access logs for its activity.
- Resume writes only after the application has passed its own read, write, list, and checksum checks.
Expected result
The application fails predictably during a storage incident and operators can prove that critical data is correct after recovery.
Common failures
Related guides
S2 object storage overview
Understand the S3-compatible S2 surface, its bucket model, security boundaries, data controls, and application workflows.
Versioning, lifecycle, and retention
Protect object history, restore earlier versions, automate aging policies, and prevent protected objects from being deleted too early.
Events, analytics, inventory, and websites
Connect object changes to application workflows, inspect bucket activity, export inventories, query catalogs, and publish static content.
Alerts view
Use alerts to focus on active operational problems instead of scanning every resource manually.