Not yet released. This guide describes implemented changes awaiting rollout. Availability requires the corresponding backend and dashboard release.
Goal
Use the new backend workflows with current authority, bounded execution and explicit quality and coverage limits.Prerequisites
- Unreleased migrations 000549–000553 and matching API and existing Assets-role worker versions. Quiesce Support consumers and delivery during migration and update every replica before resuming.
- Existing Support object storage, scanner, managed provider or encrypted BYOK configuration and workspace budgets. The worker image must include ffmpeg, ffprobe and the Linux media limiter.
- Current workspace membership, session, SSO/IP policy and the appropriate inbox, author, publisher or configuration permission. Defaults remain off; no new environment variable or Assets subscription is required.
Workflow
1
Create a staff copilot or article translation job with the existing source revision and Idempotency-Key. Prefer: respond-async returns a 202 queued job; ordinary callers await the same durable execution. Poll /ai/jobs/{jobID}. Disconnect stops waiting while the job continues.
2
For approved staff answer previews, set stream: true and consume /ai/jobs/{jobID}/events using the staff bearer token. Use its opaque event ID as Last-Event-ID. Remove provisional text on withdrawal, cancellation, interruption, permission denial, expiry or resync; reconcile customer completion with its committed message_id.
3
Inspect /ai/usage with environment_id and RFC3339 from/to. The added costs distinguish managed/BYOK, immutable dispatch-time prices, measured cache/audio/retrieval units and explicitly unknown or unpriced history. Reporting never debits credits or changes allowances.
4
Run go run ./cmd/support-evaluate -dataset internal/support/evaluation/testdata/multilingual-v1.json in the backend for offline synthetic checks. Owner/admin configuration authority can create /ai/evaluation-datasets and /ai/evaluation-runs, then inspect, cancel and submit immutable human ratings.
5
Create scoped /knowledge/glossaries as an author and approve the current revision as a publisher. Preview up to 100 article/target pairs through /knowledge/translation-batches/preview, then submit its source/target/glossary pins with Idempotency-Key. Inspect status, cancel or explicitly retry a failed item.
6
Audio/video upload completion queues the default validated derivative. Select an optional fixed profile through /attachments/{attachmentID}/conversions; inspect /media-conversions/{conversionID} and use current If-Match for cancellation or retry. Final provider delivery waits for a compatible scanned derivative.
7
A client starts or explicitly switches with PUT /workforce/focus using environment_id/conversation_id, heartbeats with POST /workforce/focus/heartbeat every 30 seconds with lease_id/If-Match and stops explicitly on exit. The server grants a 90-second lease and allows one active focus per workspace membership across sessions and environments.
Answer replay and customer opt-in
A dedicated bounded replay stream exposes approved answer text after hidden planning/tool work. The renderer has no tools or private planning context and can emit only exact approved prefixes. Tool arguments/results, reasoning and structured planning output are never replayed. Streamed summaries use public messages only and exclude private notes. Current session, SSO/IP, permissions, source/publication scope, generation and retention are rechecked. Customer provisional_streaming_enabled is a separate revision-protected setting, default false. /support/widget/v1/conversations/{conversationID}/ai/events?job_id={jobID} requires a current first-party website access token and the originating installation/brand/contact/session. Native non-website origins are denied; external providers receive only final committed messages. This backend release adds no frontend stream consumer. Events are started, text_delta, completed, withdrawn, cancelled and interrupted with job/generation/sequence and optional attempt_id. Customer completion binds the committed message_id. Replay lasts 30 minutes, at most 256 records and 32 KiB per generation, 4 KiB per delta and 64 records per read; connections end after five minutes with 15-second comments. Invalidation scrubs provisional bytes, and erasure/content expiry removes replay.Reporting and quality evidence
The existing token report gains costs and monetary_cost_available=true for reporting, with charging_policy_changed=false. Shared platform prices are captured before each provider attempt, including renderer/fallback, transcription, summaries and retrieval. Actual supplier units are priced only when a matching snapshot exists. Unknown, unpriced and historical rows are partial; older retrieval cannot be reconstructed. Financial evidence contains no prompt/answer bodies and survives content deletion/rollback. The initial dataset contains 100 distinct Support cases each in Nigerian English (en-NG), Nigerian Pidgin (pcm) and French (fr), plus 100 passages in each English/French and English/Pidgin direction and 12 audio critical-fact text fixtures. Generated annotations remain pending_human_review. Offline fixture playback is separate from live/human scores and is not recorded-speech evaluation. Deliberately naïve evidence playback produces 15 credential-leak findings in the untrusted-instruction cases, five per language; this is a negative scorer baseline, not a quality pass. Evaluation mode defaults deterministic and makes no provider calls. Live mode requires explicit live_opt_in, enabled copilot/provider configuration, request_budget through 3,000 and token_budget through 10,000,000 within existing workspace limits. Each supplier fallback counts against the request bound. No business tools, messages or publication are allowed. Baselines must share dataset and mode; ratings bind reviewers and cannot be overwritten. Run history lasts one day unless explicit history opt-in chooses configured retention. Release targets remain at least 95% grounding/fact preservation, correct handover, translation meaning and speech critical facts; 90% human language acceptance; zero critical violations; reviewed annotations and complete human coverage. The initial language/direction suites require 100 cases each. Release assessment never qualifies offline playback, and recorded-speech/provider qualification remains separately outstanding. Use synthetic datasets; immutable dataset revisions are retained independently of run expiry.Translation, conversion and focus limits
Glossaries are environment/centre/source-target scoped with at most 200 required or exact terms. Approval is publisher-only and immutable. Batches pin source article/draft, target and glossary revisions, run at most two items per workspace under lower existing AI limits, return unpublished drafts and recheck current author/team/session access. Content expires under AI retention; batch metadata lasts 30 days. Existing knowledge review/save/publish remains required. The existing Assets-role worker shares one conversion slot through the internal Support media queue. Inputs remain five attachments/message, 10 MiB ordinary/video and 20 MiB audio, with stricter provider ceilings. WAV PCM, MP3, AAC/MP4 and Ogg/WebM Opus audio plus MP4/MOV/WebM video produce fixed AAC MP4, Opus Ogg or H264/AAC MP4 derivatives. Input/output are limited to five minutes; input dimensions 4096×2160, output 720p/30 fps. Linux conversion enforces a 120-second process deadline, 1 GiB address space, bounded output/logs/file descriptors, scratch below 256 MiB, no shell/remote fetching and restricted demuxers/local protocols. Input and output scans/checksums/quota and current authority gate publication. Derivatives retain provenance and expire within 30 days or earlier audio retention; erasure/source deletion denies reads and deletes bytes asynchronously. Image builds check codec versions and limiter presence; the deployment host checks both codecs through the unchanged limiter before stopping writers or migrating. This avoids Rosetta address-space failures during Apple Silicon builds. The native check selects the worker image only, excluding Compose dependencies. If images were already published before a preflight failure, operators may retry with ./deploy.sh platform-runtime —reuse-published followed by the explicit 40-character release tag. This reuses published runtime images without local rebuilding or upload; configuration and migration/readiness checks still run. Full conversion and process-limit qualification remains separate. Focus leases close on expiry, explicit switch/stop, logout/revocation, current access loss, conversation closure and erasure. Metric active_handling_time version 2 sums recorded interval intersections by timezone and captured dimensions. Multiple staff contributions are effort and may exceed conversation elapsed time. Unobserved, expired and historical coverage stays unknown; history_complete=false. Existing SLA clocks remain unchanged. No focus-emitting UI is added in this backend release.Expected result
Stable jobs, unpublished translation drafts and validated media results can be reviewed independently of HTTP connections. Cost totals and recorded handling effort retain explicit unknown coverage; deterministic checks do not claim provider or language qualification.