> ## Documentation Index
> Fetch the complete documentation index at: https://docs.stackshift.cloud/llms.txt
> Use this file to discover all available pages before exploring further.

# Support AI, media and handling reporting

> Durable answer jobs, cost reporting, evaluation, approved translation batches, controlled conversion and explicit staff focus.

<Note>
  **Not yet released.** This guide describes implemented changes awaiting rollout. Availability requires the corresponding backend and dashboard release.
</Note>

## Goal

Use the new backend workflows with current authority, bounded execution and explicit quality and coverage limits.

## Prerequisites

* Unreleased migrations 000549–000553 and matching API and existing Assets-role worker versions. Quiesce Support consumers and delivery during migration and update every replica before resuming.
* Existing Support object storage, scanner, managed provider or encrypted BYOK configuration and workspace budgets. The worker image must include ffmpeg, ffprobe and the Linux media limiter.
* Current workspace membership, session, SSO/IP policy and the appropriate inbox, author, publisher or configuration permission. Defaults remain off; no new environment variable or Assets subscription is required.

## Workflow

<Steps>
  <Step>
    Create a staff copilot or article translation job with the existing source revision and Idempotency-Key. Prefer: respond-async returns a 202 queued job; ordinary callers await the same durable execution. Poll /ai/jobs/\{jobID}. Disconnect stops waiting while the job continues.
  </Step>

  <Step>
    For approved staff answer previews, set stream: true and consume /ai/jobs/\{jobID}/events using the staff bearer token. Use its opaque event ID as Last-Event-ID. Remove provisional text on withdrawal, cancellation, interruption, permission denial, expiry or resync; reconcile customer completion with its committed message\_id.
  </Step>

  <Step>
    Inspect /ai/usage with environment\_id and RFC3339 from/to. The added costs distinguish managed/BYOK, immutable dispatch-time prices, measured cache/audio/retrieval units and explicitly unknown or unpriced history. Reporting never debits credits or changes allowances.
  </Step>

  <Step>
    Run go run ./cmd/support-evaluate -dataset internal/support/evaluation/testdata/multilingual-v1.json in the backend for offline synthetic checks. Owner/admin configuration authority can create /ai/evaluation-datasets and /ai/evaluation-runs, then inspect, cancel and submit immutable human ratings.
  </Step>

  <Step>
    Create scoped /knowledge/glossaries as an author and approve the current revision as a publisher. Preview up to 100 article/target pairs through /knowledge/translation-batches/preview, then submit its source/target/glossary pins with Idempotency-Key. Inspect status, cancel or explicitly retry a failed item.
  </Step>

  <Step>
    Audio/video upload completion queues the default validated derivative. Select an optional fixed profile through /attachments/\{attachmentID}/conversions; inspect /media-conversions/\{conversionID} and use current If-Match for cancellation or retry. Final provider delivery waits for a compatible scanned derivative.
  </Step>

  <Step>
    A client starts or explicitly switches with PUT /workforce/focus using environment\_id/conversation\_id, heartbeats with POST /workforce/focus/heartbeat every 30 seconds with lease\_id/If-Match and stops explicitly on exit. The server grants a 90-second lease and allows one active focus per workspace membership across sessions and environments.
  </Step>
</Steps>

## Answer replay and customer opt-in

A dedicated bounded replay stream exposes approved answer text after hidden planning/tool work. The renderer has no tools or private planning context and can emit only exact approved prefixes. Tool arguments/results, reasoning and structured planning output are never replayed. Streamed summaries use public messages only and exclude private notes. Current session, SSO/IP, permissions, source/publication scope, generation and retention are rechecked.

Customer provisional\_streaming\_enabled is a separate revision-protected setting, default false. /support/widget/v1/conversations/\{conversationID}/ai/events?job\_id=\{jobID} requires a current first-party website access token and the originating installation/brand/contact/session. Native non-website origins are denied; external providers receive only final committed messages. This backend release adds no frontend stream consumer.

Events are started, text\_delta, completed, withdrawn, cancelled and interrupted with job/generation/sequence and optional attempt\_id. Customer completion binds the committed message\_id. Replay lasts 30 minutes, at most 256 records and 32 KiB per generation, 4 KiB per delta and 64 records per read; connections end after five minutes with 15-second comments. Invalidation scrubs provisional bytes, and erasure/content expiry removes replay.

## Reporting and quality evidence

The existing token report gains costs and monetary\_cost\_available=true for reporting, with charging\_policy\_changed=false. Shared platform prices are captured before each provider attempt, including renderer/fallback, transcription, summaries and retrieval. Actual supplier units are priced only when a matching snapshot exists. Unknown, unpriced and historical rows are partial; older retrieval cannot be reconstructed. Financial evidence contains no prompt/answer bodies and survives content deletion/rollback.

The initial dataset contains 100 distinct Support cases each in Nigerian English (en-NG), Nigerian Pidgin (pcm) and French (fr), plus 100 passages in each English/French and English/Pidgin direction and 12 audio critical-fact text fixtures. Generated annotations remain pending\_human\_review. Offline fixture playback is separate from live/human scores and is not recorded-speech evaluation. Deliberately naïve evidence playback produces 15 credential-leak findings in the untrusted-instruction cases, five per language; this is a negative scorer baseline, not a quality pass.

Evaluation mode defaults deterministic and makes no provider calls. Live mode requires explicit live\_opt\_in, enabled copilot/provider configuration, request\_budget through 3,000 and token\_budget through 10,000,000 within existing workspace limits. Each supplier fallback counts against the request bound. No business tools, messages or publication are allowed. Baselines must share dataset and mode; ratings bind reviewers and cannot be overwritten. Run history lasts one day unless explicit history opt-in chooses configured retention.

Release targets remain at least 95% grounding/fact preservation, correct handover, translation meaning and speech critical facts; 90% human language acceptance; zero critical violations; reviewed annotations and complete human coverage. The initial language/direction suites require 100 cases each. Release assessment never qualifies offline playback, and recorded-speech/provider qualification remains separately outstanding. Use synthetic datasets; immutable dataset revisions are retained independently of run expiry.

## Translation, conversion and focus limits

Glossaries are environment/centre/source-target scoped with at most 200 required or exact terms. Approval is publisher-only and immutable. Batches pin source article/draft, target and glossary revisions, run at most two items per workspace under lower existing AI limits, return unpublished drafts and recheck current author/team/session access. Content expires under AI retention; batch metadata lasts 30 days. Existing knowledge review/save/publish remains required.

The existing Assets-role worker shares one conversion slot through the internal Support media queue. Inputs remain five attachments/message, 10 MiB ordinary/video and 20 MiB audio, with stricter provider ceilings. WAV PCM, MP3, AAC/MP4 and Ogg/WebM Opus audio plus MP4/MOV/WebM video produce fixed AAC MP4, Opus Ogg or H264/AAC MP4 derivatives. Input/output are limited to five minutes; input dimensions 4096×2160, output 720p/30 fps.

Linux conversion enforces a 120-second process deadline, 1 GiB address space, bounded output/logs/file descriptors, scratch below 256 MiB, no shell/remote fetching and restricted demuxers/local protocols. Input and output scans/checksums/quota and current authority gate publication. Derivatives retain provenance and expire within 30 days or earlier audio retention; erasure/source deletion denies reads and deletes bytes asynchronously. Image builds check codec versions and limiter presence; the deployment host checks both codecs through the unchanged limiter before stopping writers or migrating. This avoids Rosetta address-space failures during Apple Silicon builds. The native check selects the worker image only, excluding Compose dependencies. If images were already published before a preflight failure, operators may retry with ./deploy.sh platform-runtime --reuse-published followed by the explicit 40-character release tag. This reuses published runtime images without local rebuilding or upload; configuration and migration/readiness checks still run. Full conversion and process-limit qualification remains separate.

Focus leases close on expiry, explicit switch/stop, logout/revocation, current access loss, conversation closure and erasure. Metric active\_handling\_time version 2 sums recorded interval intersections by timezone and captured dimensions. Multiple staff contributions are effort and may exceed conversation elapsed time. Unobserved, expired and historical coverage stays unknown; history\_complete=false. Existing SLA clocks remain unchanged. No focus-emitting UI is added in this backend release.

## Expected result

<Check>
  Stable jobs, unpublished translation drafts and validated media results can be reviewed independently of HTTP connections. Cost totals and recorded handling effort retain explicit unknown coverage; deterministic checks do not claim provider or language qualification.
</Check>

## Common failures

<Warning>
  * execution\_interrupted means the prior dispatched inference or conversion is not automatically repeated. Confirm current source and authority before explicit retry. HTTP reconnect uses the original job ID and idempotency intent.
  * Interrupted evaluation recovery closes abandoned child jobs and reconciles reserved usage. Translation items requested by disabled staff fail with authority\_changed, allowing later work to progress; restore authority before an explicit retry.
  * source\_or\_target\_changed or stale\_revision invalidates a translation result. Preview current source/target/glossary revisions again. Approved glossaries cannot be edited; create and approve another revision set.
  * A failed or cancelled media conversion fails pending outbound delivery. Correct and explicitly retry conversion, then use the existing delivery retry. Stricter provider limits still apply, including Microsoft inline attachments below 3 MiB.
  * Missing Linux tools, unsupported codecs, invalid frame rates, durations below one millisecond, over-limit duration/dimensions, scan failure or changed authority deny conversion/publication. macOS queue/unit checks do not qualify the production worker image.
  * A stale heartbeat cannot reclaim focus after another session explicitly switches. Starting a new lease closes an expired interval at its lease boundary without waiting for maintenance; missing telemetry stays unknown.
</Warning>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.