Skip to main content
This guide includes features awaiting release. If an option is not available in your workspace, contact StackShift Support.

What you can do

Compare AI behavior with examples and use outcomes to choose improvements. Where to find it: AI → Evaluations, Improvements and Procedure outcomes

Before you start

  • Permission to manage or review AI.
  • Representative examples with the answer or behavior your team expects.

Steps

1
Create or open an evaluation and add realistic examples, including questions AI should hand to a person.
2
Select the relevant configuration or procedure version and run the evaluation.
3
Read individual results and their evidence. A combined score can hide one serious mistake.
4
Use Improvements and Procedure outcomes to investigate repeated failure patterns; revise the knowledge, policy or procedure that caused them.
5
Run the same examples again before broadening customer-reply scope. Review real outcomes afterward.

Example

Include both a routine delivery question and a request for an unauthorized refund; expect an answer to one and escalation for the other.

Things to know

  • Evaluation results do not guarantee every future answer.
  • An improvement suggestion needs review before changing a published procedure or customer policy.

What happens next

You have a repeatable way to assess a change instead of assuming a new prompt is better.

If something goes wrong

  • Run cannot start: check provider availability, credits and inputs.
  • Good score but poor replies: include examples from the actual failing conversation and verify the evaluated version.

Let AI answer customers

Enable customer-facing replies for a chosen scope, with rules for handing work to a person.

Support procedures and resolution

Create, test and publish support procedures with approved actions and human handover.