A pilot is the only way to learn what an AI agent will resolve for you, because vendor averages come from other companies' tickets. A useful pilot takes about four weeks, covers a defined slice of real traffic, and has success criteria written down before it starts.
Before you start
- Get a baseline. Record four weeks of volume by topic, first response time, resolution time, satisfaction and cost per ticket.
- Fix your help content. The agent can only answer what is written down. Update the twenty articles behind your most common questions.
- Choose the slice. Start with one channel and the topics that are frequent, well documented and low risk. Leave out complaints, cancellations and anything involving money until you trust the agent.
- Agree the definitions. Write down what counts as resolved, using this guide, and use the same definition for every vendor you test.
What to measure
| Metric | Why it matters |
|---|---|
| Resolution rate, by your definition | The number the business case rests on. |
| Satisfaction on AI-handled conversations | A high resolution rate with falling satisfaction is a warning. |
| Repeat contacts within seven days | Shows whether "resolved" customers were really helped. |
| Wrong answers found in a manual sample | Review at least 100 conversations by hand each week. |
| Handoff quality | Does the human get the context, or does the customer start again? |
| Real cost | Usage charges plus seats plus the time your team spent tuning. |
A four-week plan
- Week 1: connect the knowledge base, test internally, write the handoff rules.
- Week 2: go live on 10 to 20% of traffic in the chosen slice. Review every conversation.
- Week 3: fix content gaps, widen to 50%, begin sampling instead of reviewing everything.
- Week 4: hold settings steady and measure. This is the week your numbers come from.
Decide the go or no-go criteria in advance
Write down the thresholds before the pilot, for example: a resolution rate of at least 40% on the chosen topics, satisfaction within a few points of your human baseline, fewer than two confidently wrong answers per hundred, and a cost per resolved conversation below your human cost. Deciding afterward is how weak results get explained away.
Frequently asked questions
How much traffic do I need for a meaningful pilot?
Aim for at least a few hundred AI-handled conversations in the measurement week. Below that, a handful of unusual tickets can move the resolution rate by several points.
Should I pilot two vendors at once?
If you can, yes: split the same slice of traffic between them and use the same definition of resolved. It is the only like-for-like comparison you will get.
Keep reading
- What counts as a resolution?guide
- Resolution rate vs deflection rateguide
- Per-resolution vs per-seat pricingguide
- Agentic customer service, explainedguide
- The best AI customer service agents in 2026best list
This page was researched and drafted with AI assistance from the sources listed on it. We have not run hands-on tests of these products. Method: How we review