The short version
- We score products on five published criteria, using rules anyone can apply and check.
- Every price and percentage is stored once, with its source and the date we read it, and every page in every language is built from that record.
- We separate what a vendor claims from what a named customer or third party reports.
- We have not run hands-on tests yet. No page on this site says or implies that we did.
- There are no affiliate links, paid placements or sponsored rankings here.
What we do, and what we do not do yet
Most "best AI agent" lists are written from vendor marketing pages and other lists, then presented as testing. We think that is the central problem in this category, so we start by being exact about our own work.
What we do today: desk research. We read vendor pricing pages, documentation, security pages and contracts; published case studies from named customers; and independent reporting and procurement data. We record what each source says, link to it, and score products on how much a buyer can verify before signing.
What we do not do yet: hands-on testing. We have not put these agents through a shared ticket set, so we publish no measured resolution rates of our own and award no points for them. When that changes, this page and the changelog will say so, and the results will be published in full.
Where our facts come from
- The vendor's own pricing pages, documentation and legal terms — for that vendor's prices, billing definitions, channels, languages, hosting and certifications.
- Case studies that name the customer — for results achieved in production, read with the knowledge that vendors choose which customers to publish.
- Independent reporting and data — established press, analyst research, and procurement data that shows what buyers actually paid.
If we cannot find a source for something, the page says "not documented" rather than guessing. Estimates are labelled as estimates, with the name of whoever made them.
One product database, so pages cannot contradict each other
Review sites often quote one price in a ranking and a different one in a comparison, because each article was written separately. Here, every fact lives in a single product record with its source and the date it was read. Reviews, pricing pages, comparisons, the calculator and all four language editions are generated from that record. When a vendor changes a price, we change it once and every page updates.
Prices are shown in the currency the vendor publishes, which in this category is almost always US dollars. We do not convert them, because an invented exchange rate is not a price anyone can buy at.
How we score
Rubric v1.0 has five criteria. Each product gets a level from 1 to 5 on each criterion, using the written rules below. The level is multiplied by the criterion's weight and the results are added up to give a score out of 10. Because the rules are public, you can check any score yourself and tell us if we got one wrong.
Pricing transparency — 25%
Can a buyer work out the bill before talking to sales?
| Level | What it means |
|---|---|
| 5/5 | The full price list is public, the billable unit is defined in writing, and minimums or overage rules are stated. |
| 4/5 | Unit prices are public and the billable unit is defined, but the total depends on seats or another product you must also buy. |
| 3/5 | Some prices are public, but a material part of the bill, such as a platform fee, usage rate or minimum, is quote-only or undefined. |
| 2/5 | Quote only, but the vendor documents its pricing model (for example per resolution or per outcome). |
| 1/5 | No prices and no public description of how the vendor charges. |
Performance evidence — 25%
What stands behind the resolution-rate claims?
| Level | What it means |
|---|---|
| 5/5 | Named-customer figures plus independent data, such as press, analyst or procurement reporting, that a buyer can set against the vendor's claim. |
| 4/5 | Figures from named customers in published case studies. |
| 3/5 | A vendor figure with a published definition or methodology, but no named customers. |
| 2/5 | A headline figure from the vendor, with no definition of how it is measured. |
| 1/5 | No performance figures published. |
Access and time to pilot — 20%
How quickly can a team try it on real tickets?
| Level | What it means |
|---|---|
| 5/5 | Self-serve sign-up or trial, and it runs on top of the helpdesk you already use. |
| 4/5 | Self-serve sign-up or trial, but only on the vendor's own helpdesk or platform. |
| 3/5 | Sales-led, but a pilot or go-live within weeks is documented. |
| 2/5 | Sales-led, with a multi-month implementation. |
| 1/5 | Enterprise sales only, with reported minimum contracts in six figures. |
Channel and language coverage — 15%
Does it work where your customers are?
| Level | What it means |
|---|---|
| 5/5 | Chat, email, voice and messaging; 40 or more languages including German and Spanish; integrations with major helpdesks, CRMs and commerce platforms. |
| 4/5 | Chat, email, and voice or messaging channels, with 30 or more languages. |
| 3/5 | Chat and email, major languages, and the main helpdesks. |
| 2/5 | Chat plus one other channel, or limited language and integration support. |
| 1/5 | One channel, with few documented languages or integrations. |
Control and compliance — 15%
Can you govern what it says, and prove it afterwards?
| Level | What it means |
|---|---|
| 5/5 | Everything in level 4, plus documented EU data residency and reporting or audit trails for AI answers. |
| 4/5 | Handoff, guardrails and testing or simulation tools; SOC 2 or ISO 27001; a GDPR data processing agreement. |
| 3/5 | Handoff and content controls are documented, with SOC 2 or ISO 27001. |
| 2/5 | Human handoff is documented; little else is. |
| 1/5 | No documented human handoff, guardrails or security certifications. |
The ranking
| Product | Pricing transparency | Performance evidence | Access and time to pilot | Channel and language coverage | Control and compliance | Score |
|---|---|---|---|---|---|---|
| Intercom Fin | 5/5 | 4/5 | 4/5 | 5/5 | 5/5 | 9.1 |
| Tidio Lyro | 3/5 | 4/5 | 5/5 | 4/5 | 4/5 | 7.9 |
| Freshworks Freddy AI Agent | 4/5 | 4/5 | 4/5 | 4/5 | 3/5 | 7.7 |
| HubSpot Customer Agent | 4/5 | 4/5 | 4/5 | 3/5 | 4/5 | 7.7 |
| Zendesk AI Agents | 3/5 | 4/5 | 4/5 | 4/5 | 4/5 | 7.5 |
| Ada | 2/5 | 5/5 | 3/5 | 4/5 | 5/5 | 7.4 |
| Help Scout AI Answers | 4/5 | 4/5 | 4/5 | 1/5 | 3/5 | 6.8 |
| Gladly AI | 3/5 | 4/5 | 4/5 | 3/5 | 2/5 | 6.6 |
| Sierra | 2/5 | 4/5 | 2/5 | 4/5 | 4/5 | 6.2 |
| Salesforce Agentforce | 3/5 | 4/5 | 3/5 | 2/5 | 2/5 | 5.9 |
| Decagon | 2/5 | 4/5 | 1/5 | 4/5 | 4/5 | 5.8 |
| Forethought | 2/5 | 4/5 | 2/5 | 3/5 | 3/5 | 5.6 |
Claimed versus reported
Every vendor in this category advertises a resolution rate, and the numbers are not comparable. Some count any conversation the customer abandons as resolved; others require the customer to confirm. So wherever we show a performance figure, we label who said it:
- Vendor claim — a number from the vendor's marketing, usually an average across customers.
- Vendor, with a published definition — the vendor also explains how the number is measured.
- Named customer — a figure a named company reported in a published case study.
- Independent third party — press, analysts or procurement data not controlled by the vendor.
We explain the measurement problem in What counts as a resolution? and Resolution rate vs deflection rate.
What the score does not tell you
The score tells you how transparent, evidenced, accessible, broad and governable a product is, judged from public information. It does not tell you what share of your tickets the agent will resolve. That depends on your knowledge base, your ticket mix and your integrations, and the only way to find out is a pilot on your own data. Our guide to running an AI agent pilot explains how.
The rubric also favours products a team can try quickly and price without a sales call. That is deliberate, because it reflects what most buyers need, but it means enterprise-only platforms score lower on access than their capability alone would suggest. Our enterprise list reads the same data with that in mind.
Planned: a hands-on benchmark
The next version of the rubric will add a measured criterion. The protocol we intend to publish before running it:
- One public test set of 50 support tickets across billing, order status, account access, product how-to and complaints, including questions the knowledge base cannot answer.
- The same knowledge base loaded into every agent that allows self-serve access.
- Scored for correct resolution, wrong answers given with confidence, appropriate handoff, and setup time.
- Full transcripts and scoring sheets published, so anyone can re-score them.
Keeping pages current
Every page shows three dates: when it was published, when it was last updated, and when its prices and figures were last checked against their sources. We re-verify the product database monthly and whenever a vendor announces a pricing change. Material changes are recorded in the changelog.
Independence and funding
If the publisher sells a product in this category, that product is labelled "Publisher's product" wherever it appears and is scored with the same rules as everything else. We do not use affiliate links, and links to vendors carry no tracking parameters.
AI assistance
Pages are researched and drafted with AI assistance, working from the sources listed at the bottom of each page. Where a page has been checked by a named editor, the byline says who. We would rather tell you this than put an invented name on the page.
Corrections
If a price is stale, a fact is wrong or a score does not follow from the rules above, use the "Report an error" link on any page. We fix confirmed errors in the product database, which corrects every page at once, and we log material corrections in the changelog.
This page was researched and drafted with AI assistance from the sources listed on it. We have not run hands-on tests of these products. Method: How we review