How we review

Our method in full: where facts come from, how the five-part rubric turns them into a score, what the score cannot tell you, and what we have not tested yet.

By the ai-agents.reviews teamPublished September 18, 2026Updated September 18, 2026Prices and figures verified September 18, 2026

The short version

  • We score products on five published criteria, using rules anyone can apply and check.
  • Every price and percentage is stored once, with its source and the date we read it, and every page in every language is built from that record.
  • We separate what a vendor claims from what a named customer or third party reports.
  • We have not run hands-on tests yet. No page on this site says or implies that we did.
  • There are no affiliate links, paid placements or sponsored rankings here.

What we do, and what we do not do yet

Most "best AI agent" lists are written from vendor marketing pages and other lists, then presented as testing. We think that is the central problem in this category, so we start by being exact about our own work.

What we do today: desk research. We read vendor pricing pages, documentation, security pages and contracts; published case studies from named customers; and independent reporting and procurement data. We record what each source says, link to it, and score products on how much a buyer can verify before signing.

What we do not do yet: hands-on testing. We have not put these agents through a shared ticket set, so we publish no measured resolution rates of our own and award no points for them. When that changes, this page and the changelog will say so, and the results will be published in full.

Where our facts come from

  1. The vendor's own pricing pages, documentation and legal terms — for that vendor's prices, billing definitions, channels, languages, hosting and certifications.
  2. Case studies that name the customer — for results achieved in production, read with the knowledge that vendors choose which customers to publish.
  3. Independent reporting and data — established press, analyst research, and procurement data that shows what buyers actually paid.
One rule we do not bend: a vendor's content is evidence about that vendor's own product only. We do not use one vendor's blog post as the source for a competitor's price or performance. If that is the only source that exists, we say so next to the figure, or we leave the figure out.

If we cannot find a source for something, the page says "not documented" rather than guessing. Estimates are labeled as estimates, with the name of whoever made them.

One product database, so pages cannot contradict each other

Review sites often quote one price in a ranking and a different one in a comparison, because each article was written separately. Here, every fact lives in a single product record with its source and the date it was read. Reviews, pricing pages, comparisons, the calculator and all four language editions are generated from that record. When a vendor changes a price, we change it once and every page updates.

Prices are shown in the currency the vendor publishes, which in this category is almost always US dollars. We do not convert them, because an invented exchange rate is not a price anyone can buy at.

How we score

Rubric v1.0 has five criteria. Each product gets a level from 1 to 5 on each criterion, using the written rules below. The level is multiplied by the criterion's weight and the results are added up to give a score out of 10. Because the rules are public, you can check any score yourself and tell us if we got one wrong.

Pricing transparency — 25%

Can a buyer work out the bill before talking to sales?

LevelWhat it means
5/5The full price list is public, the billable unit is defined in writing, and minimums or overage rules are stated.
4/5Unit prices are public and the billable unit is defined, but the total depends on seats or another product you must also buy.
3/5Some prices are public, but a material part of the bill, such as a platform fee, usage rate or minimum, is quote-only or undefined.
2/5Quote only, but the vendor documents its pricing model (for example per resolution or per outcome).
1/5No prices and no public description of how the vendor charges.

Performance evidence — 25%

What stands behind the resolution-rate claims?

LevelWhat it means
5/5Named-customer figures plus independent data, such as press, analyst or procurement reporting, that a buyer can set against the vendor's claim.
4/5Figures from named customers in published case studies.
3/5A vendor figure with a published definition or methodology, but no named customers.
2/5A headline figure from the vendor, with no definition of how it is measured.
1/5No performance figures published.

Access and time to pilot — 20%

How quickly can a team try it on real tickets?

LevelWhat it means
5/5Self-serve sign-up or trial, and it runs on top of the helpdesk you already use.
4/5Self-serve sign-up or trial, but only on the vendor's own helpdesk or platform.
3/5Sales-led, but a pilot or go-live within weeks is documented.
2/5Sales-led, with a multi-month implementation.
1/5Enterprise sales only, with reported minimum contracts in six figures.

Channel and language coverage — 15%

Does it work where your customers are?

LevelWhat it means
5/5Chat, email, voice and messaging; 40 or more languages including German and Spanish; integrations with major helpdesks, CRMs and commerce platforms.
4/5Chat, email, and voice or messaging channels, with 30 or more languages.
3/5Chat and email, major languages, and the main helpdesks.
2/5Chat plus one other channel, or limited language and integration support.
1/5One channel, with few documented languages or integrations.

Control and compliance — 15%

Can you govern what it says, and prove it afterwards?

LevelWhat it means
5/5Everything in level 4, plus documented EU data residency and reporting or audit trails for AI answers.
4/5Handoff, guardrails and testing or simulation tools; SOC 2 or ISO 27001; a GDPR data processing agreement.
3/5Handoff and content controls are documented, with SOC 2 or ISO 27001.
2/5Human handoff is documented; little else is.
1/5No documented human handoff, guardrails or security certifications.

The ranking

ProductPricing transparencyPerformance evidenceAccess and time to pilotChannel and language coverageControl and complianceScore
Intercom Fin5/54/54/55/55/59.1
Tidio Lyro3/54/55/54/54/57.9
Freshworks Freddy AI Agent4/54/54/54/53/57.7
HubSpot Customer Agent4/54/54/53/54/57.7
Zendesk AI Agents3/54/54/54/54/57.5
Ada2/55/53/54/55/57.4
Help Scout AI Answers4/54/54/51/53/56.8
Gladly AI3/54/54/53/52/56.6
Sierra2/54/52/54/54/56.2
Salesforce Agentforce3/54/53/52/52/55.9
Decagon2/54/51/54/54/55.8
Forethought2/54/52/53/53/55.6

Claimed versus reported

Every vendor in this category advertises a resolution rate, and the numbers are not comparable. Some count any conversation the customer abandons as resolved; others require the customer to confirm. So wherever we show a performance figure, we label who said it:

  • Vendor claim — a number from the vendor's marketing, usually an average across customers.
  • Vendor, with a published definition — the vendor also explains how the number is measured.
  • Named customer — a figure a named company reported in a published case study.
  • Independent third party — press, analysts or procurement data not controlled by the vendor.

We explain the measurement problem in What counts as a resolution? and Resolution rate vs deflection rate.

What the score does not tell you

The score tells you how transparent, evidenced, accessible, broad and governable a product is, judged from public information. It does not tell you what share of your tickets the agent will resolve. That depends on your knowledge base, your ticket mix and your integrations, and the only way to find out is a pilot on your own data. Our guide to running an AI agent pilot explains how.

The rubric also favors products a team can try quickly and price without a sales call. That is deliberate, because it reflects what most buyers need, but it means enterprise-only platforms score lower on access than their capability alone would suggest. Our enterprise list reads the same data with that in mind.

Planned: a hands-on benchmark

The next version of the rubric will add a measured criterion. The protocol we intend to publish before running it:

  • One public test set of 50 support tickets across billing, order status, account access, product how-to and complaints, including questions the knowledge base cannot answer.
  • The same knowledge base loaded into every agent that allows self-serve access.
  • Scored for correct resolution, wrong answers given with confidence, appropriate handoff, and setup time.
  • Full transcripts and scoring sheets published, so anyone can re-score them.
Status: not run. Vendors that sell only through sales teams cannot be tested this way without their cooperation; they will stay scored on documented evidence and be labeled accordingly.

Keeping pages current

Every page shows three dates: when it was published, when it was last updated, and when its prices and figures were last checked against their sources. We re-verify the product database monthly and whenever a vendor announces a pricing change. Material changes are recorded in the changelog.

Independence and funding

If the publisher sells a product in this category, that product is labeled "Publisher's product" wherever it appears and is scored with the same rules as everything else. We do not use affiliate links, and links to vendors carry no tracking parameters.

AI assistance

Pages are researched and drafted with AI assistance, working from the sources listed at the bottom of each page. Where a page has been checked by a named editor, the byline says who. We would rather tell you this than put an invented name on the page.

Corrections

If a price is stale, a fact is wrong or a score does not follow from the rules above, use the "Report an error" link on any page. We fix confirmed errors in the product database, which corrects every page at once, and we log material corrections in the changelog.

This page was researched and drafted with AI assistance from the sources listed on it. We have not run hands-on tests of these products. Method: How we review