The Number on the Slide

An assortment of printed numbers in different sizes and fonts, jumbled together
Photo by Nick Hillier on Unsplash

I spent some time last week looking into AI customer-support agents, the tools that handle tickets so your team doesn’t have to. I wasn’t shopping or trying to buy anything, I just wanted to see what a buyer actually finds when they try to compare these products on their merits. What I found was one product with at least three different accuracy stories, each published, each technically true, but all useless when it comes to making a choice.

The product was Intercom’s Fin, one of the bigger names in the space. I’m using it as the example because it’s one of the more open vendors in the field. Their own site cites an average resolution rate of 76% across 12,000 customers . A competitor’s pricing analysis, pulling from Intercom’s own customer case studies, puts real-world rates between 42% and 50% (the source sells a rival product, but the case studies it cites are Intercom’s own). The same vendor page also reports head-to-head testing that scores Fin at 73%, with a competing product at 49%, described as independent testing without naming who ran it.

Look at that spread again: 42, 50, 73, 76. Two of those numbers come from the vendor’s own site; the other two trace straight back to the vendor’s own customer case studies. If these were financial reports, someone would already be on the phone to a lawyer.

Nobody’s lying. Each number just answers a different question. What counts as “resolved”? A ticket the AI closed and the customer never reopened? A ticket the customer abandoned mid-conversation? Does the denominator include every ticket, or only the ones the vendor considers in scope for AI? Every one of those choices moves the number by double digits, and every vendor makes the choices differently. The deck never shows you the choices, it shows you the output.

A vendor metric is an answer to a question the vendor chose. The 76, the 73, and the 42 to 50 aren’t even answering the same question as each other, which is how one product ends up with a spread this wide. And the head-to-head has its own problem: put Fin’s 73 next to the competitor’s 49, and what you’re holding is an evaluation published by the party that won it, scored on a definition it picked, by a tester the page never names. The comparison measures marketing.

I’ve seen this pattern from the other side of my work. I build verification systems, and I can tell you that honest measurement, before anyone compresses it for a slide, never produces one number. It produces a range that depends on the test data and a dial someone has to set. When my team evaluates a model, the result is a curve and a set of failure cases, and reasonable people argue about what it means. The single confident number only appears at the end, after somebody decides which question to answer. The 34-point spread I found in those support tools is what the truth usually looks like before that decision gets made.

So what does a buyer do with this, short of becoming a data scientist? I’d ask three questions (Note: these apply to just about any AI category, not just support tools).

First: what, exactly, does this number count? Make the vendor walk through the definition, including what’s excluded from the denominator. This takes ten minutes and routinely deflates a headline figure by a third.

Second: whose data produced it? A number from the vendor’s benchmark, the vendor’s best customer, and your own sample are three different species. Only the last one predicts anything about you.

Third: will you help us produce this number on our cases? Give the vendor a couple hundred real examples from your own operation, including the ugly ones, and measure with your definitions. The good vendors say yes. I’d go further: the willingness to be measured on your ground is itself one of the strongest signals you can get, because the vendors who refuse are telling you what they think the answer would be.

You don’t have to be cynical to ask for any of this. The people publishing these numbers are mostly describing their product in the best light they can defend, which is what marketing is for. The problem is on our end: treating a marketing artifact as a real measurement, then wondering later why production doesn’t look like the deck.

Share

Get weekly insights on technology leadership

One idea per issue. No spam.