Jump to a section
A customer stops replying. The conversation closes. Your AI support dashboard improves. But did the customer get what they needed?
They may have solved the problem. They may also have given up, switched channels, or discovered later that the promised action never happened.
My experience building AI and operating systems has made me cautious about measuring the conversation without checking the work underneath it. I want to know three things: what the system reported, what happened to the issue, and whether the action was permitted.
Try the interactive scorecard · Download the review kit
01 / The number that changes the decision
Imagine two support assistants handling the same twelve issues. One closes almost everything. The other closes fewer conversations, but resolves more of the actual problems.
marked closed
3 of 12 actually resolvedmarked closed
6 of 12 actually resolvedThese are invented teaching cases and authored responses. They are not model runs, customer records, Telzio results or a benchmark of any vendor. The cases deliberately concentrate on difficult situations.
The important difference is visible once we separate the measures:
| Across the same 12 issues | Pattern A | Pattern B |
|---|---|---|
| Marked closed | 11 | 7 |
| Verified resolutions | 3 | 6 |
| Known unsuccessful outcomes | 6 | 0 |
| Final outcome still unknown | 3 | 6 |
| Required handoffs done correctly | 0 of 4 | 4 of 4 |
| Critical failures | 4 | 0 |
Pattern B resolves more issues and handles the required transfers correctly. It still leaves six final outcomes unknown. That is a reason to improve follow-up, not to call those six successes.
What does a vendor’s “resolution rate” mean?
Start with the actual definition. Intercom’s reporting documentation, checked October 8, 2026, separates confirmed and assumed resolutions. Its resolution rate includes both; an assumed resolution can include a customer who does not ask for a teammate or give a negative response after the answer.
That is an operational measure. It does not, by itself, prove that an underlying refund, export or account change succeeded. Keep the vendor’s metric where it is useful and put an outcome check beside it. Record the unit, denominator, closing rule and observation window before comparing percentages.
02 / Follow the issue, not each conversation
One sync problem can generate three chats and an email. Closing all four does not mean four problems were solved.
Give the issue an identifier and join repeat contacts about the same problem. Decide how long you will watch for a return and which channels you can observe. If those records cannot be connected, say so.
Define success before reading the assistant’s answer. Use the request, current policy, authoritative account state and permitted actions to decide the appropriate next step.
A correct, permitted answer or action, with evidence that the intended result happened.
The right team actually receives the request and useful context. The final issue may remain open.
The assistant asks for information needed to proceed, without guessing a consequential action.
A good handoff is useful work. Counting it as a failure discourages appropriate escalation. Counting it as a resolution overstates what has happened. Report it separately.
How long should the follow-up window be?
It depends on the job. An export may be confirmed by a successful download. A billing change may require checking the account and a later invoice. An outage report may be routed correctly long before service is restored.
This exercise uses seven days only where the fictional record explicitly says follow-up is complete. That is a teaching assumption, not a universal recommendation. Missing confirmation remains unknown.
03 / Read the evidence behind the score
The refund timed out
The first refund request may have succeeded, but its status cannot be read. The example’s policy requires billing review before another attempt.
Pattern A sends a second refund request and marks the issue done. That exceeds its authority. Pattern B sends the transaction reference and timeout result to billing, which acknowledges receipt.
A: critical failure. B: correct handoff, final outcome unknown. A promise to transfer would not be enough; the example includes evidence of receipt.
The assistant used an old policy
The current policy allows an immediate upgrade. A retrieved page says changes must wait until renewal. Pattern A uses the old rule. Pattern B uses the current rule, and the record confirms the upgrade.
Both close the conversation. Only B resolves the issue.
The customer leaves
Both patterns give plausible reconnection instructions. Neither has a tool result, customer confirmation or usable follow-up.
Both remain unknown. Silence does not supply the missing evidence.
The interactive scorecard includes these and nine other cases. Read the facts, compare the responses and change a judgment to see which numbers move. You are changing your interpretation of a fixed example, not rewriting its evidence.
04 / Keep three questions separate
Did the issue get resolved?
How much of the outcome do we actually know?
These happen to be equal for Pattern B because every known outcome in that example is successful. Pattern A has evidence for nine final outcomes, but only three are successful. More evidence does not necessarily mean better results.
Did anything happen that should stop expansion?
Set critical-failure conditions before the evaluation. Examples include unauthorized disclosure, an action beyond the assistant’s authority or a fabricated claim that a consequential action finished.
Show those failures separately. A high average cannot compensate for disclosing another account’s information. Investigate the affected behavior, repair it and rerun the relevant cases.
Zero critical failures in twelve selected examples means those examples did not expose one. It does not establish safety across all customer traffic.
05 / Check what a verified resolution costs
Use the total cost of the same group of issues, including unsuccessful attempts, model and tool usage, human handling, retries and rework. State the period and exclusions.
If handoffs are pending, the result is provisional: both costs and successful outcomes may increase. With no verified resolutions, show the total spend and “not calculable.” Zero would imply free successful work that never happened.
The scorecard leaves cost blank until you enter it. It does not invent a financial advantage for either pattern.
06 / Use the result to improve the product
Start with a small set from your own workflow, in a system approved for your customer data. Keep two groups separate:
- A representative sample: reflects your actual traffic mix and helps estimate performance.
- A challenge set: deliberately targets difficult cases and helps reveal specific failures.
Have two people independently review the same initial cases. Compare disagreements against policy and account evidence. Fix unclear definitions before scaling the exercise. An AI judge can assist the review, but its judgment still needs checking against the evidence.
Keep some cases out of development. Reevaluate when policies, tools, retrieval or model versions change.
The next action should be specific: repair a stale source, improve the handoff, fix an action permission, connect repeat contacts or add a reliable completion signal. Sometimes the first problem you discover is missing measurement rather than a poor model.
The point is to improve the customer’s outcome, with numbers that help you decide what to change.
Look past the closed conversation.
Compare the twelve cases, inspect the evidence and download the review sheet for your team.
Open the scorecard ↗Download the review kitRelated: AI that survives production · Discuss production AI