Practical guide · Production AI

Did the AI actually solve the problem?

A closed conversation is easy to count. A resolved customer issue takes evidence. Here is a practical way to tell them apart.

Jump to a section

A customer stops replying. The conversation closes. Your AI support dashboard improves. But did the customer get what they needed?

They may have solved the problem. They may also have given up, switched channels, or discovered later that the promised action never happened.

My experience building AI and operating systems has made me cautious about measuring the conversation without checking the work underneath it. I want to know three things: what the system reported, what happened to the issue, and whether the action was permitted.

Try the interactive scorecard · Download the review kit

01 / The number that changes the decision

Imagine two support assistants handling the same twelve issues. One closes almost everything. The other closes fewer conversations, but resolves more of the actual problems.

Pattern A92%

marked closed

3 of 12 actually resolved
Pattern B58%

marked closed

6 of 12 actually resolved

These are invented teaching cases and authored responses. They are not model runs, customer records, Telzio results or a benchmark of any vendor. The cases deliberately concentrate on difficult situations.

The important difference is visible once we separate the measures:

Across the same 12 issues Pattern A Pattern B
Marked closed 11 7
Verified resolutions 3 6
Known unsuccessful outcomes 6 0
Final outcome still unknown 3 6
Required handoffs done correctly 0 of 4 4 of 4
Critical failures 4 0

Pattern B resolves more issues and handles the required transfers correctly. It still leaves six final outcomes unknown. That is a reason to improve follow-up, not to call those six successes.

What does a vendor’s “resolution rate” mean?

Start with the actual definition. Intercom’s reporting documentation, checked October 8, 2026, separates confirmed and assumed resolutions. Its resolution rate includes both; an assumed resolution can include a customer who does not ask for a teammate or give a negative response after the answer.

That is an operational measure. It does not, by itself, prove that an underlying refund, export or account change succeeded. Keep the vendor’s metric where it is useful and put an outcome check beside it. Record the unit, denominator, closing rule and observation window before comparing percentages.

02 / Follow the issue, not each conversation

One sync problem can generate three chats and an email. Closing all four does not mean four problems were solved.

Give the issue an identifier and join repeat contacts about the same problem. Decide how long you will watch for a return and which channels you can observe. If those records cannot be connected, say so.

Define success before reading the assistant’s answer. Use the request, current policy, authoritative account state and permitted actions to decide the appropriate next step.

Resolve

A correct, permitted answer or action, with evidence that the intended result happened.

Handoff

The right team actually receives the request and useful context. The final issue may remain open.

Clarify

The assistant asks for information needed to proceed, without guessing a consequential action.

A good handoff is useful work. Counting it as a failure discourages appropriate escalation. Counting it as a resolution overstates what has happened. Report it separately.

How long should the follow-up window be?

It depends on the job. An export may be confirmed by a successful download. A billing change may require checking the account and a later invoice. An outage report may be routed correctly long before service is restored.

This exercise uses seven days only where the fictional record explicitly says follow-up is complete. That is a teaching assumption, not a universal recommendation. Missing confirmation remains unknown.

03 / Read the evidence behind the score

The refund timed out

The first refund request may have succeeded, but its status cannot be read. The example’s policy requires billing review before another attempt.

Pattern A sends a second refund request and marks the issue done. That exceeds its authority. Pattern B sends the transaction reference and timeout result to billing, which acknowledges receipt.

A: critical failure. B: correct handoff, final outcome unknown. A promise to transfer would not be enough; the example includes evidence of receipt.

The assistant used an old policy

The current policy allows an immediate upgrade. A retrieved page says changes must wait until renewal. Pattern A uses the old rule. Pattern B uses the current rule, and the record confirms the upgrade.

Both close the conversation. Only B resolves the issue.

The customer leaves

Both patterns give plausible reconnection instructions. Neither has a tool result, customer confirmation or usable follow-up.

Both remain unknown. Silence does not supply the missing evidence.

The interactive scorecard includes these and nine other cases. Read the facts, compare the responses and change a judgment to see which numbers move. You are changing your interpretation of a fixed example, not rewriting its evidence.

04 / Keep three questions separate

Did the issue get resolved?

Verified resolution share6 resolved ÷ 12 attempted = 50%Pattern B · Unknown outcomes stay in the denominator.

How much of the outcome do we actually know?

Outcome evidence coverage6 known outcomes ÷ 12 attempted = 50%Known outcomes include success and failure. A received handoff does not establish final resolution.

These happen to be equal for Pattern B because every known outcome in that example is successful. Pattern A has evidence for nine final outcomes, but only three are successful. More evidence does not necessarily mean better results.

Did anything happen that should stop expansion?

Set critical-failure conditions before the evaluation. Examples include unauthorized disclosure, an action beyond the assistant’s authority or a fabricated claim that a consequential action finished.

Show those failures separately. A high average cannot compensate for disclosing another account’s information. Investigate the affected behavior, repair it and rerun the relevant cases.

Zero critical failures in twelve selected examples means those examples did not expose one. It does not establish safety across all customer traffic.

05 / Check what a verified resolution costs

Use the total cost of the same group of issues, including unsuccessful attempts, model and tool usage, human handling, retries and rework. State the period and exclusions.

Illustrative cost per verified resolution$120 total cost ÷ 6 resolved = $20An arithmetic example, not a support-cost benchmark.

If handoffs are pending, the result is provisional: both costs and successful outcomes may increase. With no verified resolutions, show the total spend and “not calculable.” Zero would imply free successful work that never happened.

The scorecard leaves cost blank until you enter it. It does not invent a financial advantage for either pattern.

06 / Use the result to improve the product

Start with a small set from your own workflow, in a system approved for your customer data. Keep two groups separate:

  • A representative sample: reflects your actual traffic mix and helps estimate performance.
  • A challenge set: deliberately targets difficult cases and helps reveal specific failures.

Have two people independently review the same initial cases. Compare disagreements against policy and account evidence. Fix unclear definitions before scaling the exercise. An AI judge can assist the review, but its judgment still needs checking against the evidence.

Keep some cases out of development. Reevaluate when policies, tools, retrieval or model versions change.

The next action should be specific: repair a stale source, improve the handoff, fix an action permission, connect repeat contacts or add a reliable completion signal. Sometimes the first problem you discover is missing measurement rather than a poor model.

The point is to improve the customer’s outcome, with numbers that help you decide what to change.

Put the method to work

Look past the closed conversation.

Compare the twelve cases, inspect the evidence and download the review sheet for your team.

Open the scorecard ↗Download the review kit

Related: AI that survives production · Discuss production AI