What to measure on an AI front desk (and what the default dashboard won't tell you.)
Published August 2026 · by Kalina, founder of Promptive
Message volume is the most reported and least useful number in AI support.
Every tool puts it on the front page. Conversations this week, up 12%. It tells you the agent is switched on. It tells you nothing about whether it made you money, saved you a hire, or quietly lost you eleven trial bookings on a Tuesday night.
I build a custom analytics layer for every front desk I ship, because the stock dashboard in Chatwoot, Intercom or anything else reports on the tool. You need reporting on the business.
Here is what I put on it, and what each metric actually caught.
Start with one number: resolution rate
If you track nothing else, track this.
Resolution rate is the percentage of conversations the agent finished on its own, with no human touching it. Not "conversations handled." Not "messages sent." Finished, and the member did not come back with the same question.
On one gym deployment, that number was 76% in the first month across 2,326 conversations. The other 24% went to a human with the full context attached.
Why this one first: it is the only metric that converts directly into a staffing decision. 76% resolution across that volume worked out to roughly 150 staff hours the front desk did not spend. That is the number you take to a budget conversation.
Two ways people inflate it, both worth knowing so you can check your own:
Counting a conversation as resolved when the member simply stopped replying. Silence is not resolution. Sometimes it is a member giving up.
Counting deflection as resolution. Deflection rate is the share of conversations that never reached a human. Resolution is the share that got answered correctly. An agent that stonewalls everyone has perfect deflection and terrible resolution. If your dashboard only shows one of the two, it is showing you the flattering one.
The eight metrics, and what each one is for
| # | Metric | Answers | Check |
|---|---|---|---|
| 1 | Resolution rate | Is it working? | Weekly |
| 2 | Inquiry volume by hour | When do we lose people? | Monthly |
| 3 | Conversation to booking | Does it sell? | Weekly |
| 4 | Drop-off point | Where do they leave? | Monthly |
| 5 | Sentiment by branch | What's breaking on the floor? | Weekly |
| 6 | Escalation reason | What can't it do yet? | Monthly |
| 7 | Topic distribution | What should we fix upstream? | Quarterly |
| 8 | Branch comparison | Who's the outlier? | Quarterly |
1. Resolution rate
Covered above. One number, checked weekly, plotted as a line so you can see it move after every change you make to the agent.
2. Inquiry volume by hour
Plot every incoming message against the hour it arrived, then draw a vertical line where your front desk closes.
On the gym build, 60% of member messages arrived outside opening hours. The staff were answering them the next morning, by which time a share of those prospects had already joined whoever replied first.
That single chart is usually the one that ends the debate about whether an AI front desk is worth it. It is also the fastest one to produce, because you already have the timestamps.
3. Conversation to booking rate
The percentage of conversations that ended in a booked trial, a booked consultation, or a captured lead in the CRM.
The gym deployment pushed roughly 200 leads into the CRM in month one that had previously gone cold overnight.
Split this by offer if you run more than one. You will find one offer converting at two or three times the others, and it is rarely the one you promoted hardest.
4. Drop-off point
Where in the conversation people stop replying.
This is the most underused metric in the set and the one that pays back fastest. If 40% of prospects leave at the moment the agent asks for a phone number, that is not an AI problem, that is a form-field problem, and you can fix it in ten minutes.
Track it as a funnel: opened, engaged, qualified, converted. The biggest single drop between two stages is your next job.
5. Sentiment by branch
Not a sentiment score for its own sake. A sentiment score with a location attached and an alert on it.
Frustration in member messages shows up weeks before it shows up in a Google review or a cancellation. When one branch starts generating negative sentiment and the others do not, that is an operational problem on that floor, and it is usually something concrete: a broken machine, a class that keeps getting cancelled, a billing change nobody announced.
Set an alert on the spike, not a report on the average. Averages hide exactly the thing you want to catch.
6. Escalation reason
Every time the agent hands off to a human, log why.
Do not just count escalations, categorise them. After a month you get a ranked list of what the agent cannot do yet. That list is your build backlog, written by your members instead of by you.
Watch for the loop as well: conversations where the agent answers, the member rephrases, the agent answers the same way again. That pattern is a knowledge gap presenting itself as a working conversation, and volume metrics will never show it to you.
7. Topic distribution
What people actually ask about, ranked.
Billing, membership freezes, class schedules and opening hours will be most of it. That is worth knowing for the agent, but it is worth more for everything upstream. If 18% of your inbound volume is "how do I freeze my membership," the answer is not a better chatbot reply. The answer is a visible freeze button in the member app.
Quarterly is often enough here. It moves slowly.
8. Branch comparison
Only if you run more than one location.
Same metrics, side by side, one row per branch. You are not looking for a winner. You are looking for the outlier, because an outlier is either a practice worth copying or a problem worth visiting.
What the stock dashboard gives you, and what it leaves out
| Stock dashboard | What you actually need |
|---|---|
| Total conversations | Resolution rate, and the trend |
| Average response time | Inquiry volume by hour, against your opening hours |
| Agent activity | Conversation to booking rate |
| Open vs closed tickets | Drop-off point in the funnel |
| A sentiment score | Sentiment by branch, with an alert |
| Tags | Ranked escalation reasons |
The pattern is the same across every row. The stock version reports on the software. The version you want reports on the business the software is sitting in front of.
None of this needs a new platform. It is a reporting layer built on data your existing tool already collects and does not surface.
What I don't put on the dashboard
Anything nobody will act on.
I have watched clients ask for twenty-metric dashboards and then check none of them, because a screen with twenty numbers on it has no obvious first move. Eight is already generous. If you are starting from nothing, ship three: resolution rate, volume by hour, conversation to booking. Add the rest when someone asks a question the first three cannot answer.
I also leave off anything the agent cannot influence. Total membership, revenue per branch and churn belong on a different screen. Mixing them in makes the dashboard feel more serious and makes it much harder to tell what your agent changed.
Frequently asked questions
What is a good resolution rate for an AI front desk?
For a gym or clinic front desk answering routine member questions, 70% to 80% is a realistic first-month range once the agent is trained on real documents. Anything advertised above 95% is usually measuring deflection, or counting silence as success.
What's the difference between deflection rate and resolution rate?
Deflection is the share of conversations that never reached a human. Resolution is the share that were answered correctly and did not come back. An agent that refuses to escalate scores well on the first and badly on the second. Track both.
How do I know if my AI agent is losing leads?
Look at the drop-off point in your conversation funnel and at conversation to booking rate. If people engage and then disappear at a consistent step, that step is the problem. Volume metrics will show the same conversation count either way.
How often should I review these metrics?
Resolution rate, sentiment alerts and booking rate weekly. Volume by hour, drop-off and escalation reasons monthly. Topic distribution and branch comparison quarterly. Reviewing everything weekly is how dashboards stop getting opened.
Can I get these from Chatwoot or Intercom directly?
Partly. Volume, response time and open ticket counts come as standard. Resolution rate as defined here, drop-off point, escalation reasons and branch-level sentiment need a reporting layer on top of the conversation data. That layer is what I build alongside the agent.
Where to start
Pick resolution rate and volume by hour. Both come from data you already have, and between them they answer the two questions that decide whether the system stays: is it working, and is it catching the people your staff cannot.
I build this layer into every front desk I ship, so the client can tell what the system did without asking me. See the numbers from a live build → or book a call if you want to know what your own conversation data is already showing.