Back to blog
Articles

Support metrics that matter for founders: five numbers, no dashboard sprawl

FRT, resolution rate, CSAT, cost per conversation and honest AI resolution: how each is computed, the trap inside each one, and reasonable targets by stage.

Sufox TeamJune 9, 20266 min read

Key takeaways

  • Five numbers describe a support operation completely enough to make decisions: first response time, resolution rate, CSAT, cost per conversation, and honest AI resolution rate — a twenty-chart dashboard is procrastination.
  • Track FRT as a median-p90 pair, never a mean: the median describes the typical case, the p90 shows how bad 'bad' gets, and a pattern in the p90 is usually a routing or category gap, not a headcount gap.
  • Resolution rate is gameable by accident because 'resolved' is a click — pair it permanently with reopen rate, and treat low first-contact resolution as a missing-context problem before a coaching problem.
  • Cost per conversation converts hiring, automation and tooling decisions into one comparable unit, but only with a full numerator — founder hours and knowledge-base work included — and its absolute value matters less than the requirement that it falls as volume grows.
  • Count AI resolution strictly — answered, closed, stayed closed within the reopen window — because vanity deflection built on customer silence inflates the number with quiet failures and misroutes the queue.

Founders track MRR to the decimal and can quote churn from memory — then run support entirely on vibes. "The queue feels okay" is not a metric, and neither is a dashboard with twenty charts nobody reads. Between those extremes sit five numbers that describe a support operation completely enough to make decisions with: first response time, resolution rate, CSAT, cost per conversation, and AI resolution rate.

This is a founder's guide to those five — how each is actually computed, the trap hiding inside each one, and what targets are reasonable at each stage.

First response time: measure the median, not the mean

What it is: time from a customer's first message to the first substantive reply — human or AI. Auto-acknowledgments don't count; customers stopped being fooled by "we received your message" years ago.

The trap: averages. One conversation that sat over a weekend drags the mean upward until it describes nobody's experience. Track the median for the typical case and the 90th percentile for the worst realistic case; the pair tells you both what's normal and how bad "bad" gets. A median of 40 minutes with a p90 of 26 hours means your process is fine but something — a channel, a time zone, a category — is falling through. Find the pattern in the p90 before hiring anyone.

Reasonable targets: on email, under 4 business hours median for a small team; under 1 hour once support is someone's full-time job. Live chat expectations are minutes. An AI front line answers instantly, which effectively splits the metric in two — track time-to-first-answer (near zero) and time-to-human for escalated conversations, where the honest target for a small team is under 4 business hours.

Resolution rate and the reopen test

What it is: the share of conversations that actually end resolved — and, more useful day-to-day, first-contact resolution (FCR): solved in a single exchange, no back-and-forth.

The trap: "resolved" is a status an agent clicks, which makes it gameable by accident. The honest check is the reopen rate — how often customers come back on the same issue within, say, seven days. High resolution numbers with high reopens mean conversations are being closed, not solved. Pair the two permanently.

Low FCR is rarely a people problem. It usually points at missing information at first touch: agents who can't see the customer's plan, order or history will always need a second exchange. Fixing context beats coaching speed.

Reasonable targets: FCR of 65–75% is strong for a technical product; reopens under 5%.

CSAT: small samples lie

What it is: the post-resolution "how did we do?" score — typically the percentage of positive responses among all responses.

The traps are statistical. Response rates run 10–30%, and people with strong feelings respond most, so the sample skews toward extremes. Ten responses tell you nothing; a hundred begin to mean something. Second trap: CSAT measures the conversation, not the product — a customer furious about a missing feature often still rates the agent five stars. Read verbatim comments monthly; the numbers say "how good," the comments say "why."

Reasonable targets: 75–85% positive is a solid range for SMB SaaS; over 90% sustained usually means your unhappy customers have stopped filling in surveys — or stopped writing in at all, which is worse. Watch the trend and investigate step-changes; the absolute number matters less than its direction.

Cost per conversation: the metric that connects support to the P&L

What it is: everything support costs in a month — loaded salaries and the fraction of founder time, tooling, AI capacity — divided by conversations handled.

Why founders should love it: it converts "should we hire?", "should we automate?", and "is this tool worth it?" into one comparable unit. A worked hypothetical: a two-person team (say $9,000/month loaded) plus a $219 Growth workspace on Sufox handling 2,100 conversations lands at about $4.40 per conversation — and because the tooling is flat with seats unlimited, the number moves only when payroll or volume moves, which makes trends legible.

The trap: partial numerators. Count the founder's hours and the knowledge-base work, or the metric will flatter you into wrong decisions — automation ROI computed against an understated cost baseline always looks worse than it is.

Reasonable shape: the absolute value varies wildly by product; the requirement is that it falls as volume grows. Rising cost per conversation during growth means the operation is scaling linearly with something it shouldn't — usually headcount or per-seat tooling.

AI resolution rate: count it honestly or don't count it

What it is: the share of all conversations fully resolved by the AI agent — with no human touch, and no reopen or follow-up within your reopen window.

The trap: vanity deflection. Counting "AI replied and the customer went silent" inflates the number with abandoned conversations and quiet failures. The strict definition — answered, closed, stayed closed — is the only version worth steering by, and it's the one that should gate how much of the queue you route to AI in the first place.

Reasonable targets: 30–40% honest resolution soon after launch on a decent knowledge base; 50–70% after a few quarters of feeding gaps back into documentation. Watch its complement too: escalation quality. If humans regularly rescue conversations the AI mangled (rather than ones it correctly declined), tighten the AI's scope before raising its quota.

Targets by stage, in one list

  • Under ~500 conversations/month: track median FRT and reopen rate only. CSAT samples are too small; cost per conversation is dominated by fixed costs. Goal: FRT median under 4 business hours, reopens under 8%.
  • 500–3,000/month: all five metrics become meaningful. Goals: FRT median under 2 business hours, FCR above 65%, CSAT above 75% on 50+ monthly responses, cost per conversation trending down, honest AI resolution above 40%.
  • 3,000+/month: add p90s and per-category breakdowns; averages now hide entire failure modes. Goals: FCR above 70%, CSAT above 80%, AI resolution above 55%, cost per conversation falling quarter over quarter.

The five-minute weekly review

Every Monday, one screen — Sufox's analytics assembles these without spreadsheet work, but any stack that produces the five numbers qualifies: median and p90 FRT against last week; resolution rate next to reopens; CSAT trend with one verbatim comment read aloud; cost per conversation, monthly; AI resolution rate and where it missed.

Fifteen minutes a month deeper on the misses — which question types, which missing articles — is what actually moves the numbers. Metrics don't improve support; they tell you which of the three levers (documentation, process, people) to pull next. Founders who read five numbers weekly pull the right lever quarters earlier than founders who wait for the queue to hurt.

Share this article

Frequently asked questions

Under roughly 500 conversations a month, just two: median first response time (target: under 4 business hours) and reopen rate (under 8%). CSAT samples are too small to mean anything yet, and cost per conversation is dominated by fixed costs. The full set of five — FRT, resolution rate, CSAT, cost per conversation, AI resolution rate — becomes meaningful from about 500 conversations a month.

On email, a median under 4 business hours is respectable for a small team and under 1 hour is strong once support is a full-time role; chat expectations are minutes. Always track the median with the 90th percentile: a fine median with a 26-hour p90 means some channel, time zone or category is silently falling through — a routing fix, not a hiring signal. With an AI front line, also track time-to-human for escalations.

75–85% positive is a solid range for SMB SaaS. Be suspicious of sustained scores above 90% — often the unhappy customers have simply stopped answering surveys. Mind the statistics: response rates run 10–30% and skew toward strong feelings, so ten responses tell you nothing. Watch the trend rather than the absolute, and read verbatim comments monthly — the score says how good, the comments say why.

Divide everything support costs in a month by conversations handled: loaded salaries, the honest fraction of founder time, tooling, AI capacity, and knowledge-base hours. A hypothetical two-person team at $9,000 loaded plus a $219 flat Sufox workspace handling 2,100 conversations lands around $4.40. The absolute varies by product; the requirement is that it falls as volume grows.

With a strict definition — resolved with no human touch and no reopen or follow-up within about seven days — 30–40% soon after launch on a decent knowledge base is realistic, climbing to 50–70% after a few quarters of feeding missed questions back into documentation. Anything counted from customer silence is vanity deflection and will mislead every downstream decision about routing and quotas.

Five minutes weekly on one screen: median and p90 FRT versus last week, resolution rate next to reopens, the CSAT trend with one verbatim comment, cost per conversation monthly, and AI resolution rate with its misses. Then fifteen minutes a month on the misses themselves — which question types, which missing articles. The metrics don't improve support; they tell you which lever to pull: documentation, process, or people.

Ready to put AI support to work?

14 days free. Full platform. We move your data for you.