We’re operationalizing AI inside the Sentry Support team in real time, with real customers and real stakes. This series captures what works, what breaks, and what we change as it happens, with a non-zero chance it becomes a cautionary tale instead of a success story.
Measuring 1% of your angry customers via email survey is 1990s tech. We need to score 100% of conversations with AI.
The Feedback Mirage
For decades, the Customer Satisfaction (CSAT) score has been the “North Star” of Support. We chase the gold star, the “Great” rating, and the 95% satisfaction goal. Yes, the Sentry team is exceptional.
But here is the dirty secret that CS leaders hate to admit: CSAT is a lie. It’s a data set built on extremes. CSAT amplifies the loud edges. Most customers never contact support in the first place, so we learn nothing about how they feel. And even when they do, only about 10% respond. What’s left is a tiny, biased slice pretending to represent the whole business.
When you manage your business based on a 2% actual response rate, you aren’t leading; you’re squinting at a mirage. (←AI-generated phrasing)
The New North Star: CX Score
We officially retire surveyed CSAT as a main KPI.
Enter the CX Score.
This CX Score will use AI to analyze most customer conversations for resolution, sentiment, and service quality. We want to reduce reliance on manual surveys. While the exact formula will evolve over time, the approach will surface clear drivers behind the score, such as answer quality or excessive customer effort, allowing us to pinpoint specific areas for improvement rather than relying on abstract metrics. Yes, we are familiar with the trendy, pre-AI, Customer Effort Score (CES).
Instead of waiting for a biased survey, we will task our agent to evaluate conversations against a standardized rubric. The AI doesn’t care if it’s a Tuesday or if the customer had a bad cup of coffee. It looks for objective markers:
Resolution Quality: Did we actually solve the problem or just “close” the ticket?
Sentiment Shift: Did the user start frustrated and end empowered?
Effort Score: How many hoops did we make them jump through?
By scoring every single interaction, we moved from a “sample size” of X tickets to a total visibility of X^n}). That is the difference between an anecdote and an insight.
How efficient are we?
By shifting our focus to the CX Score, we prove we could grow the business while keeping the team smart and engaged. Because we aren’t measuring human effort anymore; we’re measuring Systemic Resolution. When the AI handles 80% of the volume at a “Grade A” CX Score, you don’t need to hire defenders, your team are now builders.
The “RTFM” Metric
We’ve also killed the “Ticket Volume” metric. In the old world, a high ticket volume meant the team was “busy” (which we mistakenly rewarded).
Now, we track the Automated Resolution Rate. If a customer asks a question that is clearly answered in our docs, and a human has to paste that link, we consider that a failure of the system. We call it the “RTFM” gap. Every time a human answers a “Googleable” question, it’s a signal that our AI knowledge loop is broken.
We stop rewarding ourselves for “pasting links” and start rewarding us for ensuring the AI never has to ask for that link again.
From Sentiment to Strategy
CSAT tells you how someone felt yesterday. The CX Score tells you how your product is performing today.
When you [not] score 100% of your interactions, you start seeing patterns that surveys miss. You don’t just see “support issues”; you see other friction points, documentation gaps, and UI failures in high definition.
We don’t want your five-star rating. We want zero-effort resolution for our customers.



