Every call reviewed, not a sample
A supervisor can only manage the quality of the calls somebody listens to, and nobody can listen to many. Sampling is not a methodology anyone chose — it is what listening costs.
Client anonymised; the contact centre is in Slovenia. Figures as at 4 September 2026.
The problem
A contact centre can only manage the quality of the calls someone listens to — and nobody can listen to many.
A supervisor reviews a handful of calls a week. Not the calls that went wrong: the calls that happened to be picked. Everything outside that sample is unmanaged, and the sample is far too small to be representative of anything.
That creates three problems at once, and they compound.
Coaching is generic. A supervisor who has heard four calls this week can talk about standards in general. They cannot tell an agent "on this call, at this point, you became defensive and the customer left without a resolution." The specific conversation is the one that changes behaviour, and the evidence for it is missing.
Failures stay invisible. A customer whose situation needed escalating, and did not get escalated, is a lost account waiting to happen. If that call is not in the sample, nobody learns about it — not that week, not ever.
And the reason for all of it is cost. Sampling is not a methodology anyone chose. It is what you do when listening to a call takes as long as the call.
What Red Bumerang built
A system that reviews every call, unsupervised, overnight — and hands the supervisor a ranked list of the ones worth their attention.
Each weeknight it collects the previous day's recordings from the client's telephony system, transcribes them, works out which speaker is the agent, and scores that agent's handling across eight quality dimensions. Each score carries a written reason, so the supervisor receives an argument they can coach from, not a number.
Calls that breach thresholds are flagged by severity and ranked — raised voice first, then critical findings, then lowest score. The supervisor opens the dashboard to the worst calls of the month, already in order. They are no longer choosing what to listen to. They are acting on what the evidence points at.
The rhythm is what makes it usable. The run happens overnight, unattended, every working day. Nobody starts it and nobody waits for it. A manager arrives in the morning to the previous day's flagged calls already waiting — not a report commissioned three weeks after the fact, when the customer has gone and the agent cannot remember the conversation.
Three decisions that make it usable rather than merely automatic
It knows when not to score a call. Bereavement calls, medical crises and administrative transfers are excluded before scoring. An agent is never marked down on a call that nobody could have handled well. Of the corpus, 4,787 calls were set aside this way — most of them simply too short to judge.
It knows a loud voice is not an angry one. A raised voice is registered only when the acoustic measurement and the emotion model both agree. Either signal alone would bury the supervisor in false alarms, and a supervisor who stops trusting the flags is back to sampling.
It works in the language the calls are in. Transcription, the greeting and vocabulary patterns that identify the agent, and the scoring itself all operate directly in Slovenian — including colloquial speech and its common mistranscriptions. Nothing is translated into English as an intermediate step.
What it produced
As of 4 September 2026, over a corpus running since October 2025:
Those 83 calls are the clearest answer to what full coverage buys. Each is a customer situation the system judged should have been escalated and was not. A weekly sample of a handful would have found perhaps one of them, by luck.
And the cost line is the reason any of this is possible. Sampling exists because listening is expensive. At under a dollar a night, the reason to sample disappears.
| Recordings processed | 16,400 (443.1 hours) |
|---|---|
| Calls fully scored | 11,613 |
| Calls flagged for review | 4,262 — 179 critical, 4,083 warning |
| Escalations needed and not performed | 83 |
| Quality dimensions per call | 8, each with a written justification |
| A typical night | 65–92 calls in about 25 minutes, unattended |
| Processing cost | under $1 per night |
Method
What the numbers count. 16,400 is every recording the client's system made available in the period, transcribed and passed through analysis — not a subset chosen by us. 443.1 hours is the summed duration of those recordings. Both figures grow every weeknight; they are stated as at 4 September 2026.
What is measured and what is judged. Durations, talk ratios, silence and turn counts are measured. The eight quality scores, the escalation assessment and the behavioural flags are a model's judgement of the conversation. The 83 missed escalations are calls the system assessed as mishandled — they have not been individually confirmed by a reviewer.
Scope. One Slovenian contact centre, one queue, one language. Short transactional calls averaging 1.62 minutes, inbound and outbound.
---
Bring us a process, or a market.
Every system here started as a conversation about work that was already happening — how it actually runs, where it costs the most, and what the order underneath it looks like. That is the first meeting.