Scoring every conversation against a standard the business wrote
Most contact centres judge quality by listening to a handful of conversations per adviser each month, and we could find no measured figure for how small that sample is. This article sets out what the evidence says about scoring every conversation instead: where automated scoring can be trusted, where it cannot, and what a standard must look like before a score against it means anything.
On this page
The situation
A head of claims at a motor insurer runs a team that holds tens of thousands of conversations a month. Six reviewers listen to a few calls per adviser, score them on a form, and hold a calibration meeting each quarter. The quality score is 84 per cent and has been for a year.
Then a complaint reaches the regulator. A customer says nobody told him the police report placed the other driver at fault, so he spent three weeks chasing the wrong insurer. The regulator asks what he was told, and when. The answer takes two days, because someone has to find the recording and listen to it. Nobody can say whether it happened to anyone else, because the adviser's other 99 conversations that week were never reviewed.
The score is real. But it describes a sample nobody chose carefully, against a form nobody has tested, summarised as an average that hides the conversations that reach a regulator.
What the evidence shows
Nobody knows how small the sample is
The figure repeated across the industry is that contact centres review one to two per cent of conversations. We could not find a measured source for it. Its earliest traceable appearance is an unsourced sentence in an industry body's whitepaper from around 2008 [1]. The large practitioner surveys since, including one of more than 900 executives in 2022, never asked what proportion of conversations is reviewed [2].
What is measured is evaluations per adviser. A 2013 poll of 115 contact centre professionals found two thirds reviewed six or fewer calls per adviser per month [3]. Peer-reviewed case studies of four Scottish centres in 2002 found between two and roughly eight a month [4]. This evidence is secondary and dated, but it has pointed the same way for twenty years.
The arithmetic from there is ours. If an adviser handles 400 conversations a month and five are reviewed, coverage is just over one per cent. If a serious failure, say an identity check skipped, occurs once in 200 conversations, the chance a five-call sample contains one is about 2.5 per cent. Four passes out of five gives a 95 per cent confidence interval (Wilson method) of roughly 38 to 96 per cent. The centre-wide average can still be fine: about 385 random conversations estimate a proportion to within five points at 95 per cent confidence, regardless of volume. The per-adviser score, and the rare event, are close to noise.
Sampling is also often not random. In the 2022 survey, 63 per cent of executives said transactions were selected randomly, and 75 per cent said their organisation tried to relate quality scores to customer satisfaction, so a quarter did not [2].
Human reviewers agree on actions, not on impressions
Inter-rater reliability is the degree to which two people scoring the same thing give the same score. It is usually reported as kappa, which runs from 0 (agreement no better than chance) to 1 (perfect). One widely cited analysis argues that any kappa below 0.60 is inadequate for decisions [5]; the standard reference in content analysis asks for 0.80 to rely on data and 0.67 for tentative conclusions [6].
Agreement depends far more on what reviewers judge than on who they are. In a 2026 study of 38 examiners scoring 90 nursing students, agreement on "performs chest auscultation" was kappa 0.95. On "shows empathetic attitude" it was 0.30, or 0.62 on a version of the statistic adjusted for how rarely the behaviour was marked absent [7]. A meta-analysis covering 14,650 employees found two supervisors rating the same person's overall performance correlated at 0.52, meaning they shared about half the variance in their judgements [8]. In the 2002 case studies, one adviser received three different marks for exactly the same call opening [4].
Calibration meetings may do less than their reputation. A trial of 52 medical educators, 31 of them randomised, found a rater training workshop did not improve agreement or accuracy; it improved raters' confidence [9]. Meta-analyses do find that frame-of-reference training, which uses shared examples to teach what each score level looks like, improves accuracy [10]. But a two-year study of raters observing 458 teachers found drift, the tendency for a reviewer's standard to shift over time, was large at the start and never converged; variability between raters grew [11]. Continuous measurement against a fixed reference set is the response this evidence points to, though it has not itself been tested in a controlled study.
This evidence comes from medicine, education and workplace psychology. We could find no peer-reviewed study reporting kappa for commercial contact centre scorecards. That gap is itself a finding.
Language models as reviewers: strong on easy cases, weak where it matters
A language model used to score transcripts is often called an "LLM judge". The most cited study reported that GPT-4 agreed with human raters 85 per cent of the time, above the 81 per cent at which humans agreed with each other [12]. That headline needs its footnote. The 85 per cent excludes ties and cases where the model changed its verdict when the two answers were swapped. With those included, model-human agreement was 64 to 66 per cent against 63 to 67 per cent for human pairs. On a deliberately hard set of 80 near-identical answer pairs, the 2023 model gave an order-dependent verdict 35 per cent of the time; the bias shrinks when the two answers differ clearly in quality [12].
The biases are documented and measurable:
- Position. Reordering two answers let a weaker model "beat" a stronger one in 66 of 80 test cases [13].
- Length. One model's win rate on a standard benchmark swung from 22.9 to 64.3 per cent depending only on how verbose it was told to be, until length was controlled statistically [14].
- Self-preference. Models that recognise their own output score it higher [15], and a judge from the same model family as the system it scores inflates that system's results [16].
- Rubric presentation. Reversing the order of rubric levels changed roughly a quarter of one frontier model's scores; attaching a top-scored reference answer changed nearly half [17].
- Criterion type. Across 11 classification tasks, one model's agreement with humans ranged from kappa 0.84 down to minus 0.24, with the weakest results on toxicity and safety judgements [18]. Agreement is a property of the criterion, not of the judge.
The nearest evidence to compliance scoring is a 2026 benchmark of 318 synthetic dialogues in airline, healthcare and insurance settings with injected rule violations. The strongest current models detected 78 to 95 per cent of violations turn by turn, while an older model managed 51 to 57 per cent. All of them over-penalised compliant turns while missing real violations, and a model fine-tuned for the task got every turn of a conversation right in only 39 to 51 per cent of conversations [19].
Mitigations have measured effects. Swapping order and averaging raised one evaluator's agreement with a human majority from 53 to 63 per cent; sending the hardest 20 per cent of cases to humans raised it to 74 [13]. Decomposing a rubric into yes/no questions raised agreement between different judges from 0.09 to 0.48 [20]. Scoring only where the model is confident, and escalating the rest, achieved over 80 per cent human agreement on about 80 per cent of cases [22].
Binary checks buy consistency, and something is lost
Many scorecards assume yes/no items are more reliable than 1-to-5 scales. Between human raters, a 2015 systematic review of 45 studies found inter-rater reliability of 0.81 for checklists and 0.78 for global rating scales: no meaningful difference. Global ratings generalised better from one task to the next, 0.80 against 0.69 [23]. A 1999 study found experienced clinicians scored lower than trainees on binary checklists while scoring highest on global ratings [24]; the likely reason is that experts skip steps they do not need. Checklists measure thoroughness. They can penalise skill.
For automated judges the picture is different. Decomposing a rubric into yes/no questions raised agreement between twelve different judge models from 0.09 to 0.48, and raised their correlation with human judgement by between 0.01 and 0.27 depending on the model [20]. Checklists also raised agreement between human annotators, from 0.19 to 0.26, in a related study [21]. The leading contact centre standard's own guidance asks for binary scoring and for critical errors to be tracked separately from the overall score [25].
Why the average is the wrong number
Two teams both score 82. In the first, every conversation scores between 78 and 86. In the second, nine in ten score 88 and one in ten scores 28, because an identity check was skipped. Same average; one team has a one-in-ten failure heading for a regulator. That is our illustration, not data, but the principle is a classic of statistics: four datasets with identical means, variances and correlations can look entirely different when plotted [26], and a 2017 paper produced twelve, one shaped like a dinosaur [27].
Tails also matter more than the middle for how customers remember a conversation. People judge an experience largely by its worst moment and its end [28], and negative events carry more weight than positive ones across almost every domain studied [29]. Regulators think the same way. UK conduct rules require firms to monitor the outcomes of their communications and support and to identify whether "any group of retail customers is experiencing different outcomes compared to another group" [30]. The regulator's review of insurers criticised thresholds "set by using industry averages without justification" [31], and its 2026 review found that "aggregated MI can make it harder to identify whether different groups of customers with characteristics of vulnerability have different needs or experience different barriers and outcomes" [32].
Arabic is where automated scoring is weakest
Most of the evidence above is English-language. For Gulf operations, three further facts matter.
Scoring a call means scoring a transcript, and off-the-shelf speech recognition for Gulf dialect is far worse than for Modern Standard Arabic. On a Saudi speech benchmark, the largest open general-purpose model's word error rate was 28 per cent on MSA and 60 per cent on Khaliji; models tuned for Arabic did better, but still worse on dialect than on MSA [33]. On Emirati-English code-switched speech, the same family of models produced more errors than words, while scoring 12 per cent on the English-only segments [34]. A transcript with every second word wrong cannot be scored for empathy or a required disclosure.
Language models also judge Arabic less well than English. In a 23-language benchmark of preference judgement, Arabic was the lowest-scoring language, at 62.8 per cent averaged across the models tested [35]. The team behind one Arabic leaderboard found its best candidate judge agreed with a single human evaluator at kappa 0.46 and wrote that this was "relatively low to base a decision on" [36].
And human raters disagree more as speech becomes more dialectal, including on sentiment [37]. The human reference used to calibrate an automated scorer is itself noisier in dialect.
What Gulf regulators ask for
In the UAE, the Central Bank's Consumer Protection Standards require licensed financial institutions to document monthly call-backs on a sample of consumers to detect inappropriate staff conduct, to run regular mystery shopping, and to monitor staff performance on fair treatment [38]. Those Standards are addressed to financial institutions licensed under the Central Bank law, so for an insurer they are best read as a reference point rather than a direct obligation. The 2010 insurance code of conduct, still in force under the Central Bank, requires each complaint to be decided within 15 days and the complaints register to be open to inspectors [39]. The Central Bank's 2026 guidance on AI, which expressly covers insurers, asks for periodic bias testing, continuous monitoring, meaningful human oversight and a route for customers to request human review [40]. In Saudi Arabia, the Insurance Authority began operations in November 2023 and prior rules remain in force until replaced [41]; we could not verify the current conduct instruments on a regulator domain and do not describe them here.
Outside the region, the EU's AI Act has prohibited inferring the emotions of people "in the areas of workplace and education institutions" since February 2025 [42]: a line worth noting between analysing a customer's sentiment and analysing an employee's.
What that means
Sampling does not miss patterns at the centre level; a few hundred random conversations estimate the average well. What it cannot see is the individual adviser, the rare failure, and the group of customers receiving a different outcome. Those are exactly what regulators ask about.
Scoring every conversation fixes coverage. It does not fix the standard. An automated scorer inherits every weakness of the form it is given, adds biases of its own, and in Arabic inherits the errors of the transcript underneath. The hypothesis we started with, that criteria should be observable, calibrated against humans and reported as a distribution, survives with three amendments.
Observable criteria are necessary but not sufficient. Binary checks make automated scores consistent and auditable, and they are the right form for the non-negotiables: identity verified, disclosure given, no claims decision implied. They are the wrong form for judgement, which needs a few well-anchored scale items kept separate from the checks rather than blended into one number.
Calibration against humans has a ceiling. Where experienced reviewers themselves agree below kappa 0.6 on a criterion, no automated score on it should carry consequences for a person. The useful comparison is not "model versus human" but "model versus a fixed set of conversations scored by several reviewers, measured continuously".
Distributions must be cut by language and channel. An Arabic-dialect call and an English email are not scored with the same reliability, and one league table penalises whoever serves the harder language.
There is a trade-off. Scoring everything raises the intensity of monitoring, and a 2002 survey of 347 advisers in two UK centres found perceived monitoring intensity strongly associated with lower well-being, moderated by job control and supervisor support [43]. Coverage should serve the team.
How we design for it
Every conversation is scored, whether an agent or a member of your team held it, and the criteria are yours. The standard is written during implementation as two separate lists: binary checks for observable requirements, and a small set of anchored judgement items with a description of what each level looks like. The two are never combined into one percentage.
The scorer sees the standard in one fixed form: the criteria, their order and their anchors do not change between conversations, and it is never shown a reference answer with a score attached. Every score carries the transcript passage it rests on, so a reviewer can see why without replaying the call. A score without a reason is treated as a defect.
Scores are reported as a distribution against your standard, with the share below it stated first, and cut by language, channel and reason for contact. Reason-for-contact classification tags each conversation against a category set you define, so a tail can be traced to what customers were asking about. Outcome classification records whether each conversation was resolved, unresolved or escalated, which gives the scoring an independent check: a high score on an unresolved conversation is a question, not a result.
Compliance flagging checks conversations against rules you configure and marks those where something was said that should not have been. A flag is a queue for a person, not a verdict. Given how often models over-penalise compliant turns and miss real violations, that queue is built to be reviewed, and the review feeds back into the criteria.
Sentiment analysis scores how the customer felt about the interaction. It is applied to the customer, not to your team, and reported separately from quality, because a customer can be unhappy with an answer that was correct and complete.
Coaching insights surface patterns from high and low scoring conversations for your team leaders: where people are finding conversations hard, and what good looks like in your own operation, against the standard you wrote. They exist to support the people who handle the conversations that need a person.
Scoring is calibrated against a fixed set of your conversations scored by your reviewers, separately for Arabic and English. We report agreement per criterion, not per scorecard, and say plainly which criteria fall below the threshold for automated scoring. Where speech recognition confidence on a call is low, the score is withheld and the call goes to a person.
What we don't know yet
No published study measures how transcription errors degrade a language model's agreement with human reviewers on the same conversations. The two effects are documented separately; their combination is not.
No published study measures agreement between language models and human raters on Arabic customer service conversations. The nearest evidence is from general-purpose benchmarks.
No published study reports kappa for commercial contact centre scorecards. The evidence on human agreement is borrowed from medicine, education and workplace psychology, and we have said so where it appears.
Whether continuous measurement against a reference set prevents drift in production, rather than only detecting it, has not been tested in a controlled study.
And the coverage figure the industry has repeated for nearly twenty years has, as far as we could find, never been measured. If you quote one, compute it from your own volumes.
Sources
- 1.ICMI, Discover why contact center quality doesn't measure up, and what you can do about it, whitepaper, c. 2008. Unsourced industry estimate.https://www.icmi.com/~/media/files/resources/whitepapers/quality-whitepaper.ashx
- 2.COPC Inc., Global Benchmarking Series 2022: Contact Center Quality Assurance, survey of more than 900 executives, fieldwork September to December 2021, pp. 11, 21 and 37. Secondary, self-reported.https://cx.copc.com/hubfs/Global%20Benchmarking%20Series%202022_Contact%20Center%20Quality%20Assurance.pdf
- 3.Call Centre Helper, Poll: how many calls do you monitor per agent per month?, webinar poll, n=115, March 2013. Secondary.https://www.callcentrehelper.com/poll-how-many-calls-do-you-monitor-per-agent-per-month-43247.htm
- 4.Bain, P., Watson, A., Mulvey, G., Taylor, P. and Gall, G., Taylorism, targets and the pursuit of quantity and quality by call centre management, New Technology, Work and Employment 17(3), 2002 (page references are to the open-access manuscript, pp. 10 to 15).https://strathprints.strath.ac.uk/4165/6/strathprints004165.pdf
- 5.McHugh, M.L., Interrater reliability: the kappa statistic, Biochemia Medica 22(3), 2012.https://biochemia-medica.com/en/journal/22/3/10.11613/BM.2012.031/fullArticle
- 6.Krippendorff, K., Reliability in content analysis: some common misconceptions and recommendations, Human Communication Research 30(3), 2004. Foundational.https://www.academia.edu/21851693/
- 7.Yayama et al., inter-rater agreement in a nursing OSCE, Sage Open Nursing, 2026 (90 students, 38 examiners).https://journals.sagepub.com/doi/10.1177/23779608261417794
- 8.Viswesvaran, C., Ones, D.S. and Schmidt, F.L., Comparative analysis of the reliability of job performance ratings, Journal of Applied Psychology 81(5), 1996 (N=14,650). Foundational.https://www.academia.edu/20224570/
- 9.Cook, D.A. et al., Effect of rater training on reliability and accuracy of mini-CEX scores: a randomized, controlled trial, Journal of General Internal Medicine, 2009 (52 preceptors, 31 randomised).https://link.springer.com/article/10.1007/s11606-008-0842-3
- 10.Roch, S.G., Woehr, D.J., Mishra, V. and Kieszczynska, U., Rater training revisited: an updated meta-analytic review of frame-of-reference training, Journal of Occupational and Organizational Psychology 85(2), 2012.https://researchconnect.suny.edu/en/publications/rater-training-revisited-an-updated-meta-analytic-review-of-frame/
- 11.Casabianca, J.M., Lockwood, J.R. and McCaffrey, D.F., Trends in classroom observation scores, Educational and Psychological Measurement, 2015 (observations of 458 teachers over two years).https://journals.sagepub.com/doi/10.1177/0013164414539163
- 12.Zheng, L. et al., Judging LLM-as-a-judge with MT-Bench and Chatbot Arena, NeurIPS 2023, Tables 2, 5 and 6.https://arxiv.org/abs/2306.05685
- 13.Wang, P. et al., Large language models are not fair evaluators, 2023, abstract and Table 4.https://arxiv.org/abs/2305.17926
- 14.Dubois, Y. et al., Length-controlled AlpacaEval: a simple way to debias automatic evaluators, 2024, Figure 3.https://arxiv.org/abs/2404.04475
- 15.Panickssery, A., Bowman, S.R. and Feng, S., LLM evaluators recognize and favor their own generations, 2024.https://arxiv.org/abs/2404.13076
- 16.Li, D. et al., Preference leakage: a contamination problem in LLM-as-a-judge, ICLR 2026.https://arxiv.org/abs/2502.01534
- 17.Li et al. (Ant Group), Evaluating scoring bias in LLM-as-a-judge, 2025, Table 3.https://arxiv.org/abs/2506.22316
- 18.Bavaresco, A. et al., LLMs instead of human judges? A large scale empirical study across 20 NLP evaluation tasks, ACL 2025, Table 1 (Cohen's kappa for GPT-4o on the 11 categorical datasets).https://arxiv.org/abs/2406.18403
- 19.Yang et al., CompliBench: benchmarking LLM judges for compliance violation detection in dialogue systems, 2026, Table 1 and Figure 6.https://arxiv.org/abs/2604.12312
- 20.Lee et al., CheckEval, EMNLP 2025, Tables 1 and 2.https://arxiv.org/abs/2403.18771
- 21.Cook, J. et al., TICK: targeted instruct-evaluation with checklists, 2024, Tables 4 and 6.https://arxiv.org/abs/2410.03608
- 22.Jung, J., Brahman, F. and Choi, Y., Trust or escalate: LLM judges with provable guarantees for human agreement, 2024.https://arxiv.org/abs/2407.18370
- 23.Ilgen, J.S. et al., A systematic review of validity evidence for checklists versus global rating scales in simulation-based assessment, Medical Education 49(2), 2015 (45 studies).https://pubmed.ncbi.nlm.nih.gov/25626747/
- 24.Hodges, B., Regehr, G., McNaughton, N., Tiberius, R. and Hanson, M., OSCE checklists do not capture increasing levels of expertise, Academic Medicine, 1999. Foundational.https://pubmed.ncbi.nlm.nih.gov/10536636/
- 25.COPC Inc., What are the guidelines for QA form design and attributes?, 2024, and Measure quality using three metrics instead of one overall score, 2016. Secondary, practitioner guidance.https://www.copc.com/clearly-cx/what-are-the-guidelines-for-qa-form-design-and-attributes/
- 26.Anscombe, F.J., Graphs in statistical analysis, The American Statistician 27(1), 1973. Foundational.https://www.tandfonline.com/doi/abs/10.1080/00031305.1973.10478966
- 27.Matejka, J. and Fitzmaurice, G., Same stats, different graphs, CHI 2017.https://www.research.autodesk.com/publications/same-stats-different-graphs/
- 28.Kahneman, D., Fredrickson, B.L., Schreiber, C.A. and Redelmeier, D.A., When more pain is preferred to less, Psychological Science 4(6), 1993. Foundational.https://journals.sagepub.com/doi/10.1111/j.1467-9280.1993.tb00589.x
- 29.Baumeister, R.F., Bratslavsky, E., Finkenauer, C. and Vohs, K.D., Bad is stronger than good, Review of General Psychology 5(4), 2001. Foundational.https://journals.sagepub.com/doi/abs/10.1037/1089-2680.5.4.323
- 30.Financial Conduct Authority, Handbook PRIN 2A.9.8R, 2A.9.10R and 2A.9.11R, in force from 31 July 2023.https://handbook.fca.org.uk/handbook/prin2a/prin2as9
- 31.Financial Conduct Authority, Insurance multi-firm review of outcomes monitoring under the Consumer Duty, 26 June 2024.https://www.fca.org.uk/publications/multi-firm-reviews/insurance-multi-firm-review-outcomes-monitoring-under-consumer-duty
- 32.Financial Conduct Authority, Outcomes monitoring: good practice and areas for improvement, 2026.https://www.fca.org.uk/publications/good-and-poor-practice/outcomes-monitoring-good-practice-and-areas-improvement
- 33.ELM Company, Open universal Arabic ASR leaderboard, 2024, Tables 1 and 3 (Whisper-large-v3, zero-shot, on the SADA dataset). Industry technical report.https://arxiv.org/abs/2412.13788
- 34.Al Ali, M. and Aldarmaki, H., Mixat: a data set of bilingual Emirati-English speech, SIGUL 2024, Table 3 (Whisper medium, zero-shot).https://arxiv.org/abs/2405.02578
- 35.Gureja, S. et al., M-RewardBench: evaluating reward models in multilingual settings, ACL 2025.https://arxiv.org/abs/2410.15522
- 36.Inception and MBZUAI, AraGen leaderboard and 3C3H, Hugging Face, December 2024. Lab technical write-up, sample size not disclosed.https://huggingface.co/blog/leaderboard-3c3h-aragen
- 37.Keleg, A. and Magdy, W., on annotator agreement and Arabic level of dialectness, 2024.https://arxiv.org/abs/2405.11282
- 38.Central Bank of the UAE, Consumer Protection Standards, Notice 1158/2021, Articles 4.1.1.8, 4.1.1.9 and 5.2.2.4.https://rulebook.centralbank.ae/en/rulebook/consumer-protection-standards
- 39.Insurance Authority Board Resolution No. 3 of 2010, Instructions concerning the code of conduct and ethics, Article 10, in force under the Central Bank of the UAE.https://rulebook.centralbank.ae/en/rulebook/insurance-authority-board-resolution-no-3-2010-instructions-concerning-code-conduct-and
- 40.Central Bank of the UAE, Guidance note on the consumer protection and responsible adoption and use of AI and ML by licensed financial institutions, 2026.https://rulebook.centralbank.ae/en/rulebook/guidance-note-consumer-protection-and-responsible-adoption-and-use-artificial-intelligence
- 41.Saudi Press Agency, The Insurance Authority officially commences its operations, 23 November 2023.https://www.spa.gov.sa/en/N2002759
- 42.Regulation (EU) 2024/1689 (AI Act), Article 5(1)(f) and Recital 44.https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-5
- 43.Holman, D., Chissick, C. and Totterdell, P., The effects of performance monitoring on emotional labor and well-being in call centers, Motivation and Emotion, 2002 (n=347).https://link.springer.com/article/10.1023/A:1015194108376
Get new research as it's published
Occasional emails when we publish.