A Chinese podcast · Same story, 3 levels

一个人工智能客服顶了七百个人,平均满意度和人一样:一年后公司把人招了回来,因为平均值把最糟的那批客户藏起来了
One AI Did the Work of Seven Hundred Agents and Scored the Same on Average — a Year Later the Company Hired the Humans Back, Because the Average Was Hiding the Worst Customers
About this story
Klarna launched an AI customer service agent in February 2024: 2.3 million conversations in month one, the work of about 700 agents, two minutes per query against eleven. By mid-2025 it was rehiring humans. Average satisfaction had looked level the whole time, because the distribution was bimodal. Chinese listening practice at four levels. HSK 5-6 Chinese listening practice.
This is an HSK 5-6 Chinese listening episode that runs about 5 minutes. The full Mandarin script is shown with tap-for-pinyin and a line-by-line English translation, so you can listen and read at once — comprehensible input in the sense of Stephen Krashen's i+1 theory. It teaches 12 key vocabulary words such as 指标、结构、信号 and walks through 3 grammar patterns, each explained in English with examples. The same news story is retold at 3 difficulty levels — use the level selector above to find the version that is challenging but still understandable for you.
Read at your level
原文Read the complete story in Chinese. Reveal pinyin and English only when you need them.
English transcript reference
This episode covers a complete cycle that ran from 2024 to 2025, and an old problem in statistics: what an average can answer, and what it cannot.
The facts first.
In February 2024, the Swedish payments company Klarna rolled out a large-model-based AI customer service system globally.
It handled 2.3 million conversations in its first month.
The company's stated conversion: a workload equivalent to about seven hundred full-time agents.
The operating metrics published alongside it included average resolution time falling from eleven minutes to two, repeat query rate down twenty-five percent, and projected additional profit for the year of about forty million dollars.
That data set was widely cited at the time and became the most frequently offered empirical case for the "AI replaces customer service" narrative.
In the same period, a billboard appeared on the streets of San Francisco carrying a single line: stop hiring humans.
On the public information available, the matter appeared to be settled.
Move forward to mid-2025, and the company began rehiring human agents.
The founder was strikingly candid in an interview with Bloomberg.
He said the company had "gone too far", and that an excessive focus on cost had driven quality down.
The easiest reading here is "the AI wasn't good enough".
That reading does not hold, because none of the efficiency data was ever overturned.
Two minutes against eleven was real; the fall in repeat queries was real.
The question actually worth analysing is: why did a drop in quality take more than a year to reach management's field of view?
The answer lies in the choice of monitoring metric.
The core metric the company tracked was the average of customer satisfaction.
And measured on the average, the AI and human agents performed about the same.
In other words, there was no anomalous signal anywhere on the dashboard.
The statistics need spelling out here.
An average is a measure of central tendency; its function is to compress a distribution into a single number.
Compression necessarily loses information, and what it loses is precisely the shape of the distribution.
When a data set is approximately normal, unimodal and not highly variable, the average represents the whole reasonably well.
When the data is bimodal, the average lands in the trough between the two peaks — a position where almost no samples sit.
This distribution was bimodal.
Highly standardised, low-context queries — order lookups, address changes, refund status — were handled quickly and accurately, and satisfaction for those users was close to the ceiling.
Queries requiring context, involving exceptions, or coming from users already unhappy were answered at the same two-minute speed, and answered wrong.
Satisfaction for those users was close to the floor.
More to the point is who those users were: they were typically the ones genuinely in need of help.
One end near the ceiling, one near the floor; weight them together and you land precisely on "level with human agents".
The result is a characteristic state: the metric is green and the experience is broken.
Had the company been monitoring the satisfaction distribution, or the tenth percentile, or the complaint escalation rate, the signal would have appeared within weeks.
But it was monitoring the average, so the signal was structurally erased.
The subsequent adjustment came in two parts.
Operationally, traffic was split: standardised queries handled by the model, complex ones routed to humans, with explicit handoff triggers defined.
On employment, the model became flexible — no fixed shifts, but a ride-hailing-style job-acceptance mechanism, with participants including students, people in rural areas, and the company's own users.
Finally, back to the numbers.
This needs stating clearly: not one of them was false.
Seven hundred, two minutes, twenty-five percent, forty million — every one can be checked, and every one holds.
"Average satisfaction roughly level" holds equally well.
The problem was never falsified data. It was a mismatch between the question and the metric.
What these metrics answer is "what happened on average".
What a service system actually needs to know is "how bad is the worst part, and who is in it".
Those are two different questions, and one metric cannot answer both.
Two questions to leave you with.
The first is concrete: the number you currently use to judge whether something is going normally — is it an average?
If so, what has it flattened out?
The second is bigger: under what conditions does an organisation go and look at the data it was never asked to look at?
Listen again
Try it without the transcript and notice what sounds clearer.
What vocabulary does this episode teach?
词汇HSK 5. 答案在监测指标的选择上 — the whole diagnosis in one line.
HSK 5. 信号被结构性地抹掉了 — not hidden by anyone, erased by the design.
HSK 5. It would have appeared in weeks under a different metric.
HSK 5. Implicit throughout: what the efficiency gains actually bought.
HSK 6. 那些效率数据本身并没有被推翻 — the numbers still stand.
问题从来不在数据造假 — the point is that honest numbers did this.
What an average measures, and why compression loses the shape.
The average lands in the trough — a position where almost no samples sit.
The tenth percentile would have surfaced the problem within weeks.
指标是绿的,体验是崩的 — the characteristic state this produces.
提问与指标之间的错配 — not bad data, the wrong question.
The explicit handoff rules they added when they split the traffic.
* beyond level超纲词
What grammar patterns appear in this episode?
语法换句话说,……
In other words. A hinge that restates a technical finding in operational terms — used here to turn a statistic into "no alarm went off".
换句话说,仪表盘上不存在任何异常信号。
如果当时……,……会……
Had they done X, Y would have happened. Counterfactual, used to isolate exactly which choice caused the outcome.
如果当时监测的是满意度分布、或者第十百分位数、或者投诉升级率,信号会在几周内出现。
这是两个不同的问题,用同一个……回答不了
Two different questions; one instrument cannot answer both. A compact closing formulation for a category error.
这是两个不同的问题,用同一个指标回答不了。
Proper Nouns
专有名词Sources
来源Free account
Keep learning from this story
Create a free account to keep saved words and your preferred level together.
- Keep words with their story context
- Remember your preferred level
- Build your vocabulary over time