A Chinese podcast · Same story, 4 levels

一个人工智能客服顶了七百个人,平均满意度和人一样:一年后公司把人招了回来,因为平均值把最糟的那批客户藏起来了
Klarna Rehires Humans After AI Replaces 700 Agents
About this story
Klarna launched an AI customer service agent in February 2024: 2.3 million conversations in month one, the work of about 700 agents, two minutes per query against eleven. By mid-2025 it was rehiring humans. Average satisfaction had looked level the whole time, because the distribution was bimodal. Chinese listening practice at four levels. HSK 5-6 Chinese listening practice.
This is an HSK 5-6 Chinese listening episode that runs about 5 minutes. The full Mandarin script is shown with tap-for-pinyin and a line-by-line English translation, so you can listen and read at once — comprehensible input in the sense of Stephen Krashen's i+1 theory. It teaches 12 key vocabulary words such as 效率、问题、公司 and walks through 3 grammar patterns, each explained in English with examples. The same news story is retold at 4 difficulty levels — use the level selector above to find the version that is challenging but still understandable for you.
Read at your level
原文Read the complete story in Chinese. Reveal pinyin and English only when you need them.
English transcript reference
This episode covers a complete cycle that ran from 2024 to 2025, and an old problem in statistics: what an average can answer, and what it cannot.
The facts first.
In February 2024, the Swedish payments company Klarna rolled out a large-model-based AI customer service system globally.
It handled 2.3 million conversations in its first month.
The company's stated conversion: a workload equivalent to about seven hundred full-time agents.
The operating metrics published alongside it included average resolution time falling from eleven minutes to two, repeat query rate down twenty-five percent, and projected additional profit for the year of about forty million dollars.
That data set was widely cited at the time and became the most frequently offered empirical case for the "AI replaces customer service" narrative.
In the same period, a billboard appeared on the streets of San Francisco carrying a single line: stop hiring humans.
On the public information available, the matter appeared to be settled.
Move forward to mid-2025, and the company began rehiring human agents.
The founder was strikingly candid in an interview with Bloomberg.
He said the company had "gone too far", and that an excessive focus on cost had driven quality down.
The easiest reading here is "the AI wasn't good enough".
That reading does not hold, because none of the efficiency data was ever overturned.
Two minutes against eleven was real; the fall in repeat queries was real.
The question actually worth analysing is: why did a drop in quality take more than a year to reach management's field of view?
The answer lies in the choice of monitoring metric.
The core metric the company tracked was the average of customer satisfaction.
And measured on the average, the AI and human agents performed about the same.
In other words, there was no anomalous signal anywhere on the dashboard.
The statistics need spelling out here.
An average is a measure of central tendency; its function is to compress a distribution into a single number.
Compression necessarily loses information, and what it loses is precisely the shape of the distribution.
When a data set is approximately normal, unimodal and not highly variable, the average represents the whole reasonably well.
When the data is bimodal, the average lands in the trough between the two peaks — a position where almost no samples sit.
This distribution was bimodal.
Highly standardised, low-context queries — order lookups, address changes, refund status — were handled quickly and accurately, and satisfaction for those users was close to the ceiling.
Queries requiring context, involving exceptions, or coming from users already unhappy were answered at the same two-minute speed, and answered wrong.
Satisfaction for those users was close to the floor.
More to the point is who those users were: they were typically the ones genuinely in need of help.
One end near the ceiling, one near the floor; weight them together and you land precisely on "level with human agents".
The result is a characteristic state: the metric is green and the experience is broken.
Had the company been monitoring the satisfaction distribution, or the tenth percentile, or the complaint escalation rate, the signal would have appeared within weeks.
But it was monitoring the average, so the signal was structurally erased.
The subsequent adjustment came in two parts.
Operationally, traffic was split: standardised queries handled by the model, complex ones routed to humans, with explicit handoff triggers defined.
On employment, the model became flexible — no fixed shifts, but a ride-hailing-style job-acceptance mechanism, with participants including students, people in rural areas, and the company's own users.
Finally, back to the numbers.
This needs stating clearly: not one of them was false.
Seven hundred, two minutes, twenty-five percent, forty million — every one can be checked, and every one holds.
"Average satisfaction roughly level" holds equally well.
The problem was never falsified data. It was a mismatch between the question and the metric.
What these metrics answer is "what happened on average".
What a service system actually needs to know is "how bad is the worst part, and who is in it".
Those are two different questions, and one metric cannot answer both.
Two questions to leave you with.
The first is concrete: the number you currently use to judge whether something is going normally — is it an average?
If so, what has it flattened out?
The second is bigger: under what conditions does an organisation go and look at the data it was never asked to look at?
Listen again
Try it without the transcript and notice what sounds clearer.
What vocabulary does this episode teach?
词汇这个解读站不住,因为那些效率数据本身并没有被推翻。 — That reading does not hold, because none of the efficiency data was ever overturned.
这期讲一个二〇二四到二〇二五年之间完成的完整循环,以及一个统计学上的老问题:平均值能回答什么,不能回答什么。 — This episode covers a complete cycle that ran from 2024 to 2025, and an old problem in statistics: what an average can answer, and what it cannot.
二〇二四年二月,瑞典支付公司克拉纳全球上线了一套基于大模型的人工智能客服系统。 — In February 2024, the Swedish payments company Klarna rolled out a large-model-based AI customer service system globally.
这组数据在当时被广泛引用,成为"人工智能替代客服"这一叙事最常被举出的实证案例。 — That data set was widely cited at the time and became the most frequently offered empirical case for the "AI replaces customer service" narrative.
二〇二四年二月,瑞典支付公司克拉纳全球上线了一套基于大模型的人工智能客服系统。 — In February 2024, the Swedish payments company Klarna rolled out a large-model-based AI customer service system globally.
同时公布的运营指标包括:平均问题解决时长从十一分钟降至两分钟;重复咨询率下降百分之二十五;预计当年增加利润约四千万美元。 — The operating metrics published alongside it included average resolution time falling from eleven minutes to two, repeat query rate down twenty-five percent, and projected additional profit for the year of about forty million dollars.
同时公布的运营指标包括:平均问题解决时长从十一分钟降至两分钟;重复咨询率下降百分之二十五;预计当年增加利润约四千万美元。 — The operating metrics published alongside it included average resolution time falling from eleven minutes to two, repeat query rate down twenty-five percent, and projected additional profit for the year of about forty million dollars.
但监测的是平均值,所以信号被结构性地抹掉了。 — But it was monitoring the average, so the signal was structurally erased.
换句话说,仪表盘上不存在任何异常信号。 — In other words, there was no anomalous signal anywhere on the dashboard.
问题从来不在数据造假,而在于提问与指标之间的错配。 — The problem was never falsified data. It was a mismatch between the question and the metric.
当数据呈现双峰分布时,平均值会落在两个峰之间的低谷处——一个几乎没有样本的位置。 — When the data is bimodal, the average lands in the trough between the two peaks — a position where almost no samples sit.
平均值是一个集中趋势指标,它的功能是把一个分布压缩成一个数。 — An average is a measure of central tendency; its function is to compress a distribution into a single number.
* beyond level超纲词
What grammar patterns appear in this episode?
语法换句话说,……
In other words. A hinge that restates a technical finding in operational terms — used here to turn a statistic into "no alarm went off".
换句话说,仪表盘上不存在任何异常信号。
如果当时……,……会……
Had they done X, Y would have happened. Counterfactual, used to isolate exactly which choice caused the outcome.
如果当时监测的是满意度分布、或者第十百分位数、或者投诉升级率,信号会在几周内出现。
这是两个不同的问题,用同一个……回答不了
Two different questions; one instrument cannot answer both. A compact closing formulation for a category error.
这是两个不同的问题,用同一个指标回答不了。
Proper nouns
专有名词Sources
来源Free account
Keep learning from this story
Create a free account to keep saved words and your preferred level together.
- Keep words with their story context
- Remember your preferred level
- Build your vocabulary over time