A Chinese podcast · Same story, 3 levels

一个人工智能客服顶了七百个人,平均满意度和人一样:一年后公司把人招了回来,因为平均值把最糟的那批客户藏起来了
One AI Did the Work of Seven Hundred Agents and Scored the Same on Average — a Year Later the Company Hired the Humans Back, Because the Average Was Hiding the Worst Customers
About this story
Klarna launched an AI customer service agent in February 2024: 2.3 million conversations in month one, the work of about 700 agents, two minutes per query against eleven. By mid-2025 it was rehiring humans. Average satisfaction had looked level the whole time, because the distribution was bimodal. Chinese listening practice at four levels. HSK 3 Chinese listening practice.
This is an HSK 3-4 Chinese listening episode that runs about 4 minutes. The full Mandarin script is shown with tap-for-pinyin and a line-by-line English translation, so you can listen and read at once — comprehensible input in the sense of Stephen Krashen's i+1 theory. It teaches 12 key vocabulary words such as 指标、分布、发现 and walks through 3 grammar patterns, each explained in English with examples. The same news story is retold at 3 difficulty levels — use the level selector above to find the version that is challenging but still understandable for you.
Read at your level
原文Read the complete story in Chinese. Reveal pinyin and English only when you need them.
English transcript reference
The conclusion first: the real subject of this episode is not artificial intelligence. It is an average.
In February 2024, the Swedish payments company Klarna launched an AI customer service agent.
In its first month it handled 2.3 million conversations.
The company's own conversion: equivalent to seven hundred full-time agents.
Several other metrics were published at the same time.
A human agent took an average of eleven minutes to resolve a query; the AI took two.
Repeat queries about the same issue fell by twenty-five percent.
The company projected forty million dollars in additional profit that year.
That set of numbers travelled extremely widely at the time.
At roughly the same moment, a billboard went up on the streets of San Francisco with three words on it: stop hiring humans.
On the data, the question looked settled.
Then, in 2025, the company began rehiring human agents.
Its founder was very direct in an interview.
He said: we went too far.
He said: we focused too much on cost, and the result was lower quality.
At this point the usual telling is "AI isn't up to it".
But that telling is neither accurate nor interesting.
What actually deserves unpacking is why the drop in quality took more than a year to surface.
The core metric the company was monitoring was customer satisfaction.
On the average, the AI and the human agents performed about the same.
Which is to say: there was no anomaly anywhere on the dashboard.
The problem lies in the average itself as a tool.
An average compresses a distribution.
Take two groups of wildly different data, average them, and the result can land in the middle — a position where neither group actually is.
The distribution here looked roughly like this.
Simple, standard, high-frequency queries were handled by the AI quickly and accurately; those customers were extremely satisfied.
Complex, edge-case queries requiring context were answered just as fast, and answered wrong.
Customers in that group were extremely dissatisfied — and they were often the ones who most needed help.
One end high, one end low; average them and you land exactly on "about the same as humans".
So the dashboard was green and the experience was broken.
It took the company more than a year to learn this from other channels.
The current approach is to split the traffic: standard queries go to the machine, complex ones to a person.
The employment model changed too — no fixed shifts, but something closer to ride-hailing: you log on and take jobs when you are free.
The people taking those jobs include students, people living in the countryside, and the company's own users.
Finally, back to those numbers.
Seven hundred is real. Two minutes is real. Twenty-five percent is real.
Every one of them holds up, and every one can be checked.
"Average satisfaction was roughly level" is equally real.
The problem was never that a number was falsified.
The problem is that these numbers answer "what happened on average", and what the company needed to know was "how bad is the worst case".
Those are two questions, and one metric cannot answer both.
Think about it:
Someone shows you an average and tells you the situation is normal.
Before you believe it, which sentence do you press on?
Listen again
Try it without the transcript and notice what sounds clearer.
What vocabulary does this episode teach?
词汇HSK 5. The choice of metric is the whole explanation for the year-long delay.
HSK 5. 平均值压缩了分布 — what an average throws away.
HSK 3. 质量下降这件事为什么花了一年多才被发现。
HSK 4. 需要判断上下文的问题 — exactly what the model could not do.
HSK 6. 我们过于关注成本 — the founder's own account.
它们全部成立,而且都能核对 — none of the numbers was false.
压缩了分布 — and a compressed distribution is a hidden one.
The one number on the dashboard, and it stayed level all year.
仪表盘是绿的,体验是崩的 — the sentence the episode exists for.
Seven hundred of them replaced, then partially restored.
标准问题交给机器,复杂问题转人工 — the fix.
The employment model they rehired under: log on when free, take jobs.
* beyond level超纲词
What grammar patterns appear in this episode?
语法既不……,也没意思
Neither accurate nor interesting. Dismisses an easy reading on two grounds at once before offering a better one.
但这个讲法既不准确,也没意思。
于是出现了一个典型状态:……
And so you get a characteristic state. Names a recurring pattern rather than treating the case as unique.
于是出现了一个典型状态:指标是绿的,体验是崩的。
问题不在于……,而在于……
The problem is not X; it is Y. Redirects from where people look to where the fault actually is.
问题不在于任何一个数字造假。
问题在于,这些数字回答的是"平均发生了什么"。
Proper Nouns
专有名词Sources
来源Free account
Keep learning from this story
Create a free account to keep saved words and your preferred level together.
- Keep words with their story context
- Remember your preferred level
- Build your vocabulary over time