Fluentide

A Chinese podcast · Same story, 4 levels

Klarna Rehires Humans After AI Replaces 700 Agents — episode cover art
4:52

一个人工智能客服顶了七百个人,平均满意度和人一样:一年后公司把人招了回来,因为平均值把最糟的那批客户藏起来了

Klarna Rehires Humans After AI Replaces 700 Agents

About this story

Klarna launched an AI customer service agent in February 2024: 2.3 million conversations in month one, the work of about 700 agents, two minutes per query against eleven. By mid-2025 it was rehiring humans. Average satisfaction had looked level the whole time, because the distribution was bimodal. Chinese listening practice at four levels. HSK 5-6 Chinese listening practice.

This is an HSK 5-6 Chinese listening episode that runs about 5 minutes. The full Mandarin script is shown with tap-for-pinyin and a line-by-line English translation, so you can listen and read at once — comprehensible input in the sense of Stephen Krashen's i+1 theory. It teaches 12 key vocabulary words such as 效率、问题、公司 and walks through 3 grammar patterns, each explained in English with examples. The same news story is retold at 4 difficulty levels — use the level selector above to find the version that is challenging but still understandable for you.

Read at your level

原文

Read the complete story in Chinese. Reveal pinyin and English only when you need them.

这期讲一个二〇二四到二〇二五年之间完成的完整循环,以及一个统计学上的老问题:平均值能回答什么,不能回答什么。
先给事实。
二〇二四年二月,瑞典支付公司克拉纳全球上线了一套基于大模型的人工智能客服系统。
上线首月处理对话两百三十万次。
公司对外给出的换算是:工作量相当于约七百名全职客服。
同时公布的运营指标包括:平均问题解决时长从十一分钟降至两分钟;重复咨询率下降百分之二十五;预计当年增加利润约四千万美元。
这组数据在当时被广泛引用,成为"人工智能替代客服"这一叙事最常被举出的实证案例。
同期,旧金山街头出现了一块广告牌,文案只有一句话:别再雇人。
从公开信息看,这件事在当时似乎已经有了结论。
时间推到二〇二五年年中,公司开始重新招聘人工客服。
公司创始人在接受彭博社采访时的表述相当坦率。
他表示公司"走得太远了",称过度聚焦成本导致质量下降。
到这里,最容易得出的解读是"人工智能达不到要求"。
这个解读站不住,因为那些效率数据本身并没有被推翻。
两分钟对十一分钟是真实的,重复咨询率下降也是真实的。
真正值得分析的问题是:质量下降为什么花了一年多才进入管理层视野。
答案在监测指标的选择上。
公司当时追踪的核心指标是客户满意度的平均值。
而按平均值衡量,人工智能与人工客服的表现基本持平。
换句话说,仪表盘上不存在任何异常信号。
这里需要把统计学的部分说清楚。
平均值是一个集中趋势指标,它的功能是把一个分布压缩成一个数。
压缩必然损失信息,而损失掉的恰恰是分布形态。
当一组数据近似正态、单峰、方差不大时,平均值能较好地代表整体。
当数据呈现双峰分布时,平均值会落在两个峰之间的低谷处——一个几乎没有样本的位置。
这次的分布正是双峰的。
标准化程度高、语境依赖低的问题——查订单、改地址、退款进度——人工智能处理得又快又准,这部分用户满意度接近满分。
而需要理解上下文、涉及例外情形、或者用户情绪已经不佳的问题,人工智能同样以两分钟的速度给出了回应,但答非所问。
这部分用户的满意度极低。
更关键的是这部分用户的构成:他们往往正处在真正需要帮助的处境里。
一端接近满分,一端接近零分,加权平均之后,恰好落在"与人工客服持平"。
于是出现了一个典型状态:指标是绿的,体验是崩的。
如果当时监测的是满意度分布、或者第十百分位数、或者投诉升级率,信号会在几周内出现。
但监测的是平均值,所以信号被结构性地抹掉了。
公司后来的调整分两部分。
业务上改为分流:标准化问题由模型处理,复杂问题转人工,并设置了明确的转接触发条件。
用工上改为弹性模式,不再要求固定坐班,采用类似网约车的接单机制,参与者包括学生、居住在乡村地区的人员,以及公司自身的用户。
最后回到那些数字。
需要说清楚:没有一个数字是假的。
七百、两分钟、百分之二十五、四千万,每一个都可以核对,每一个都成立。
"平均满意度基本持平"这句话同样成立。
问题从来不在数据造假,而在于提问与指标之间的错配。
这些指标回答的问题是"平均而言发生了什么"。
而一个服务系统真正需要知道的是"最差的那部分有多差,以及他们是谁"。
这是两个不同的问题,用同一个指标回答不了。
留两个问题给你。
第一个具体些:你现在用来判断某件事进展是否正常的那个数字,是一个平均值吗?
如果是,它压掉了什么?
第二个大一些:一个组织在什么条件下,才会主动去看它没有被要求看的那部分数据?
English transcript reference

This episode covers a complete cycle that ran from 2024 to 2025, and an old problem in statistics: what an average can answer, and what it cannot.

The facts first.

In February 2024, the Swedish payments company Klarna rolled out a large-model-based AI customer service system globally.

It handled 2.3 million conversations in its first month.

The company's stated conversion: a workload equivalent to about seven hundred full-time agents.

The operating metrics published alongside it included average resolution time falling from eleven minutes to two, repeat query rate down twenty-five percent, and projected additional profit for the year of about forty million dollars.

That data set was widely cited at the time and became the most frequently offered empirical case for the "AI replaces customer service" narrative.

In the same period, a billboard appeared on the streets of San Francisco carrying a single line: stop hiring humans.

On the public information available, the matter appeared to be settled.

Move forward to mid-2025, and the company began rehiring human agents.

The founder was strikingly candid in an interview with Bloomberg.

He said the company had "gone too far", and that an excessive focus on cost had driven quality down.

The easiest reading here is "the AI wasn't good enough".

That reading does not hold, because none of the efficiency data was ever overturned.

Two minutes against eleven was real; the fall in repeat queries was real.

The question actually worth analysing is: why did a drop in quality take more than a year to reach management's field of view?

The answer lies in the choice of monitoring metric.

The core metric the company tracked was the average of customer satisfaction.

And measured on the average, the AI and human agents performed about the same.

In other words, there was no anomalous signal anywhere on the dashboard.

The statistics need spelling out here.

An average is a measure of central tendency; its function is to compress a distribution into a single number.

Compression necessarily loses information, and what it loses is precisely the shape of the distribution.

When a data set is approximately normal, unimodal and not highly variable, the average represents the whole reasonably well.

When the data is bimodal, the average lands in the trough between the two peaks — a position where almost no samples sit.

This distribution was bimodal.

Highly standardised, low-context queries — order lookups, address changes, refund status — were handled quickly and accurately, and satisfaction for those users was close to the ceiling.

Queries requiring context, involving exceptions, or coming from users already unhappy were answered at the same two-minute speed, and answered wrong.

Satisfaction for those users was close to the floor.

More to the point is who those users were: they were typically the ones genuinely in need of help.

One end near the ceiling, one near the floor; weight them together and you land precisely on "level with human agents".

The result is a characteristic state: the metric is green and the experience is broken.

Had the company been monitoring the satisfaction distribution, or the tenth percentile, or the complaint escalation rate, the signal would have appeared within weeks.

But it was monitoring the average, so the signal was structurally erased.

The subsequent adjustment came in two parts.

Operationally, traffic was split: standardised queries handled by the model, complex ones routed to humans, with explicit handoff triggers defined.

On employment, the model became flexible — no fixed shifts, but a ride-hailing-style job-acceptance mechanism, with participants including students, people in rural areas, and the company's own users.

Finally, back to the numbers.

This needs stating clearly: not one of them was false.

Seven hundred, two minutes, twenty-five percent, forty million — every one can be checked, and every one holds.

"Average satisfaction roughly level" holds equally well.

The problem was never falsified data. It was a mismatch between the question and the metric.

What these metrics answer is "what happened on average".

What a service system actually needs to know is "how bad is the worst part, and who is in it".

Those are two different questions, and one metric cannot answer both.

Two questions to leave you with.

The first is concrete: the number you currently use to judge whether something is going normally — is it an average?

If so, what has it flattened out?

The second is bigger: under what conditions does an organisation go and look at the data it was never asked to look at?

Listen again

Try it without the transcript and notice what sounds clearer.

What vocabulary does this episode teach?

词汇
xiàolǜefficiency

这个解读站不住,因为那些效率数据本身并没有被推翻。 — That reading does not hold, because none of the efficiency data was ever overturned.

wèntíquestion; problem

这期讲一个二〇二四到二〇二五年之间完成的完整循环,以及一个统计学上的老问题:平均值能回答什么,不能回答什么。 — This episode covers a complete cycle that ran from 2024 to 2025, and an old problem in statistics: what an average can answer, and what it cannot.

gōngsīcompany

二〇二四年二月,瑞典支付公司克拉纳全球上线了一套基于大模型的人工智能客服系统。 — In February 2024, the Swedish payments company Klarna rolled out a large-model-based AI customer service system globally.

shùjùdata

这组数据在当时被广泛引用,成为"人工智能替代客服"这一叙事最常被举出的实证案例。 — That data set was widely cited at the time and became the most frequently offered empirical case for the "AI replaces customer service" narrative.

zhìnéngintelligence (as in 人工智能)

二〇二四年二月,瑞典支付公司克拉纳全球上线了一套基于大模型的人工智能客服系统。 — In February 2024, the Swedish payments company Klarna rolled out a large-model-based AI customer service system globally.

fēnzhōngminute

同时公布的运营指标包括:平均问题解决时长从十一分钟降至两分钟;重复咨询率下降百分之二十五;预计当年增加利润约四千万美元。 — The operating metrics published alongside it included average resolution time falling from eleven minutes to two, repeat query rate down twenty-five percent, and projected additional profit for the year of about forty million dollars.

zhǐbiāometric; indicator

同时公布的运营指标包括:平均问题解决时长从十一分钟降至两分钟;重复咨询率下降百分之二十五;预计当年增加利润约四千万美元。 — The operating metrics published alongside it included average resolution time falling from eleven minutes to two, repeat query rate down twenty-five percent, and projected additional profit for the year of about forty million dollars.

jiégòustructure

但监测的是平均值,所以信号被结构性地抹掉了。 — But it was monitoring the average, so the signal was structurally erased.

xìnhàosignal

换句话说,仪表盘上不存在任何异常信号。 — In other words, there was no anomalous signal anywhere on the dashboard.

zàojiǎto fake; to falsify

问题从来不在数据造假,而在于提问与指标之间的错配。 — The problem was never falsified data. It was a mismatch between the question and the metric.

shuāngfēng fēnbùa two-peaked distribution

当数据呈现双峰分布时,平均值会落在两个峰之间的低谷处——一个几乎没有样本的位置。 — When the data is bimodal, the average lands in the trough between the two peaks — a position where almost no samples sit.

jízhōng qūshìcentral tendency

平均值是一个集中趋势指标,它的功能是把一个分布压缩成一个数。 — An average is a measure of central tendency; its function is to compress a distribution into a single number.

* beyond level超纲词

What grammar patterns appear in this episode?

语法

换句话说,……

In other words. A hinge that restates a technical finding in operational terms — used here to turn a statistic into "no alarm went off".

换句话说,仪表盘上不存在任何异常信号。

如果当时……,……会……

Had they done X, Y would have happened. Counterfactual, used to isolate exactly which choice caused the outcome.

如果当时监测的是满意度分布、或者第十百分位数、或者投诉升级率,信号会在几周内出现。

这是两个不同的问题,用同一个……回答不了

Two different questions; one instrument cannot answer both. A compact closing formulation for a category error.

这是两个不同的问题,用同一个指标回答不了。

Proper nouns

专有名词
瑞典RuìdiǎnSweden克拉纳KèlānàKlarna美国Měiguóthe United States旧金山JiùjīnshānSan Francisco彭博社PéngbóshèBloomberg

Sources

来源

Free account

Keep learning from this story

Create a free account to keep saved words and your preferred level together.

  • Keep words with their story context
  • Remember your preferred level
  • Build your vocabulary over time

Keep learning