Fluentide

A Chinese podcast · Same story, 3 levels

One AI Did the Work of Seven Hundred Agents and Scored the Same on Average — a Year Later the Company Hired the Humans Back, Because the Average Was Hiding the Worst Customers — episode cover art
4:52

One AI Did the Work of Seven Hundred Agents and Scored the Same on Average — a Year Later the Company Hired the Humans Back, Because the Average Was Hiding the Worst Customers

About this story

Klarna launched an AI customer service agent in February 2024: 2.3 million conversations in month one, the work of about 700 agents, two minutes per query against eleven. By mid-2025 it was rehiring humans. Average satisfaction had looked level the whole time, because the distribution was bimodal. Chinese listening practice at four levels. HSK 5-6 Chinese listening practice.

This is an HSK 5-6 Chinese listening episode that runs about 5 minutes. The full Mandarin script is shown with tap-for-pinyin and a line-by-line English translation, so you can listen and read at once — comprehensible input in the sense of Stephen Krashen's i+1 theory. It teaches 12 key vocabulary words such as 指标、结构、信号 and walks through 3 grammar patterns, each explained in English with examples. The same news story is retold at 3 difficulty levels — use the level selector above to find the version that is challenging but still understandable for you.

Read at your level

原文

Read the complete story in Chinese. Reveal pinyin and English only when you need them.

一个五年之间完成完整循环以及一个统计问题平均值回答什么不能回答什么
事实
四年二月瑞典支付公司克拉全球上线一套基于模型人工智能系统
上线处理对话两百三十
公司对外换算工作相当七百全职
同时公布运营指标包括平均问题解决十一分钟降至两分钟重复咨询下降百分之二十五预计当年增加利润千万美元
数据当时广泛引用成为"人工智能替代"叙事最常实证案例
同期旧金山街头出现一块广告文案只有一句别再雇人
公开信息当时似乎已经有结论
时间推到五年公司开始重新招聘人工
公司创始接受采访表述相当坦率
表示公司"走得"过度聚焦成本导致质量下降
这里容易得出解读"人工智能不到要求"
这个解读站不住因为那些效率数据本身没有推翻
两分钟十一分钟真实重复咨询下降也是真实
真正值得分析问题质量下降为什么花了一年多才进入管理视野
答案监测指标选择
公司当时追踪核心指标客户满意平均值
平均值衡量人工智能人工表现基本持平
换句话说仪表不存在任何异常信号
这里需要统计学的部分说清楚
平均值一个集中趋势指标它的功能一个分布压缩一个
压缩必然损失信息损失恰恰分布形态
数据近似不大平均值较好代表整体
数据呈现双峰分布平均值落在之间一个几乎没有样本位置
这次分布正是双峰
标准程度依赖问题订单地址退款进度人工智能处理部分用户满意接近
需要理解上下文涉及例外情形或者用户情绪已经不佳问题人工智能同样两分钟速度出了回应答非所问
部分用户满意
关键的是部分用户构成他们往往真正需要帮助处境
一端接近一端接近零分平均之后恰好落在"人工持平"
于是出现一个典型状态指标绿体验
如果当时监测的是满意分布或者第十百分或者投诉升级信号出现
监测的是平均值所以信号结构抹掉
公司后来调整部分
业务分流标准问题模型处理复杂问题人工设置明确接触条件
弹性模式不再要求固定采用类似机制参与包括学生居住乡村地区人员以及公司自身用户
最后回到那些数字
需要说清楚没有一个数字假的
七百两分钟百分之二十五千万一个都可以核对一个成立
"平均满意基本持平"同样成立
问题从来不在数据造假在于提问指标之间
这些指标回答问题"平均而言发生什么"
一个服务系统真正需要知道的是"部分以及他们"
不同问题一个指标回答不了
问题
一个具体现在判断进展是否正常那个数字一个平均值
如果掉了什么
第二个一些一个组织什么条件下主动没有要求部分数据
English transcript reference

This episode covers a complete cycle that ran from 2024 to 2025, and an old problem in statistics: what an average can answer, and what it cannot.

The facts first.

In February 2024, the Swedish payments company Klarna rolled out a large-model-based AI customer service system globally.

It handled 2.3 million conversations in its first month.

The company's stated conversion: a workload equivalent to about seven hundred full-time agents.

The operating metrics published alongside it included average resolution time falling from eleven minutes to two, repeat query rate down twenty-five percent, and projected additional profit for the year of about forty million dollars.

That data set was widely cited at the time and became the most frequently offered empirical case for the "AI replaces customer service" narrative.

In the same period, a billboard appeared on the streets of San Francisco carrying a single line: stop hiring humans.

On the public information available, the matter appeared to be settled.

Move forward to mid-2025, and the company began rehiring human agents.

The founder was strikingly candid in an interview with Bloomberg.

He said the company had "gone too far", and that an excessive focus on cost had driven quality down.

The easiest reading here is "the AI wasn't good enough".

That reading does not hold, because none of the efficiency data was ever overturned.

Two minutes against eleven was real; the fall in repeat queries was real.

The question actually worth analysing is: why did a drop in quality take more than a year to reach management's field of view?

The answer lies in the choice of monitoring metric.

The core metric the company tracked was the average of customer satisfaction.

And measured on the average, the AI and human agents performed about the same.

In other words, there was no anomalous signal anywhere on the dashboard.

The statistics need spelling out here.

An average is a measure of central tendency; its function is to compress a distribution into a single number.

Compression necessarily loses information, and what it loses is precisely the shape of the distribution.

When a data set is approximately normal, unimodal and not highly variable, the average represents the whole reasonably well.

When the data is bimodal, the average lands in the trough between the two peaks — a position where almost no samples sit.

This distribution was bimodal.

Highly standardised, low-context queries — order lookups, address changes, refund status — were handled quickly and accurately, and satisfaction for those users was close to the ceiling.

Queries requiring context, involving exceptions, or coming from users already unhappy were answered at the same two-minute speed, and answered wrong.

Satisfaction for those users was close to the floor.

More to the point is who those users were: they were typically the ones genuinely in need of help.

One end near the ceiling, one near the floor; weight them together and you land precisely on "level with human agents".

The result is a characteristic state: the metric is green and the experience is broken.

Had the company been monitoring the satisfaction distribution, or the tenth percentile, or the complaint escalation rate, the signal would have appeared within weeks.

But it was monitoring the average, so the signal was structurally erased.

The subsequent adjustment came in two parts.

Operationally, traffic was split: standardised queries handled by the model, complex ones routed to humans, with explicit handoff triggers defined.

On employment, the model became flexible — no fixed shifts, but a ride-hailing-style job-acceptance mechanism, with participants including students, people in rural areas, and the company's own users.

Finally, back to the numbers.

This needs stating clearly: not one of them was false.

Seven hundred, two minutes, twenty-five percent, forty million — every one can be checked, and every one holds.

"Average satisfaction roughly level" holds equally well.

The problem was never falsified data. It was a mismatch between the question and the metric.

What these metrics answer is "what happened on average".

What a service system actually needs to know is "how bad is the worst part, and who is in it".

Those are two different questions, and one metric cannot answer both.

Two questions to leave you with.

The first is concrete: the number you currently use to judge whether something is going normally — is it an average?

If so, what has it flattened out?

The second is bigger: under what conditions does an organisation go and look at the data it was never asked to look at?

Listen again

Try it without the transcript and notice what sounds clearer.

What vocabulary does this episode teach?

词汇
zhǐbiāometric

HSK 5. 答案在监测指标的选择上 — the whole diagnosis in one line.

jiégòustructure

HSK 5. 信号被结构性地抹掉了 — not hidden by anyone, erased by the design.

xìnhàosignal

HSK 5. It would have appeared in weeks under a different metric.

zīyuánresources

HSK 5. Implicit throughout: what the efficiency gains actually bought.

xiàolǜefficiency

HSK 6. 那些效率数据本身并没有被推翻 — the numbers still stand.

zàojiǎto falsify

问题从来不在数据造假 — the point is that honest numbers did this.

jízhōng qūshìcentral tendency

What an average measures, and why compression loses the shape.

shuāngfēng fēnbùbimodal distribution

The average lands in the trough — a position where almost no samples sit.

bǎifēnwèishùpercentile

The tenth percentile would have surfaced the problem within weeks.

yíbiǎopándashboard

指标是绿的,体验是崩的 — the characteristic state this produces.

cuòpèimismatch

提问与指标之间的错配 — not bad data, the wrong question.

chùfā tiáojiàntrigger condition

The explicit handoff rules they added when they split the traffic.

* beyond level超纲词

What grammar patterns appear in this episode?

语法

换句话说,……

In other words. A hinge that restates a technical finding in operational terms — used here to turn a statistic into "no alarm went off".

换句话说,仪表盘上不存在任何异常信号。

如果当时……,……会……

Had they done X, Y would have happened. Counterfactual, used to isolate exactly which choice caused the outcome.

如果当时监测的是满意度分布、或者第十百分位数、或者投诉升级率,信号会在几周内出现。

这是两个不同的问题,用同一个……回答不了

Two different questions; one instrument cannot answer both. A compact closing formulation for a category error.

这是两个不同的问题,用同一个指标回答不了。

Proper Nouns

专有名词
瑞典RuìdiǎnSweden克拉纳KèlānàKlarna美国Měiguóthe United States旧金山JiùjīnshānSan Francisco彭博社PéngbóshèBloomberg

Sources

来源

Free account

Keep learning from this story

Create a free account to keep saved words and your preferred level together.

  • Keep words with their story context
  • Remember your preferred level
  • Build your vocabulary over time

Continue with Fluentide