Fluentide

A Chinese podcast · Same story, 3 levels

One AI Did the Work of Seven Hundred Agents and Scored the Same on Average — a Year Later the Company Hired the Humans Back, Because the Average Was Hiding the Worst Customers — episode cover art
3:46

One AI Did the Work of Seven Hundred Agents and Scored the Same on Average — a Year Later the Company Hired the Humans Back, Because the Average Was Hiding the Worst Customers

About this story

Klarna launched an AI customer service agent in February 2024: 2.3 million conversations in month one, the work of about 700 agents, two minutes per query against eleven. By mid-2025 it was rehiring humans. Average satisfaction had looked level the whole time, because the distribution was bimodal. Chinese listening practice at four levels. HSK 3 Chinese listening practice.

This is an HSK 3-4 Chinese listening episode that runs about 4 minutes. The full Mandarin script is shown with tap-for-pinyin and a line-by-line English translation, so you can listen and read at once — comprehensible input in the sense of Stephen Krashen's i+1 theory. It teaches 12 key vocabulary words such as 指标、分布、发现 and walks through 3 grammar patterns, each explained in English with examples. The same news story is retold at 3 difficulty levels — use the level selector above to find the version that is challenging but still understandable for you.

Read at your level

原文

Read the complete story in Chinese. Reveal pinyin and English only when you need them.

结论真正主角不是人工智能一个平均
四年二月瑞典支付公司克拉上线一个人工智能
上线处理两百三十对话
公司换算相当七百全职工作
同时公布还有几个指标
人工平均解决一个问题需要十一分钟人工智能需要两分钟
重复一个问题比例下降百分之二十五
公司预计当年因此增加千万美元利润
数字当时传播非常广
差不多同一时期旧金山街头出现一块广告上面只有别再雇人
数据似乎已经有答案
然后五年公司开始重新招聘人工
公司创始接受采访说法直接
我们走得
我们过于关注成本结果质量下降
这里一般"人工智能不行"
这个既不准确没意思
真正值得拆开质量下降为什么花了一年多才发现
公司当时监测核心指标客户满意
平均值人工智能人工表现基本持平
也就是说仪表没有任何异常
问题平均值这个工具本身
平均值压缩分布
差异极大数据在一起平均得到结果可能落在中间中间这个位置数据都不
这次分布大致这样
简单标准问题人工智能处理部分客户满意
复杂边缘需要判断上下文问题人工智能同样很快答非所问
遇到这类情况客户满意而且往往是需要帮助那批
一头很高一头很低平均正好落在"人工差不多"
所以仪表绿体验
公司看了一年别的渠道意识
现在做法分流标准问题交给机器复杂问题人工
方式改了不再是固定而是类似模式有空上线
学生住在乡下居民还有公司自己用户
最后那个数字
七百这个数字真的两分钟也是真的百分之二十五也是真的
它们全部成立而且都能核对
平均满意"基本持平"同样是真的
问题在于任何一个数字造假
问题在于这些数字回答的是"平均发生什么"公司需要知道的是"最糟情况"
问题一个指标回答不了
你想一想
有人一个平均告诉情况正常
相信之前你要追问的是一句
English transcript reference

The conclusion first: the real subject of this episode is not artificial intelligence. It is an average.

In February 2024, the Swedish payments company Klarna launched an AI customer service agent.

In its first month it handled 2.3 million conversations.

The company's own conversion: equivalent to seven hundred full-time agents.

Several other metrics were published at the same time.

A human agent took an average of eleven minutes to resolve a query; the AI took two.

Repeat queries about the same issue fell by twenty-five percent.

The company projected forty million dollars in additional profit that year.

That set of numbers travelled extremely widely at the time.

At roughly the same moment, a billboard went up on the streets of San Francisco with three words on it: stop hiring humans.

On the data, the question looked settled.

Then, in 2025, the company began rehiring human agents.

Its founder was very direct in an interview.

He said: we went too far.

He said: we focused too much on cost, and the result was lower quality.

At this point the usual telling is "AI isn't up to it".

But that telling is neither accurate nor interesting.

What actually deserves unpacking is why the drop in quality took more than a year to surface.

The core metric the company was monitoring was customer satisfaction.

On the average, the AI and the human agents performed about the same.

Which is to say: there was no anomaly anywhere on the dashboard.

The problem lies in the average itself as a tool.

An average compresses a distribution.

Take two groups of wildly different data, average them, and the result can land in the middle — a position where neither group actually is.

The distribution here looked roughly like this.

Simple, standard, high-frequency queries were handled by the AI quickly and accurately; those customers were extremely satisfied.

Complex, edge-case queries requiring context were answered just as fast, and answered wrong.

Customers in that group were extremely dissatisfied — and they were often the ones who most needed help.

One end high, one end low; average them and you land exactly on "about the same as humans".

So the dashboard was green and the experience was broken.

It took the company more than a year to learn this from other channels.

The current approach is to split the traffic: standard queries go to the machine, complex ones to a person.

The employment model changed too — no fixed shifts, but something closer to ride-hailing: you log on and take jobs when you are free.

The people taking those jobs include students, people living in the countryside, and the company's own users.

Finally, back to those numbers.

Seven hundred is real. Two minutes is real. Twenty-five percent is real.

Every one of them holds up, and every one can be checked.

"Average satisfaction was roughly level" is equally real.

The problem was never that a number was falsified.

The problem is that these numbers answer "what happened on average", and what the company needed to know was "how bad is the worst case".

Those are two questions, and one metric cannot answer both.

Think about it:

Someone shows you an average and tells you the situation is normal.

Before you believe it, which sentence do you press on?

Listen again

Try it without the transcript and notice what sounds clearer.

What vocabulary does this episode teach?

词汇
zhǐbiāometric, indicator

HSK 5. The choice of metric is the whole explanation for the year-long delay.

fēnbùdistribution

HSK 5. 平均值压缩了分布 — what an average throws away.

fāxiànto discover

HSK 3. 质量下降这件事为什么花了一年多才被发现。

pànduànjudgement

HSK 4. 需要判断上下文的问题 — exactly what the model could not do.

chéngběncost

HSK 6. 我们过于关注成本 — the founder's own account.

héduìto verify, check

它们全部成立,而且都能核对 — none of the numbers was false.

píngjūnzhíthe mean

压缩了分布 — and a compressed distribution is a hidden one.

mǎnyìdùsatisfaction score

The one number on the dashboard, and it stayed level all year.

yíbiǎopándashboard

仪表盘是绿的,体验是崩的 — the sentence the episode exists for.

kèfúcustomer service

Seven hundred of them replaced, then partially restored.

fēnliúto split traffic, triage

标准问题交给机器,复杂问题转人工 — the fix.

wǎngyuēchēride-hailing

The employment model they rehired under: log on when free, take jobs.

* beyond level超纲词

What grammar patterns appear in this episode?

语法

既不……,也没意思

Neither accurate nor interesting. Dismisses an easy reading on two grounds at once before offering a better one.

但这个讲法既不准确,也没意思。

于是出现了一个典型状态:……

And so you get a characteristic state. Names a recurring pattern rather than treating the case as unique.

于是出现了一个典型状态:指标是绿的,体验是崩的。

问题不在于……,而在于……

The problem is not X; it is Y. Redirects from where people look to where the fault actually is.

问题不在于任何一个数字造假。

问题在于,这些数字回答的是"平均发生了什么"。

Proper Nouns

专有名词
瑞典RuìdiǎnSweden克拉纳KèlānàKlarna美国Měiguóthe United States旧金山JiùjīnshānSan Francisco

Sources

来源

Free account

Keep learning from this story

Create a free account to keep saved words and your preferred level together.

  • Keep words with their story context
  • Remember your preferred level
  • Build your vocabulary over time

Continue with Fluentide