Blog · Industry Signals

Klarna Replaced 700 Agents' Worth of Work, Then Rehired Humans

Mirai360 Team · Jul 20, 2026 · 8 min read

Klarna's AI assistant handled 2.3 million chats in a month, equal to ~700 agents' workload. By mid-2025 it was rehiring humans. The CEO's own words on why, and the triage pattern that fixed it.

What did Klarna actually do?

In February 2024, Klarna's AI customer service assistant handled 2.3 million chats in its first month. Klarna equated that workload to approximately 700 full-time human agents and reported roughly $40 million in annualized savings (Klarna Q1 2024 release, the company's own figure; OpenAI's published case study on this deployment is a vendor-adjacent source and should be read with that in mind). Klarna's own framing of the number was more measured than the headline that circulated: the company stated this was not 700 layoffs, but avoided incremental hiring it would otherwise have needed for that volume of chat traffic.

The deployment looked, on paper, like the fastest and cleanest agentic AI rollout on record. A single assistant absorbed a workload equivalent to hundreds of employees inside one month, with a savings figure large enough to headline a quarterly release.

Why did the reversal happen?

By mid-2025, Klarna was rehiring human agents. Chief Executive Officer Sebastian Siemiatkowski conceded that the company “focused too much on efficiency and cost… the result was lower quality, and that's not sustainable” (Forbes, May 18, 2025).

That statement names the actual failure directly. The AI assistant could handle chat volume; it handled 2.3 million chats in a month. The failure is that Klarna measured and reported volume and cost, but did not measure customer outcome quality with the same rigor before it scaled. Klarna judged the system a success on throughput and cost, then found the quality problem only after the scale-up was complete and customers had already experienced it.

This is a distinct failure mode from a model that does not work. The underlying assistant functioned; it processed chats and reduced cost exactly as reported. What broke was the sequence: Klarna scaled a customer-facing system before it measured quality against customer outcomes, then had to partially reverse the decision once the gap surfaced.

This pattern is not unique to Klarna. Separate research on generative AI deployments found that tools and organizations which do not retain feedback or adapt to workflows over time are the common thread behind projects that fail to show measurable business impact. Researchers labeled this mechanism a “learning gap” (MIT NANDA initiative, “The GenAI Divide,” August 2025, via Fortune, August 18, 2025). Klarna's assistant had the opposite of a learning problem on volume: it processed 2.3 million chats, and the company tracked that number closely. What it lacked, on the evidence of the CEO's own statement, was a customer-outcome measurement running at the same rigor as the cost measurement, from the start.

What does the reversal actually mean for a mid-market buyer?

A full reversal is rare and expensive: it means unwinding a public deployment, rehiring and retraining staff, and absorbing a quality gap that already reached customers before it was caught. For a company the size of Klarna, that is a visible, board-level correction. For a mid-market business without Klarna's balance sheet or its tolerance for a public misstep, the same sequence (deploy fast on a cost-and-volume metric, then discover a quality problem only after customers hit it) does not leave time to fix the story later. The cost lands directly on customer retention and support headcount, with no headline savings figure to offset it.

What is the triage pattern Klarna landed on?

The configuration Klarna settled into after the correction is triage: the AI assistant handles simple, routine queries, and human agents handle disputes, fraud, and hardship cases (CX Dive / Forbes, mid-2025). That split does not retreat from AI so much as match automation to risk, because not every support ticket carries the same danger if the automated answer is wrong. A question about a delivery date is low-stakes if the AI gets it wrong once. A dispute, a fraud claim, or a hardship request is high-stakes: a wrong or tone-deaf automated response there does measurable damage to the relationship, and in some cases to regulatory exposure.

The lesson generalizes past customer support. Any workflow with a mix of routine, low-risk cases and rare, high-stakes exceptions is a candidate for the same split: automate the volume, route the exceptions to a human, and use the volume data to know which cases are actually routine before deciding that.

How does Mirai360 avoid the deploy-then-reverse cycle?

The Klarna sequence is avoidable, and the fix is procedural rather than technical. Four disciplines, applied in order, are what keep a deployment from needing a public correction a year later.

Why start with a narrow workflow instead of a broad rollout?

A deployment scoped to one well-defined workflow (quote generation, invoice matching, a single support-ticket category) produces a result that can be measured and compared against a baseline within weeks. A broad rollout across an entire support queue, as Klarna's assistant effectively became within its first month, produces a large number fast, but that number describes volume and cost, not whether individual customer outcomes held up. Scope determines what gets measured; a narrow scope is what makes a real quality check possible before scale.

Why set a signed baseline before measuring anything?

Before a workflow goes live, Mirai360 agrees, in writing and signed off by the business owner, on what "working" means: an error rate ceiling, a customer satisfaction floor, a cost-per-resolution target. Mirai360 sets that baseline before the deployment begins, not afterward to justify a decision already made. Klarna could report its $40 million savings figure because it measured chat volume and agent-equivalent workload from the start (Klarna Q1 2024 release). The gap: it did not hold quality against customer outcomes to the same standard at the same time.

Why keep humans on exceptions rather than full automation?

Mirai360 deploys agents on the routine, high-volume share of a workflow and keeps a human in the loop on disputes, exceptions, and anything with regulatory or relationship risk. This is the same split Klarna adopted after its correction, applied from the first deployment rather than after a public reversal. The agent's own output on flagged or low-confidence cases routes to a person before it reaches a customer.

Why scale only against pre-set numbers?

Mirai360 makes scale-up decisions against the signed baseline, not against a volume or cost number in isolation. If the error rate or customer satisfaction floor set at the outset holds after a defined review period, the workflow scales. If it does not, Mirai360 fixes or narrows the workflow before it goes further. That decision happens on schedule, not after a quality problem has already reached enough customers to force a public correction.

This is the same operating discipline behind Mirai360's platform: an LLM gateway for cost and routing control, evals and guardrails to check output quality against a standard before and after launch, and a deployment path (self-hosted, managed, or custom-built) chosen to match how much control a given workflow's risk profile requires.

What should a business ask before it scales its own AI agent?

A team that cannot answer these before scaling is repeating the sequence Klarna had to reverse: a workload measured accurately, a quality gap measured too late.

Ready to deploy without a reversal?

Mirai360 AI builds agentic AI deployments scoped to one workflow, measured against a signed baseline, with humans kept on exceptions from day one. Book a free 30-minute call to talk through where that applies to your business.

FAQ

What did Klarna actually do?
Klarna's AI assistant handled 2.3 million chats in its first month (Feb 2024), work the company equated to ~700 full-time agents, reporting ~$40M in annualized savings (Klarna Q1 2024 release).
Why did the reversal happen?
Klarna's CEO said the company focused too much on efficiency and cost, and quality suffered. By mid-2025 it was rehiring human agents (Forbes, May 18, 2025).
What is the triage pattern Klarna landed on?
AI handles routine, simple queries; humans handle disputes, fraud, and hardship cases.
How does Mirai360 avoid the deploy-then-reverse cycle?
Narrow workflow scope, a signed quality baseline before launch, humans kept on exceptions, and scale-up decisions made only against pre-set numbers.

Ready to put agents to work?

Tell us how your business runs today. We will show you which path — self-hosted, managed, or custom-built — gets you to production fastest.

Talk to us