More than one industry

AI copilots started as developer tools but, between 2023 and 2025, became routine in law, consulting, customer service, manufacturing, and logistics. This article deliberately avoids anonymous anecdotes and estimates: every case below comes from a company's own announcement or a published study, and every figure carries its primary source.

#IndustryCasePublic evidence
1Software developmentGitHub Copilot RCT + METR studyTwo papers (arXiv)
2LegalAllen & Overy × HarveyOfficial announcement (2023)
3ConsultingBCG × Harvard field experimentHBS working paper (2023)
4Customer serviceKlarna AI assistant + NBER studyPress release + paper
5Manufacturing & logisticsSiemens × thyssenkrupp, DHLPress releases (2024)

1. Software development — the most measured, and the most split

The best-known controlled experiment is GitHub's 2022 study: 95 professional developers were asked to implement a JavaScript HTTP server, randomly assigned to use Copilot or not. The Copilot group finished 55.8% faster (1h 11m vs 2h 41m) with a higher completion rate (78% vs 70%), with the largest gains among less-experienced developers. (arXiv:2302.06590)

The opposite result also exists. In METR's 2025 randomized trial, 16 experienced open-source developers — averaging 5+ years on repositories of 1M+ lines — worked through 246 real issues. Tasks where AI tools (mostly Cursor + Claude) were allowed took 19% longer. The perception gap is the striking part: participants forecast a 24% speedup beforehand, and still believed they had been 20% faster afterward. (METR, 2025)

Read together, a pattern emerges: small, self-contained tasks plus lower experience yield the biggest gains; large familiar codebases plus deep expertise can shrink or reverse them.

2. Legal — Allen & Overy's firmwide Harvey rollout

In February 2023, global law firm Allen & Overy (now A&O Shearman) announced a firmwide deployment of Harvey, a legal AI platform built on OpenAI models. During the trial that began in November 2022, around 3,500 lawyers across 43 offices ran roughly 40,000 queries in day-to-day client work — contract drafting, case-law research, due diligence support. (A&O Shearman announcement)

The notable part is the policy: A&O made lawyer verification of Harvey's output mandatory. In an industry where the cost of error is high, the entry point was not replacement but draft generation plus human verification.

Related workflow guide: /tools/lawyer-ai-tools

3. Consulting — the BCG × Harvard "jagged frontier" experiment

In 2023, researchers from Harvard, Wharton, and MIT ran a field experiment with 758 BCG consultants on 18 realistic consulting tasks, randomly assigning GPT-4 access. The results cut both ways. (Dell'Acqua et al., HBS Working Paper 24-013)

  • Tasks inside the AI's capability frontier: +12.2% tasks completed, +25.1% speed, +40% quality
  • Tasks outside the frontier: consultants using AI were 19 percentage points less likely to reach the correct answer

The researchers called this the "jagged technological frontier": the boundary between what AI does well and badly is unintuitive, so using it outside the frontier produces confident errors. The core implication: knowing which tasks sit inside the frontier is itself a new job skill.

4. Customer service — Klarna's full rollout and partial retreat

In February 2024, fintech Klarna announced that its OpenAI-powered assistant handled 2.3 million conversations — two-thirds of all customer chats — in its first month, equivalent to the work of about 700 full-time agents. Resolution time fell from 11 minutes to under 2, repeat inquiries dropped 25%, and Klarna projected a $40M profit improvement for 2024. (Klarna press release, OpenAI case study)

The real lesson came later. In 2025, CEO Sebastian Siemiatkowski said publicly that the company had over-prioritized cost cutting and would rehire human agents, after quality gaps showed up in complex and emotionally loaded cases. The endpoint was not full automation but a redesigned AI-human division of labor.

The academic evidence points the same way. The NBER study (Brynjolfsson, Li & Raymond) analyzed 5,179 support agents and found the AI assistant raised issues resolved per hour by 14% on average — concentrated at 34% for novice agents, with minimal effect on the most skilled. (NBER w31161)

5. Manufacturing & logistics — code and documents are the entry point

In manufacturing, Siemens announced in November 2024 that thyssenkrupp Automation Engineering had adopted its Industrial Copilot, a generative AI assistant for industrial engineering, with global rollout from 2025. Uses: generating PLC control code (SCL), TIA Portal integration, automatic machine visualization — even in manufacturing, the entry point is a code draft. (Siemens press release)

In logistics, DHL Supply Chain announced in October 2024, with BCG X, generative AI applications for customer-data cleansing and initial analysis, logistics solution design support, query summarization, and legal-document processing. No quantitative figures were published; the stated core benefit was shorter lead time for solution-design proposals. (DHL Group press release)

Four patterns the public cases actually support

PatternEvidence
Gains concentrate among novicesNBER (+34% vs minimal for experts), GitHub RCT (largest gains for less-experienced devs)
Experts in complex domains can see reversalsMETR (-19%), BCG outside-frontier (-19pp)
Full replacement gets walked backKlarna's 2025 rehiring of human agents
Across industries, the entry point is "draft + human verification"A&O (contract drafts), Siemens (PLC code), DHL (proposals)

A checklist for applying this to your own role

The public cases yield three practical questions. These are questions, not data — the answers must be measured in your own work.

  1. Is my task inside or outside the frontier? Repetitive, self-contained, easily verified tasks (drafts, classification, summaries) tend to be inside; tasks needing deep domain context may be outside
  2. Have I priced in verification cost? METR's lesson: perceived speed and measured speed diverge. Measure before and after with the same metric
  3. Is the division of labor written down? Like A&O's mandatory lawyer verification, organizations that codify the AI-human line pay less to walk things back

Sources

  • Peng et al., The Impact of AI on Developer Productivity: Evidence from GitHub Copilot (2023) — arxiv.org/abs/2302.06590
  • METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (2025) — metr.org
  • Dell'Acqua et al., Navigating the Jagged Technological Frontier (HBS Working Paper 24-013, 2023) — aiinstitute.hbs.edu
  • Klarna, AI assistant handles two-thirds of customer service chats in its first month (2024) — klarna.com
  • Brynjolfsson, Li & Raymond, Generative AI at Work (NBER w31161, 2023; published in QJE 2025) — nber.org/papers/w31161
  • A&O Shearman, A&O announces exclusive launch partnership with Harvey (2023) — aoshearman.com
  • Siemens, Industrial Copilot expanded, adopted by thyssenkrupp (2024) — press.siemens.com
  • DHL Group, DHL Supply Chain implements Generative AI (2024) — group.dhl.com