Editorial note: this brief collates public numbers from primary sources. Company-published deployment metrics are labeled as company-published evidence. They are useful signals, but they are not independent audits or universal benchmarks.
Evidence snapshot
| Source | Reported number | Operational reading | Caveat |
|---|---|---|---|
| Microsoft Work Trend Index 2025 | 82% of leaders say they are confident they will use digital labor to expand workforce capacity in the next 12-18 months. Microsoft also reports 45% of leaders considering digital labor for team-capacity expansion as a top workforce strategy. | Digital labor is moving into workforce planning, not just experimentation. | Survey expectation; not proof of implemented ROI. |
| IBM CEO study on AI ROI and enterprise scaling | 61% of surveyed CEOs say they are actively adopting AI agents; only 25% of AI initiatives have delivered expected ROI; 16% have scaled enterprise-wide; 50% say recent investment created disconnected technology. | Agent interest is high, but scaling and ROI remain constrained by data, architecture, and operating model. | CEO survey; not a workflow-level benchmark. |
| Deloitte State of AI in the Enterprise 2026 | Worker access to AI rose 50% in 2025; only one in five companies reports a mature governance model for autonomous AI agents; agentic AI is expected to have the highest impact in customer support. | Enterprise AI access is expanding faster than governance, with customer support the clearest agentic entry point. | Enterprise survey and executive expectations; not independent deployment measurement. |
| OECD: Generative AI and the SME Workforce | The OECD surveyed more than 5,000 SMEs and states that SMEs account for over 99% of companies and 60% of business-sector employment in OECD economies. | SME adoption matters because smaller firms carry a large share of employment and face labor shortages and skill gaps. | Cross-country survey; effects vary by firm, sector, and adoption quality. |
| World Economic Forum Future of Jobs Report 2025 | 63% of employers identify skill gaps as a major barrier to business transformation over 2025-2030; 85% plan to prioritize workforce upskilling. | Skill pressure is a primary reason companies look for AI-assisted capacity and role redesign. | Employer survey; not evidence that AI solves the gap by itself. |
| NBER: Generative AI at Work | Access to a generative AI assistant increased customer-support issues resolved per hour by 14% on average, with larger gains for novice and low-skilled workers. | The strongest public productivity evidence is in customer support and agent-assist workflows. | Single task environment; results should not be generalized without internal baselines. |
| Klarna AI assistant launch results | Klarna reports 2.3 million AI-assistant conversations in the first month, covering two-thirds of customer-service chats and doing work equivalent to 700 full-time agents. | Company-published evidence that high-volume customer operations can shift meaningful workload to AI. | Company-published case study; not an independent audit. |
| PwC 2025 Global AI Jobs Barometer | Industries more exposed to AI show 3x higher growth in revenue per employee; PwC also reports 66% faster skill change in AI-exposed jobs. | AI exposure correlates with productivity and skill-change pressure at industry level. | Macro correlation; not proof that a specific operations deployment will produce the same result. |
What the numbers support
- Capacity expansion is the clearest executive framing: Microsoft and IBM both show leaders planning around digital labor and AI agents.
- Workflow redesign matters more than adding a chatbot to the old process: Deloitte separates surface-level AI use from process redesign and business reimagination.
- Service operations are the most evidenced entry point: NBER, Klarna, and Deloitte all point toward customer support or service work as a high-signal area.
- Skill-gap pressure is real: WEF and OECD both frame skills and labor constraints as part of the adoption context.
What the numbers do not prove
- They do not prove universal ROI. IBM's CEO study is a warning that most AI initiatives still miss expected ROI or fail to scale enterprise-wide.
- They do not prove blanket headcount replacement. Microsoft and Deloitte both describe human-led or human-reviewed operating models, not unmanaged automation.
- They do not turn vendor or company case studies into independent benchmarks. Klarna's numbers are useful because they are public and specific, but they remain company-published.
- They do not replace internal baselines. A support metric from one company or a macro PwC industry correlation cannot be copied into another firm's business case.
Function-by-function status
| Function | Public evidence status | Measure first |
|---|---|---|
| Customer support | Strongest public evidence today. NBER measures productivity in customer support, Klarna publishes high-volume customer-service deployment numbers, and Deloitte identifies customer support as the highest-impact agentic AI area. | Containment, repeat inquiries, average handle time, first response time, SLA attainment, escalation quality, CSAT. |
| HR and employee service | Promising, but less broadly evidenced in the source set used here. Treat as an internal-baseline opportunity rather than a market-proven benchmark. | Ticket deflection, employee effort, policy-answer accuracy, resolution time, escalation rate. |
| Finance and back office | Promising where work is structured: reconciliation, document handling, close support, exception routing. Public independent evidence is thinner than in support. | Cycle time, first-pass yield, rework, exception backlog, approval latency, unit cost. |
| Procurement and supply chain | Emerging use case category. Public numbers are less consistent, so pilots should start with bounded workflows and clear review points. | Supplier review time, purchase-order exception rate, cycle time, backlog, policy compliance. |
| Cross-functional operations | Useful where work crosses inboxes, CRM, ticketing, documents, and internal systems. Evidence should be built from the company's own baseline. | Throughput, handoff latency, stale cases, manual follow-up rate, SLA misses. |
Measurement checklist
| Metric | Why it matters |
|---|---|
| Cycle time | Shows whether the end-to-end workflow is actually faster. |
| Throughput | Shows whether the team handles more work with the same human capacity. |
| Unit cost | Connects operational volume to financial impact. |
| Backlog | Shows whether AI is reducing accumulated work or only improving visible response speed. |
| SLA attainment | Tests whether service commitments improve. |
| Error and rework rate | Protects against speed gains that create downstream cleanup. |
| Containment or deflection | Measures work completed without human handling, but must be paired with quality checks. |
| Repeat inquiries | Catches false automation where the same issue returns through another channel. |
Implementation reality
The evidence supports a narrow implementation pattern: pick a measurable workflow, redesign routing and ownership, connect the agent to approved systems, restrict permissions, keep human review for exceptions, and instrument the workflow before expanding it.
- Workflow redesign: define what the AI owns, what it drafts, what it routes, and what remains human-reviewed.
- Integrations: connect CRM, ticketing, HRIS, ERP, knowledge base, document store, or inbox systems only where the workflow requires them.
- Governance: apply least-privilege access, audit logs, approved knowledge sources, escalation thresholds, and rollback paths.
- Measurement: compare against the pre-AI baseline before claiming ROI.
Source limitations
Updated 2026-05-26. Several visible deployment numbers come from vendor or company-published material. This brief uses them as directional evidence and labels the caveat beside the number. Claims from the research notes that could not be traced to a primary source were excluded.