
In 2023 we published an article about large language models that promised transformed communications, faster development, and global scale. Three years later, that promise has a measurable track record — so this is the honest follow-up. The question is no longer what AI LLMs for business could do. Instead, it is what they demonstrably have done, according to regulators, randomised trials, and court records rather than vendor decks.
Adoption is far lower than the conversation suggests
Start with the least glamorous number. US Census Bureau research covering November 2025 to January 2026 found that only 18% of US firms used AI in a business function. Official Census surveys put overall US business AI use between 17% and 20% through May 2026. Among firms with 250 or more employees it reaches 37%. Among small firms it stays far lower.
Europe looks similar. Eurostat recorded 19.95% of EU enterprises using AI technologies in 2025, rising to 55% among large enterprises. Closer to home, Singapore’s IMDA reported SME adoption tripling from 4.2% to 14.5%, while larger firms moved from 44.0% to 62.5%.
Consequently, if your organisation has not deployed an LLM yet, you are not behind. You are in the majority.
The two failure statistics everyone quotes are both wrong
A no-hype article has to correct the anti-hype hype as well. Two numbers dominate the “AI is failing” conversation, and neither says what people think.
“95% of AI pilots fail.” The MIT NANDA report actually said 95% of organisations are getting zero return. That is a different claim. Furthermore, it rests on 52 structured interviews rather than a large survey. Its defensible core is a funnel. For custom and vendor-sold enterprise GenAI systems, 60% of organisations evaluated them, 20% reached pilot, and just 5% reached production.
“RAND found 80% of AI projects fail.” RAND did not measure that. The report wrote “by some estimates” and cited others. However, RAND’s own finding is more useful anyway. RAND interviewed 65 experienced practitioners, 50 of them in industry. Among those 50, 84% named leadership-driven failures: wrong problem, unclear success criteria, no data. Model quality was not the culprit.
The gap between feeling faster and being faster
The most important finding of the past two years is not about capability. Rather, it is about measurement.
In a randomised controlled trial run in early 2025, METR had 16 experienced open-source developers use AI tools on real tasks. They took 19% longer. Beforehand, they had predicted the tools would make them 24% faster. They still believed they had been sped up afterwards. Notably, METR later surveyed 349 technical workers. The median respondent reported a 3× change in raw speed. Yet the same respondent reported only a 1.4–2× change in output value.
Meanwhile, DORA’s 2025 report surveyed nearly 5,000 professionals. It found 90% using AI at work, and over 80% believing it raised their productivity. Yet 30% reported little or no trust in AI-generated code. Self-reported productivity, in short, is not evidence. Therefore treat any vendor number derived from a satisfaction survey as marketing.
Where AI LLMs for business genuinely earn their keep
The measured wins are real, and they share a shape. In a study of 5,172 customer-support agents, a generative AI assistant raised issues resolved per hour by 15%. However, the gains concentrated heavily among less-experienced staff. The most skilled saw little effect.
Similarly, a field experiment put 758 BCG consultants on real tasks. Those using GPT-4 completed 12.2% more tasks about 25% faster inside the model’s capability frontier. Outside it, they performed worse than colleagues without AI. The UK Department for Transport built a consultation-analysis tool. It processed 200,000 responses and over 8 million words. Moreover, it agreed with human analysts more than 92% of the time on theme classification.
The pattern is consistent. LLMs lift the floor rather than the ceiling, and they do it on bounded, high-volume, language-shaped work.
The eligibility trap: why a 17% win becomes a 3% win
This is the single most useful number in this article for anyone building a business case.
A randomised field experiment on Alibaba’s Taobao platform covered 680,676 chats. An agentic AI system cut chat duration by 16.8% on the chats it could handle. Impressive — except only about 5.8% of chats were eligible. Consequently, the firm-wide effect on average chat duration was just 3.2%.
Nothing about the AI underperformed. The eligibility rate did the damage. Therefore, when you evaluate any LLM proposal, ask what percentage of your actual volume qualifies before you ask how well it performs. A pilot measures the numerator; your business case depends on the denominator.
What the aggregate data says about payoff
Two large studies temper the enthusiasm considerably. Researchers linked adoption surveys of 25,000 workers across 7,000 Danish workplaces to administrative payroll records. They found precise null effects of AI chatbots on earnings and hours worked.
Additionally, a survey reached nearly 6,000 senior executives across the US, UK, Germany and Australia. It found 69% saying their firm actively uses AI. Meanwhile, more than 90% reported no measurable effect on the outcomes they track.
Crucially, this does not mean LLMs do not work. It means task-level gains are not automatically reaching the income statement. Usually the eligible share of work is small, or the surrounding process never changed.
Real costs are not the token price
Inference prices genuinely collapse over time. Epoch AI measured the price of reaching a fixed performance milestone falling between 9× and 900× per year across six benchmarks. Independent MIT work puts it nearer 5–10× per year for a given benchmark level.
However, your bill is not the headline rate. Reasoning tokens bill as output tokens even though you never receive them. Tokenizer changes can add roughly 30% more tokens for identical text, quietly offsetting a price cut. Furthermore, real deployments carry line items beyond tokens — web-search calls, agent runtime, retries, evaluation, and human review.
Meanwhile, usage grows faster than prices fall. We cover the mechanics in Reducing AI Costs Without Reducing AI Power.
What changed legally in 2026
Two 2026 developments belong in any serious business case, and both are widely misreported.
First, the EU AI Act’s high-risk obligations did not begin on 2 August 2026. The Digital Omnibus on AI, published in the Official Journal in July 2026, moved those dates. What is live is Article 50 transparency: you must tell people they are interacting with an AI system and machine-readably mark synthetic media.
Second, chatbot output is now a documented liability. A German appellate court ruled in May 2026 that a company is liable for its own chatbot’s hallucinations. That echoes the earlier Moffatt v Air Canada tribunal decision. Notably, one public database has catalogued over 1,900 court decisions worldwide involving AI-hallucinated material.
A Southeast Asia footnote that changes procurement
For regulated and public-sector work in our region, one constraint outranks capability entirely. Singapore government agencies operate a hard data-residency ceiling: overseas-hosted GenAI API services may only be used with data up to defined classification limits.
Additionally, Singapore’s standard requires a legally binding commitment that the provider does not train on agency data. Most commercial LLM procurement never asks for that clause. Therefore, in this market, contract terms and data classification decide your architecture long before benchmark scores do.
How we would evaluate an LLM proposal in 2026
Four questions, in this order, before any model comparison:
- What share of the real volume is eligible? Measure the denominator first. A 17% improvement on 6% of your work is a rounding error.
- Is the task inside the capability frontier? Bounded, language-shaped, high-volume work pays. Judgement-heavy work outside the frontier reliably does not.
- How will you measure it without self-reporting? If your evidence is a satisfaction survey, you have no evidence.
- Who owns the output when it is wrong? Answer this before launch, not after the first hallucination reaches a customer.
How Pegotec helps
We size the eligible volume before we scope a build, and we instrument the result so the gain is measured rather than felt. Furthermore, we design around the constraints that actually bind in this region — data classification, residency, and no-training clauses.
If you are weighing an LLM investment and want an honest read on whether it will pay, talk to us.
Read next
- AI Chatbot ROI: When Does It Make Financial Sense? — the eligibility question applied to one specific use case.
- AI Agents One Year Later: What Businesses Actually Built — the same honesty test applied to agents.
- AI Model Selection Guide: Comparing Leading LLM Providers — once you know the work is eligible, pick the model.
Yes, but narrowly and unevenly. Controlled studies show real gains on bounded, high-volume, language-shaped work — 15% more issues resolved per hour across 5,172 customer-support agents, and 12.2% more tasks completed by consultants working inside the model’s capability frontier. However, large aggregate studies find precise null effects on earnings and hours, and most executives report no measurable effect on tracked outcomes. Task-level gains do not automatically reach the income statement.
No — that figure is a misquote. The MIT NANDA report said 95% of organisations are getting zero return, which is a different claim, and it rests on 52 structured interviews rather than a large survey. The related claim that RAND found 80% of AI projects fail is also a misattribution: RAND wrote “by some estimates” and cited others. RAND’s own finding was that 84% of interviewed practitioners blamed leadership-driven failures such as choosing the wrong problem, not model quality.
Usually because of eligibility. In a randomised experiment on Alibaba’s Taobao platform covering 680,676 chats, an agentic AI system cut chat duration by 16.8% on the chats it could handle — but only about 5.8% of chats were eligible, so the firm-wide effect was just 3.2%. The AI performed exactly as advertised; the addressable share of work was small. Measure what percentage of your real volume qualifies before you measure how well the model performs.
Probably not. US Census Bureau research covering November 2025 to January 2026 found only 18% of US firms using AI in a business function, and Eurostat recorded 19.95% of EU enterprises in 2025. Singapore’s IMDA reported SME adoption at 14.5%. Adoption is concentrated in large firms — 37% of US firms with 250 or more employees, and 55% of large EU enterprises. The typical business has not deployed an LLM in production.
The business is. A German appellate court ruled in May 2026 that a company is liable for its own chatbot’s hallucinations, echoing the earlier Moffatt v Air Canada tribunal decision, which rejected the argument that the chatbot was a separate entity. One public database has catalogued over 1,900 court and tribunal decisions worldwide involving reliance on AI-hallucinated material. Decide who owns wrong output before launch, not after it reaches a customer.
Let's Talk About Your Project
Enjoyed reading about AI LLMs for Business in 2026: A No-Hype Reality Check on What Actually Delivers Value? Book a free 30-minute call with our consultants to discuss your project. No obligation.