
Why AI Agents Fail Without Organizational Intent
Key Takeaways
- 74% of companies report no tangible value from AI, and the problem is not the models. It is the absence of structured organizational intent.
- Klarna's AI saved $60 million and still failed: the agent improved for speed, not retention, because no one encoded what Klarna actually valued.
- Prompt engineering and context engineering are necessary but not sufficient. Intent engineering, telling agents what to want, is the missing third layer.
- Most organizations are broken at all three structural layers: unified context infrastructure, shared AI tooling, and agent-readable goal translation.
- A mediocre model with strong organizational intent will outperform a frontier model running on fragmented, unaligned knowledge. Every time.
The Real Reason Your AI Investments Aren't Working Has Nothing to Do With the Models
Most companies are spending millions on AI and getting almost nothing back. Not because the models are bad. The models are extraordinary. The problem is that nobody told the AI what the company actually wants.
That's not a technology problem. It's a leadership problem. And it has a name most organizations haven't heard yet: intent engineering.
I've watched this pattern play out across every major technology cycle for 25 years. A powerful new capability arrives. Companies rush to deploy it. The ones who fail aren't the ones with bad technology. They're the ones who never translated their organizational purpose into something the technology could act on.
This time the stakes are higher. Unlike previous tools, AI agents don't just process information. They make decisions. And a decision-making system without clear intent is not neutral. It's dangerous.
What Does Klarna's $60 Million Mistake Actually Teach Us?
In early 2024, Klarna launched an AI customer service agent that handled 2.3 million conversations in its first month across 23 markets and 35 languages. Resolution times dropped from 11 minutes to 2 minutes. The CEO projected $40 million in savings. By early 2025, the company claimed the AI did the work of 853 full-time employees and saved $60 million.
Then customers pushed back. Generic answers. Robotic tone. Zero ability to handle anything requiring judgment.
By mid-2025, CEO Sebastian Siemiatkowski told Bloomberg that "while cost was a predominant evaluation factor, the result was lower quality." Klarna started rehiring the human agents it had let go.
Here's what most people miss about this story. The AI agent wasn't bad at its job. It was extraordinarily good at resolving tickets fast. Resolving tickets fast was the wrong goal.
Klarna's actual organizational intent was to build lasting customer relationships that drive lifetime value in a competitive fintech market. A senior human agent with five years at the company knows intuitively when to bend a policy, when to spend extra time because a customer's tone signals they're about to churn, and when efficiency is the right call versus when generosity is. That knowledge lived in daily management decisions, in stories veterans told new hires, in unwritten rules about which metrics leadership actually cared about.
The AI agent had a prompt and context. It did not have intent.
There's a sharper observation worth sitting with here. The AI may have inadvertently reflected Klarna's real values, cost savings first, rather than its stated ones. Customer pushback forced the company back toward what it claimed to believe. The $60 million in savings wasn't close to enough to offset becoming the poster child for AI-driven customer service failure.
Why Is 74% of AI Investment Producing No Tangible Value?
The numbers tell a consistent story across every major research firm. 74% of companies globally report they have yet to see tangible value from AI. McKinsey found 30% of AI pilots failed to achieve scaled impact. Deloitte's 2026 State of AI report, covering 3,000+ leaders across 24 countries, found that 84% of companies have not redesigned jobs around AI capabilities. Only 21% have a mature model for agent governance.
The money is flowing. Deloitte's Tech Value Survey showed 57% of respondents putting 21 to 50% of their digital transformation budgets into AI automation. One in five companies invested over half, averaging $700 million for a company with $13 billion in revenue. Gartner predicts that by 2028, 15% of day-to-day work decisions will be made autonomously by agents. I think that estimate is low.
Capital is abundant. Capability is abundant. What's scarce?
Organizational clarity about what the AI should actually be doing.
Look at Microsoft Copilot. They embedded AI into every Office application and launched an aggressive enterprise push. 85% of Fortune 500 companies adopted it. But only 5% of organizations moved from a Copilot pilot to larger-scale deployment, according to Gartner. Only about 3% of the total Microsoft 365 user base adopted Copilot as paid users. Bloomberg reported Microsoft slashing internal sales targets after most salespeople missed their goals.
The standard explanation focuses on UX and model quality. I think the real issue is different. Deploying an AI tool across an organization without intent alignment is like hiring 40,000 new employees and never telling them what the company does, what it values, or how to make decisions. You get AI usage metrics on a dashboard and almost no measurable impact on organizational objectives.
What's the Difference Between Prompt, Context, and Intent Engineering?
These three disciplines represent the evolution of how humans communicate with AI systems. Understanding where each one breaks down matters.
Prompt engineering was the first discipline. Individual, synchronous, session-based. A person crafts an instruction and iterates on the output. The value is personal and non-transferable.
Context engineering is the discipline the industry is currently grappling. Anthropic published a foundational piece defining it as "the shift from crafting isolated instructions to crafting the entire information state that an AI system operates within." LangChain's Harrison Chase described it in a Sequoia Capital interview as encompassing "everything we've done at LangChain without knowing the term existed." Context engineering includes building RAG pipelines, wiring up MCP servers, structuring organizational knowledge so agents can access it. Necessary. Not sufficient.
Intent engineering is the third discipline, and it's largely unbuilt. Where context engineering tells agents what to know, intent engineering tells agents what to want.
It's the practice of encoding organizational purpose into infrastructure. Not as prose in a system prompt, but as structured, practical parameters that shape how agents make decisions autonomously. It's the layer that would have told Klarna's agent: "Yes, you can resolve this ticket in 90 seconds, but this customer has been with us for years and their tone indicates frustration. Spend the extra time, offer a specialist. The goal is retention."
Where Exactly Does the Intent Gap Show Up?
The gap operates across three structural layers, and most organizations are broken at all three.
Layer 1: Unified Context Infrastructure
Right now, every team building agents constructs its own context stack independently. One team pipes Slack data through a custom RAG pipeline. Another exports Google Docs into a vector store. A third built an MCP server connecting to Salesforce but not Jira. A fourth doesn't know the other three exist.
This mirrors the shadow IT crisis of the early cloud era, but with higher stakes. Agents don't just access data. They act on it.
Anthropic introduced the Model Context Protocol (MCP) in late 2024 and donated it to the Linux Foundation in December 2025. OpenAI, Google, Microsoft, and more than 50 enterprise partners have committed to it. Monthly SDK downloads are approaching 100 million. Protocol adoption and organizational setup are different things, though. Having a standard doesn't determine which ports to install, who maintains them, or what gets plugged in.
The real questions are architectural and political. Which systems become agent-accessible? Who decides what context an agent can see across departments? How do you version organizational knowledge so agents aren't operating on stale information? Deloitte's 2025 survey found that nearly half of organizations cited data searchability and reusability as top challenges blocking AI automation.
Layer 2: Coherent AI Worker Toolkit
Across most organizations, individuals are using incompatible, non-transferable AI workflows. One person uses one model for research and a different one for drafting. Someone else built a custom agent chain. Another is copy-pasting into a chat window. None of these workflows are transferable, measurable, or improvable by others.
There's an important distinction between AI activity and AI fluency. AI activity, bolting AI onto existing workflows, produces roughly 30% gains. AI fluency, rethinking the workflow itself around AI capabilities, produces gains closer to 300%. Fluency doesn't scale through training alone. It scales through shared infrastructure.
Deloitte's 2026 report found that workforce access to sanctioned AI tools expanded by 50% in a year. Access alone is insufficient without the organizational context and data that allow tools to deliver real value.
Layer 3: Intent Engineering Proper
This layer almost certainly doesn't exist in your organization. It's the one that matters most.
OKRs were designed for people. They assume human judgment about prioritization, trade-offs, values, and exceptions. A manager can tell a direct report "here's what matters this quarter" and trust they'll interpret that guidance through institutional context, professional norms, and judgment developed over months and years.
Agents have none of that. An agent doesn't know your company's OKRs unless they're placed in the context window. It doesn't know which trade-offs leadership would prefer unless those preferences are encoded in an specific way. Company culture can't be absorbed through osmosis, through all-hands meetings, hallway conversations, or watching senior people handle ambiguous situations.
What does intent engineering actually require?
Goal structures agents can interpret and act on. Not "increase customer satisfaction," which is a human-readable aspiration, but agent-practical objectives. What signals indicate satisfaction in this context? What data sources contain those signals? What actions is the agent authorized to take, and where are the hard limits?
Delegation frameworks. Organizational principles translated into decision boundaries. Amazon's "customer obsession" leadership principle works for humans because humans interpret it through contextual judgment. An agent needs it decomposed: when customer request X conflicts with policy Y, here is the resolution hierarchy.
Feedback mechanisms that close the loop. When an agent makes a decision, was it aligned with organizational intent? How is that measured and corrected over time? At Klarna, the agent tuned for resolution speed because that was the objective it could measure. The objectives that actually mattered lived in the heads of the human agents who were let go.
Why Hasn't Anyone Built This Yet?
Three reasons, and they're all structural.
It's genuinely new. Before agents could run autonomously over long time horizons, humans served as the intent layer. Long-running agents now operate over weeks, soon over months. That breaks the model entirely.
The two-cultures problem. The people who understand organizational strategy are not the people who build agents. MIT found that AI investment is still viewed primarily as a technology challenge for the CIO rather than a business issue requiring leadership across the organization. Intent comes from the entire leadership team working together, not from one function.
It's genuinely hard. Most organizations have never had to make their intent explicit and structured. Goals live in slide decks, OKR documents referenced once a year at performance reviews, and the tacit knowledge of experienced employees who know what to do in ambiguous situations even though they've never been formally told.
This is exactly the kind of structural challenge we designed the AI readiness note to address. Not another technology assessment. A clear-eyed look at whether your organizational intent is discoverable, structured, and specific by the AI systems you're deploying.
What Does a Real Solution Look Like?
It operates at three levels, and they have to be built together.
At the infrastructure level, you need a composable, vendor-agnostic architecture that lets agents operate across systems, tools, and models securely and at scale. MCP provides a protocol layer, but organizational build requires decisions about data governance, access controls, freshness guarantees, and semantic consistency. Treat this as a core strategic investment comparable to your data warehouse strategy. Not an IT project.
At the workflow level, you need an organizational capability map for AI. A shared, living understanding of which workflows are agent-ready, which are agent-extended with human-in-the-loop, and which remain human-only. This isn't a static document. It's an operating system that evolves as agent capabilities improve. I expect a new organizational role to emerge from this, something like an AI Workflow Architect sitting between engineering, operations, and strategy.
At the alignment level, you need goal translation infrastructure that converts human-readable organizational objectives into agent-specific parameters. Decision boundaries. Escalation protocols. Value hierarchies for how agents resolve trade-offs. Feedback loops for measuring and correcting alignment drift over time.
Early technical work is pointing in the right direction. Google's Agent Development Kit separates agent context into distinct layers: working context, session memory, long-term memory, and artifacts, each with specific governance. Academic work from Google DeepMind researchers proposes five levels of AI agent autonomy, from Operator to Observer, each with different intent alignment requirements and human oversight models.
The Race Changed and Most Companies Didn't Notice
For three years, the AI race was framed as an intelligence race. Best model, best benchmarks, biggest context window. That framing made sense when models were the bottleneck.
Models are no longer the bottleneck for most organizational use cases.
The frontier models available today are all extraordinarily capable. The differences between them matter far less than the differences between organizations that give them clear, structured, goal-aligned intent and organizations that don't. Full stop.
A company with a mediocre model and strong organizational intent infrastructure will outperform a company with a frontier model and fragmented, inaccessible, unaligned organizational knowledge. Every single time.
If OKRs were the management innovation that allowed Intel to align thousands of humans to shared objectives in the 1970s, intent engineering is the management innovation that allows organizations to align hundreds or thousands of agents to those same objectives in 2026, while those agents operate at speeds and scales no human manager can supervise.
OKRs took decades to become standard management practice. We don't have 20 years.
The most important AI investment you'll make this year is not a model subscription or another Copilot license. It's organizational intent architecture: making your company's goals, values, decision frameworks, and trade-off hierarchies discoverable, structured, and agent-practical.
That's not a technology project. It's a leadership decision. And the companies that make it now will be the ones that actually get returns from the AI investments everyone else is still trying to justify.
Infographic

Frequently Asked Questions
- What is intent engineering and how is it different from prompt engineering?
- Prompt engineering is personal and session-based: you write an instruction, you get an output. Intent engineering is organizational and persistent. It is the practice of encoding your company's goals, decision boundaries, and value hierarchies into infrastructure so agents can act on them autonomously, not just in a single conversation. Context engineering tells agents what to know. Intent engineering tells agents what to want.
- Why did Klarna's AI customer service project fail despite saving $60 million?
- The AI did exactly what it was built to do: resolve tickets fast. But fast resolution was the wrong goal. Klarna's actual goal was customer retention and lifetime value. The agent had no way to know when to bend a policy, spend extra time, or escalate to a specialist. That judgment lived in experienced human agents who were let go. The $60 million in savings could not offset the customer pushback and reputational damage.
- Why are most AI pilots failing to scale beyond proof of concept?
- Gartner found that only 5% of organizations moved from a Copilot pilot to larger-scale deployment. The usual explanation focuses on UX or model quality. The real issue is that deploying AI without intent alignment is like hiring thousands of employees and never telling them what the company does or how to make decisions. You get usage metrics on a dashboard and almost no measurable impact on business objectives.
- What are the three structural layers where the intent gap shows up?
- First, unified context infrastructure: most organizations have teams building disconnected agent pipelines with no shared architecture. Second, coherent AI worker tooling: individuals run incompatible workflows that cannot be shared, measured, or improved across the org. Third, intent engineering proper: the layer that translates OKRs, values, and decision trade-offs into parameters agents can actually interpret and act on. Almost no organization has built this third layer.
- What does intent engineering actually require in practice?
- Three things. First, goal structures agents can interpret: not aspirational prose like 'increase satisfaction' but specific signals, data sources, and authorized actions. Second, delegation frameworks: organizational principles decomposed into decision hierarchies agents can follow when inputs conflict. Third, feedback loops: mechanisms that measure whether agent decisions aligned with organizational intent and correct drift over time.
- How does intent engineering compare to OKRs as a management practice?
- OKRs were the management innovation that let Intel align thousands of humans to shared objectives in the 1970s. Intent engineering is the equivalent for agents in 2026. The difference is that humans interpret OKRs through judgment, context, and institutional memory. Agents cannot do that. They need goals, trade-off hierarchies, and decision boundaries made explicit and structured before they can operate reliably at scale.