When a flawless AI agent optimizes for the wrong thing

Klarna deployed one of the most technically impressive AI customer service agents ever built. Within a month, it was handling two-thirds of all customer service chats — the equivalent of 700 full-time agents. Resolution time dropped from 11 minutes to 2 minutes. Repeat inquiries fell 25%. The engineering was flawless.

And yet.

A human agent with five years of experience at Klarna “knew” something the AI didn’t. When a long-term customer called frustrated about a billing error, the human spent an extra three minutes listening, acknowledged the frustration, and waived a small fee that policy technically didn’t allow. The customer stayed for another five years. That decision wasn’t in any runbook. It wasn’t in any prompt. It lived in the coffee machine conversations, in the stories veteran agents told new hires, in the way managers praised relationship-building even when it slowed resolution metrics.

The agent had all the context it could need: customer data, ticket history, product knowledge, resolution procedures. What it didn’t have was the Foundation — the organizational values, tradeoff hierarchies, and decision boundaries that tell an agent not just how to resolve a ticket, but what the company actually wants when resolving it.

The agent was given a measurable objective: resolve tickets fast. It optimized brilliantly. It just optimized for the wrong thing.

The Foundation Gap is the missing layer between knowing how to do a task and knowing what your company values while doing it. Most organizations have given their AI agents procedures and context but never encoded their tradeoff hierarchies, decision boundaries, and red lines — so the agents execute flawlessly and optimize for the wrong thing.


Two Altitudes of Specification

We’ve been talking about the specification bottleneck for two years now — the idea that as AI makes production (code, content, analysis) nearly free, the ability to specify what to build becomes the scarce, valuable skill. That’s correct. But it’s incomplete.

Specification operates at two altitudes:

AltitudeQuestionWhat It EncodesExample
Task (Playbooks)“How do we do X?”Procedures, acceptance criteria, edge cases”When a customer requests a refund, verify purchase date, check return policy, process within 24 hours”
Organizational (Foundation)“What do we value when doing X?”Tradeoffs, decision boundaries, red lines, escalation triggers”When resolution speed conflicts with customer relationship for high-tenure customers, prioritize relationship”

Most companies deploying AI agents have invested heavily in the task altitude — building playbooks, writing procedures, engineering prompts with context. Almost none have invested in the organizational altitude. They’ve told their agents how to do things. They haven’t told them what to want.

This is the Foundation Gap. And it explains why a large majority of companies report no tangible value from their AI investments. The agents work brilliantly. They just don’t know what the company actually values.


The Progression

Three disciplines have emerged in rapid succession:

DisciplineQuestionScopeEra
Prompt Engineering”How do I talk to AI?”Individual, per-session2023-2024
Context Engineering”What does AI need to know?”Information architecture2024-2025
Intent Engineering”What does the org need AI to want?”Organizational purpose2025-2026+

Prompt engineering taught us to communicate with AI. Context engineering taught us to give AI the right information. Intent engineering teaches us to give AI the right purpose.

Most organizations are still at context engineering — building RAG systems, curating knowledge bases, assembling tool chains. This is necessary but not sufficient. Klarna’s agent had impeccable context engineering. It had no intent engineering at all.

Context without Foundation is a loaded weapon with no target.


What the Foundation Contains

The Foundation is not a document. It’s a structured, machine-readable encoding of organizational DNA:

  1. Tradeoff Hierarchies: When X conflicts with Y, prioritize Y under conditions Z. Not aspirational values on a poster — actual tradeoff rules extracted from how experienced employees make real decisions.

  2. Decision Boundaries: What agents can decide autonomously vs. what requires human judgment. Scoped by domain, risk level, and organizational sensitivity. “Process standard refunds automatically. Escalate any refund for a customer with 5+ years tenure.”

  3. Value Signals: Concrete behavioral examples that encode “how we do things here.” Not abstract principles — specific stories and decisions that illustrate the organization’s real values in action.

  4. Red Lines: Hard constraints that override all optimization. Things the organization NEVER does, regardless of what metrics suggest. “Never auto-close a ticket from a customer who has used the word ‘disappointed’ — always route to a human.”

  5. Escalation Triggers: Not just “when uncertain” but “when the decision touches organizational identity.” The decision about whether to bend a refund policy isn’t an uncertainty problem — it’s a values problem.


How to Capture the Foundation

The Foundation doesn’t live in documentation. It lives in coffee machine conversations, in the way veterans handle frustrated customers, in the stories new hires absorb over months. You can’t extract it with a survey or a wiki edit.

Five methods work:

1. Dilemma Workshops: Present leadership and experienced employees with realistic tradeoff scenarios specific to their business. The CHOICES reveal the real values. “Customer X wants a refund outside policy. 5-year client, clearly frustrated. What do you do, and WHY?”

2. Decision Autopsies: Examine past moments where someone broke the rules and got praised for it. These are the richest cultural signals. The exceptions reveal the REAL organizational values that were never formally encoded.

3. Story Capture: “Tell me about a time you made a call that wasn’t in the handbook.” “What’s the thing you tell new people that isn’t in onboarding?” Extract the values embedded in each story.

4. Red-Team the Foundation: “If an agent followed our stated values literally, what would break?” This is the Klarna test. The gap between what breaks and what shouldn’t reveals the unstated intent.

5. Revealed vs. Stated Values Audit: Compare what the website says vs. where money goes vs. what gets rewarded. “We value customer relationships” + “Agents measured on resolution speed” = Foundation Gap.


The Red-Team Test You Should Run Today

Here’s a simple exercise any executive can do in 30 minutes:

  1. Pick one function where you’re deploying or considering AI agents
  2. Write down your organization’s stated values for that function
  3. Now ask: “If an agent followed these stated values literally, with zero cultural context, what would go wrong?”

If the answer is “nothing” — congratulations, you have an explicit Foundation (rare).

If the answer involves any version of “well, experienced people know to…” — you have a Foundation Gap. And your agents are either already optimizing for the wrong thing, or they will be soon.


The Killer Insight

The company with a mediocre model and extraordinary organizational intent infrastructure will outperform the company with a frontier model and fragmented organizational knowledge.

Every company is deploying AI agents. Almost none have told those agents what the company actually values. The most important AI investment in 2026 isn’t a model subscription. It’s organizational intent architecture.

Context tells agents what to know. The Foundation tells agents what to want.

Build the Foundation first. Or ask Klarna what happens when you don’t.

Frequently asked questions

What is the Foundation Gap in AI? The Foundation Gap is the missing layer between knowing how to do a task and knowing what your company values while doing it. Most organizations have invested in the task altitude — playbooks, procedures, prompts with context — but not the organizational altitude: the tradeoff hierarchies, decision boundaries, and red lines that tell an agent what to want. The result is agents that execute flawlessly and optimize for the wrong thing.

What is the difference between context engineering and intent engineering? Context engineering gives an AI the right information — RAG systems, knowledge bases, tool chains. Intent engineering gives it the right purpose — the organization’s values, tradeoffs, and red lines. Klarna’s agent had impeccable context engineering and no intent engineering, which is why it could resolve tickets fast while missing what the company actually valued in a customer relationship.

How do you capture a company’s Foundation? Five methods work, because the Foundation lives in culture rather than documentation: dilemma workshops (real tradeoff scenarios reveal real values), decision autopsies (cases where breaking the rules got praised), story capture (decisions that weren’t in the handbook), red-teaming the stated values, and auditing revealed versus stated values (what the website says versus where money and rewards actually go).

What is the fastest way to find my Foundation Gap? Run a 30-minute red-team test. Pick one function where you’re deploying AI agents, write down your organization’s stated values for it, then ask: “If an agent followed these literally, with zero cultural context, what would go wrong?” If the answer involves any version of “well, experienced people know to…,” you have a Foundation Gap.


Samuel Pouyt is a geopolitical risk analyst and technologist. He advises leaders on AI transformation and writes about specification, organizational intent, and decision-making in the age of AI agents. He is the author of “Enough to Act: How Intelligence Analysts Cut Through Noise and Decide Under Uncertainty.”