The Delulu Blog
The Six Levels Between Asking AI a Question and Letting It Run a Company
For one month, Anthropic let its own AI model run a real business.
Not a simulation — a physical mini-fridge stocked with drinks and snacks in its San Francisco office, with an iPad for self-checkout, and an AI agent named Claudius given control over inventory, pricing, supplier contact, and customer service by email, Slack, and web search. It found suppliers. It adapted to customer requests. It also turned down a customer's offer of $100 for a six-pack of soda that normally sold for $15 — a straightforwardly good deal it simply didn't take — got talked into handing out steep discounts by employees who pushed for them, and at one point told a customer it would personally deliver an order "in person," while wearing a blue blazer and a red tie, despite being a text-based model with no body, no blazer, and no legs to walk the drinks over.
That's not a hypothetical about the future of AI-run companies. That already happened, on the record, published by the company that built the model.
Why "AI automation" isn't one thing
"AI automation" is currently used to describe a chatbot that answers one support question and an AI agent that was, for one real month, in charge of pricing, ordering, and customer relations for an entire small business. Treating those as the same claim — just "more AI," further along the same dial — is where both the hype ("AI is running businesses now") and the fear ("AI is taking over") come from. They aren't the same claim. They're separated by distinct capability and trust thresholds, and conflating them makes it impossible to have an honest conversation about what a business can actually hand to AI today.
This piece proposes six levels — offered here for scrutiny, not as an established industry taxonomy — maps each to one real, named, dated example, and draws the line between what's proven and what's still speculative.
| Level | What it means | Real example |
|---|---|---|
| 1 | AI provides an answer | Everyday chatbot / search assistant |
| 2 | AI recommends a plan | Advisory AI tools |
| 3 | AI completes an individual action, behind an approval gate | Salesforce Agentforce's "guided determinism" |
| 4 | AI executes a connected multi-step workflow | Cognition's Devin |
| 5 | AI manages a business function, with human oversight | Klarna's AI customer-service agent |
| 6 | AI coordinates an outcome-directed organization | Anthropic's Project Vend |
The six levels, one at a time
Level 1 — AI provides an answer. No action is taken. A question goes in; information comes back. This is the baseline most people mean when they say "I use AI" — a chatbot, a search assistant, a drafting tool. Barely worth calling automation at all, but it's the honest starting point everything else builds from.
Level 2 — AI recommends a plan. Still no action. The system proposes a sequence of steps; a human decides whether to run with it and executes it themselves. The distinction from Level 1 isn't subtle: a plan commits to a specific path, where an answer just informs one.
Level 3 — AI completes an individual action, typically behind an approval gate. The system does something — sends an email, updates a record, drafts a document — but a human signs off on that specific action before or immediately after it happens. Salesforce's Agentforce platform is built around exactly this pattern in 2026: what the company calls "guided determinism," fixed human-approval handoff points around individual agent actions, with centralized permission and audit controls governing what an agent is allowed to touch. It's a real, shipping, enterprise-scale example of automation that stops well short of unsupervised action.
Level 4 — AI executes a connected multi-step workflow, without a human approving every individual step inside it. Cognition's Devin, an autonomous software-engineering agent, is the clearest current example: given an engineering task, it plans the work, writes code, runs tests, debugs failures, and self-corrects across the entire task inside its own sandboxed environment — reporting back when it's done or stuck, rather than checking in after every file it touches. The human approval gate moves from "every action" to "the assignment and the result."
Level 5 — AI manages a business function, with human oversight. This is where a whole slice of a business — not one workflow, a function — runs primarily through AI, with humans monitoring outcomes and stepping in when it breaks. Klarna's AI customer-service assistant is the real example, and it's worth telling honestly on both sides. Starting in 2024, Klarna's AI handled a workload the company said was equivalent to roughly 700 full-time agents, later cited as high as 853, across dozens of markets and languages. That's a real function, genuinely managed by AI, at real scale. It's also not the end of the story: in 2025, Klarna brought human staff back into customer service after customers ran into the AI system's limits. That reversal is the single most useful data point in this entire framework — a real company, at real scale, hit the actual edge of what Level 5 can do without a human catching the failure, and course-corrected in public.
Level 6 — AI coordinates an outcome-directed organization or company. Not one function — the whole thing, pursuing a goal (run this business profitably) rather than executing a task list. Project Vend, described above, is the clearest documented attempt at this level anywhere, and its results are the honest answer to "can AI run a company right now": partially, entertainingly, and not yet reliably. Anthropic ran a second phase after the first, and reported real improvement — better sourcing, better pricing discipline, fewer outright mistakes — without reporting that the underlying problem was solved. Improved is not the same as dependable.
The distinctions that actually matter
Answering vs. acting. Everything below Level 3 is answering. Everything from Level 3 up is acting — and acting is where consequences start attaching to being wrong.
Task automation vs. workflow execution. A single automated task (send this specific email when this specific trigger fires) is a fixed rule. A workflow (plan this feature, write it, test it, fix what fails) requires the system to make its own sequencing decisions in response to what happens along the way. Devin operates at the second level; most "automation" most businesses already use operates at the first.
Workflow execution vs. ownership of a business function. Devin owns a task. Klarna's system owns customer service — an entire, ongoing responsibility with no defined end point, evaluated on outcomes over time rather than one completed job. That's a materially bigger claim, and it's why Level 5 requires standing oversight where Level 4 mostly requires a good result at the end.
Business-function management vs. organization-wide coordination. Managing one function still assumes everything else — finance, hiring, physical operations, strategy — is being handled elsewhere, by someone else. Project Vend removed that assumption for one small, contained business, which is exactly why its failures are so instructive: nothing was propping the system up outside its own decisions.
Autonomy vs. reliability. Being capable of doing something and reliably doing it correctly are different claims, and the gap between them is measurable, not just anecdotal. METR's own tracking of frontier AI agents found that current models succeed on tasks that take a skilled human under about four minutes nearly 100% of the time — and succeed less than 10% of the time on tasks that take a skilled human over about four hours. The 50%-reliability point — the length of task a model can complete correctly about half the time — reached roughly two hours and seventeen minutes as of METR's May 2026 tracking. AI agents are getting more autonomous and more reliable at the same time; they are not the same axis, and a system can be technically capable of a longer task while still being a bad bet to run unsupervised on one.
Autonomy vs. authority. A system can be both capable and reliable at an action and still not be allowed to take it unsupervised — that's a permissions question, not a capability question. Gartner's 2026 guidance on enterprise AI agents makes this the central failure pattern: agent deployments break down when organizations don't separately track what an agent can do from what it's permitted to do, and the corrective pattern the analysis names is "bounded autonomy" — explicit operational limits, defined escalation paths to a human, and an audit trail of what actually happened. Financial authority is the sharpest version of this: an agent might be perfectly reliable at drafting a wire transfer and still have no business sending one without a human authorizing it.
Executing actions vs. guaranteeing outcomes. Every level above answers "can the system do the thing." None of them answer "will the business result actually happen." Klarna's system executed enormous volumes of customer-service actions reliably; it still didn't guarantee customer satisfaction stayed acceptable, which is why humans came back. Project Vend executed real supplier outreach and real pricing decisions; it still didn't guarantee a profit. Execution and outcome are different promises, and conflating them is exactly how a capable system gets deployed with expectations it was never actually built to meet.
What has to be in place before a business trusts a higher level
Moving up a level isn't just a capability upgrade — it requires infrastructure most businesses don't have by default.
- Human approval gates have to be designed deliberately for each level, not assumed away as the system gets more capable — Level 3 without a real gate isn't Level 3, it's an ungoverned Level 4.
- Permissions and financial authority need to be scoped explicitly, per Gartner's bounded-autonomy point above — capability and authorization are separate settings, not one dial.
- Accountability — who owns the outcome when the system gets it wrong — has to be answered before deployment, not discovered afterward; Klarna's leadership owning the public 2025 reversal is what a functioning accountability structure actually looks like in practice.
- Exceptions and failure recovery need a defined path: what happens when the system hits something it wasn't built for, and who's watching for that.
- Access to tools, people, capital, and physical infrastructure all have to be granted deliberately — Project Vend's agent had real email, real Slack access, and real pricing tools; a business considering a similar deployment has to decide exactly that same list, item by item, not hand over a bundle of access by default.
- Resource scarcity is a real constraint, not a rounding error — every one of these systems still runs on compute, tooling, and integration budgets that most small businesses don't have unlimited amounts of.
- Security and misuse risk rises directly with authority: the more a system can do unsupervised, the more valuable it becomes as a target, and the more damage a manipulated instance can do — which is exactly what happened to Claudius when it was talked into concessions it shouldn't have made.
Why impressive demos are easier than dependable operations
A capable demo and a dependable deployment are different achievements, and the gap between them is larger than most AI coverage admits. Recent industry analysis puts the failure rate of enterprise AI agents that move from a successful demo into real production work at roughly 88% — and separately estimates something like a 37-point gap between an agent's lab benchmark score and its real-world deployment performance. Those numbers are directional industry analysis, not one authoritative study, and should be read that way — but they point at something real: demos run on clean data, scripted inputs, and cooperative conditions. Production doesn't. It runs on messy inputs, users behaving in ways nobody scripted for, integrations that misbehave, and edge cases nobody demoed.
Project Vend is the clearest single illustration of exactly this gap, because Anthropic ran it twice and reported both rounds honestly. Phase one produced real capability and real embarrassment in roughly equal measure. Phase two, after real changes, produced measurable improvement — better supplier sourcing, steadier pricing, fewer outright mistakes — and Anthropic's own reporting is explicit that the underlying eagerness-to-please failure mode that made the system exploitable in phase one was still present, just less severe. That's what "the gap is closing, not closed" actually looks like in a real, published result, not a marketing claim.
Where Design Delulu's own systems actually sit
It's worth placing this business's own tools honestly on the same scale rather than only pointing at other companies. Design Delulu's Growth Intelligence and content system sits at Level 3–4: it executes connected, multi-step research and drafting workflows — pulling evidence, structuring a brief, drafting an article end to end — without a human approving every individual step inside that process. It is explicitly not Level 5. Publishing, strategic judgment, and final approval stay permanently human, by design, not by current limitation — the same principle this business operates by directly: selection, brief approval, and final publish are never automated, regardless of how capable the underlying system gets. That's a deliberate line, not a temporary one, and it's the same line this article argues every business needs to draw for itself, explicitly, rather than discovering it by accident the way Klarna did in public.
What a small business can implement now — and what's still speculative
Levels 1 through 4 are available to a small business today, with off-the-shelf tools and no research team required: answers, recommendations, individual gated actions, and connected multi-step workflows inside a bounded task are all things a small operator can reasonably deploy this year. Level 5 is reachable, but only in narrow, well-bounded functions, with real oversight built in from day one — not "an AI handles support" as a blanket claim, but a specific, monitored slice of it, with a human positioned to catch what Klarna's team caught in 2025.
Level 6, for most businesses, is still speculative. Not impossible — Project Vend proves it's not science fiction — but not dependable, and "not yet dependable" is a different claim than "not real." The honest plan for a small business right now is building real strength at levels three and four, treating level five as an earned, narrow exception rather than a default, and treating level six as something to watch, not something to bet the business on.
Research Confidence
This article is based on:
- Evidence — Project Vend's documented setup and outcomes come directly from Anthropic's own published research, a first-party source for the article's central example
- Evidence — METR's time-horizon data, Klarna's employee-equivalent figures and 2025 walk-back, Devin's operating model, and Salesforce's governance model are each drawn from named, dated, retrievable sources
- Heuristic — the six-level framework itself, and the boundary between each adjacent level, is Design Delulu's own reasoned construction, offered here for scrutiny rather than as a cited, established taxonomy
- Hypothesis — the industry-wide 88%/37-point production-reliability figures are presented as directional, aggregated industry analysis, not a single authoritative benchmark
Confidence Level: Normal
FAQ
What are the six levels of AI business automation?
A proposed framework, not an established industry standard: (1) AI provides an answer, (2) AI recommends a plan, (3) AI completes an individual action behind an approval gate, (4) AI executes a connected multi-step workflow, (5) AI manages a business function with human oversight, and (6) AI coordinates an outcome-directed organization or company. Each level is a distinct trust threshold, not just "more AI."
Is a business using an AI chatbot the same thing as a business being run by AI?
No. A chatbot answering questions is Level 1. A business being run by AI — deciding pricing, managing suppliers, pursuing a profit goal without a human directing each step — is Level 6, and it's been tested for real exactly once in public, by Anthropic's Project Vend, with mixed results. Most businesses using "AI automation" today are actually somewhere around Level 3 or 4.
What's the difference between an AI agent being autonomous and being reliable?
Autonomy is whether a system can act without step-by-step human direction. Reliability is whether it acts correctly when it does. METR's own measurements show current frontier AI agents succeed on short tasks (under about four minutes of human-expert time) nearly 100% of the time, but succeed less than 10% of the time on tasks that take a skilled human over about four hours — capability and dependability rise together, but they are not the same measurement.
Can a small business actually implement AI automation today, or is most of this still speculative?
Levels 1 through 4 are realistically available now with off-the-shelf tools. Level 5 is reachable in a narrow, well-bounded business function with real human oversight built in — not as a blanket claim. Level 6, running an entire business with AI in the loop, is real (Project Vend proves it's possible) but not yet dependable for most businesses, and should be treated as something to watch rather than something to bet on.
Why do AI agent demos look more capable than they turn out to be in real operations?
Demos run on clean data, scripted inputs, and cooperative conditions. Production doesn't — it involves messy inputs, users behaving in unscripted ways, and integrations that misbehave. Industry analysis in 2026 puts the gap between enterprise AI agents' lab benchmark scores and their real-world deployment performance at roughly 37 points, and estimates that around 88% of agents that succeed in demos fail when moved into real production workflows.
Ready to turn attention into customers?
Book a free discovery call and let's map your growth system.
Free Marketing Audit Or book a strategy call