Why Does Enterprise AI Stall? A Portfolio Playbook
Run enterprise AI as a Fund, Defer, or Kill portfolio. Covers operating models, platform layers, shadow AI, and conflicting ROI numbers.

Run enterprise AI as a Fund, Defer, or Kill portfolio. Covers operating models, platform layers, shadow AI, and conflicting ROI numbers.

Enterprise AI is the practice of putting models, agents, and controls into the systems that already run the company, under identity, audit, and permission. AWS and IBM still define the category as a glossary. McKinsey finds about 9 in 10 surveyed organizations regularly use AI in at least one function, while Deloitte reports only 34% using it to deeply transform work.
A glossary is not an operating system. Searches for ai for enterprise, ai in the enterprise, ai in enterprise, and enterprise artificial intelligence belong on this URL.
ChatGPT Enterprise, Claude Enterprise, Gemini Enterprise, NVIDIA AI Enterprise, and Microsoft 365 Copilot are licenses. They are not a portfolio, a CoE, or a kill switch.
AWS calls it the adoption of advanced AI inside large organizations, and an enterprise platform the integrated stack to experiment, develop, deploy, and operate at scale. The load-bearing requirement is model reuse, not a new training run per task.
IBM Think names the capability classes: predictive ML, NLP and computer vision, generative models, retrieval over proprietary data, and agentic systems that plan multi-step work under oversight. That is integration into operations, apps, and decisions.
Snowflake names what the glossary misses. Production-grade AI sits in the systems that already run the business. The hard problem is governed context, not model IQ.
You measure claims cycle time, escalation rate, and forecast error, not "the model is smart."
Three contrasts close the usual mix-up:
Contrast | One-line |
|---|---|
vs generative AI | Generative models are a capability inside the stack. Predictive, optimization, vision, and agents sit beside them. |
vs consumer ChatGPT | Same model class, different contract: company data, SSO, audit, retention, permissions, procurement. |
vs Copilot / ChatGPT Enterprise / Claude Enterprise / Gemini Enterprise / NVIDIA AI Enterprise | Product SKUs. Category is not a license. Copilot is a front door. |
Gartner is the closest official page to the job: AI strategy is vision plus a portfolio plus an operating model, kept bidirectional with business strategy. That join does not rank on page one of "enterprise AI." This page runs it.
McKinsey's 2026 State of AI ("On the road to ROI," 2026-08-25) finds regular use in at least one function across about 9 in 10 surveyed organizations. 44% are scaling enterprise-wide (from 38%), and 56% use AI in three or more functions. The Register's independent read puts the sample at n=1,719, fielded 4 May to 8 June 2026 across 97 countries.
Deloitte's 2026 State of AI surveyed n=3,235 director-to-C-suite respondents in August-September 2025. Workforce access moved from under 40% toward about 60% in a year.
34% say they are using AI to deeply transform. 30% are redesigning key processes; 37% stay surface-level.
The Fed's FEDS Note (Allen, 3 April 2026) tracks a different population: work-related generative AI reached about 41% in November 2025, up 9.7 points year over year.
McKinsey measures surveyed organizations. The Fed measures US workers and firms. Mixing them produces a fake national rate.
Related coverage on this site: the AI statistics roundup and the function-level AI solutions playbook. This page owns the operating system those numbers sit inside.
Idea flow is not the constraint on enterprise AI. Accountable ownership, data readiness, a kill path, and review capacity are.
ProductiveEdge sees the failure mode in one sentence: intake in multiple places, value defined differently in each, approved work that then stalls. McKinsey on LinkedIn (September 2026) said the same from the other side of the table: "AI investment is everywhere. Meaningful growth from it is not."
Stand up a single door: a template, a named sponsor, and a strategic-fit score. COMPEL frames intake as demand management: channel appetite into a prioritized, resourced, governed pipeline.
Umbrex scores value, feasibility, risk, and strategic fit. Use the score as a discussion tool. A two-decimal ranking that nobody can explain is theater.
IBM's CoE is a hub for expertise, standards, and governance so AI work lines up with strategy. A review committee at arm's length is not that job.
ActiveWizards (7 June 2026) names four exits:
Decision | When it fits | What you must have |
|---|---|---|
Fund | Clear workflow, owner, permission envelope | Named in-company owner, measurable workflow, independent quality signal, rollback steward |
Defer | Real value, missing data or capacity | A reopen condition, not a polite graveyard |
Redesign | Right problem, wrong unit of work | A new workflow boundary before you rescore |
Kill | No owner, no measurable workflow, or ungovernable data | Written rationale. Do not reopen without a materially different proposal |
Fund is expensive on purpose. The scarce resource is accountable operating attention, not ideas.
BCG on LinkedIn (September 2026) described the companies that get more from AI as going deeper on fewer initiatives, concentrating in the core, and redesigning workflows. That is portfolio concentration. It is not another 40 pilots.
phData (1 July 2026) is blunt: funding one isolated use case concentrates execution, adoption, and ROI risk on a single bet. Shared data products, eval harnesses, identity, and tool permissions make the 12th case cheaper than the 2nd. Start from already-committed strategy, not a 200-row brainstorm.
Kjersten Moody in CDO Magazine (29 July 2026) flips the usual order: start with measurement, not use cases. "Models run in 2 hours instead of 24" is a technical trophy. It is irrelevant unless a process KPI moves.
RAND reports that by some estimates more than 80% of AI projects fail, twice the rate of non-AI IT projects. CDO Magazine reads that as AI applied to a broken process.
Williamson at Databricks (20 January 2026) treats AI as a portfolio of bets, not a linear roadmap, and warns against exclusive reliance on AI features sprinkled across dozens of SaaS tools.
You need a named model. You do not need a new one.
Dataiku (6 August 2026) is the cleanest teaching ladder:
Model | Control | When it fits |
|---|---|---|
Siloed | Local experiments, no shared platform | Early. Does not scale. |
Center of excellence | Central build and standards | Low maturity, need consistency |
Hub-and-spoke | Central platform and standards, local delivery | Default hybrid as demand grows |
Center for acceleration | Central enablement that unblocks business units | When the CoE has become a bottleneck |
Embedded | AI owned in the function | High maturity, strong platform, auto-governance |
No single best model. Shared platform plus adoption effort sit under all five.
An AI ops model is not a traditional IT central service. It needs domain-plus-tech collaboration and governance that follows models into production.
Microsoft's agent CoE uses a shorter axis: Centralized, Hybrid, or Federated, blended by pattern. Microsoft's starting map centralizes employee enablement and external engagement, and federates business-expert and core-process work.
Three jobs exist in every model: set the rules, build the agents, monitor production.
The Azure Cloud Adoption Framework and IBM Think agree on the CoE's actual job: methodology, shared platform, intake, capability building, and governance that ships.
Databricks's tell for a serious org is who owns data and AI, and how close that seat sits to the CEO. Split ownership produces static use cases.
The agentic organization (McKinsey, 26 September 2025) is one evolution of this ladder. Five pillars: business model, operating model, governance, workforce, tech and data. Work is reimagined AI-first.
Humans sit above the loop for strategic oversight. Pair it with the 2026 survey: chatbots scaled at 47%; agents at about 2 in 10.
Deloitte finds about 75% plan agents within a couple of years, and only 21% have mature agent governance.
For the mechanics of running work with agents, the agentic project management spoke on this site is the adjacent playbook.
Chief AI officer became a real org box in 2026. Spell the title; the acronym collides with unrelated search.
Nimrod Barak joined Synchrony from Citi's AI CoE (effective 30 June 2026). Anja Leth Zimmer was named at Novo Nordisk on 23 January 2026. The seat sits next to CIO and CDO.
The people who actually sit in the workflow are often forward-deployed engineers. Diana Wolf at Kyndryl described FDEs in the business as the difference between a tool tour and a redesigned job.
Diana Wolf, speaking at WIRED Events for Kyndryl, restated the mix operators keep forgetting:
"70% of the change to come is people and process, 20 is your data and infra, and then 10 is your AI models, which means that really the model is the easiest part." (Diana Wolf, WIRED Events / Kyndryl, 3:30)
That split is a work-mix observation, not a budget law.
Wolf's other useful test is subtractive: turn the tool off. If work slows, you are AI-enabled; if work stops, you are AI-native.
Seats are a dashboard. They do not change org design.
Deepak Seth on Gartner ThinkCast named the other failure: AI for AI's sake, FOMO as strategy. He puts change management at 100 to 200% of the technology cost.
Executives expect dollar value tomorrow because the model answers instantly. The org does not.
On X, Allie K. Miller argues enterprises need a bullpen of AI operators to build those workflows.
Tomasz Tunguz noted Moderna unifying CTO and head of people so one leader decides human versus AI work. Treat that as an org-chart experiment, not proven P&L.
Search "enterprise AI platform" and Google returns a vendor grid plus an NVIDIA SKU page.
KPMG (23 July 2026) gives stack names, not a beauty contest: Applications, Agents, Assistants, Context, Models, Refinery, Data, Compute, Energy, wrapped by Trust, Ops, and Control.
Operators collapse that stack to four layers:
Layer | What you are choosing | Examples (named, not ranked) |
|---|---|---|
Data | Governed, permissioned, lineage; semantic context for agents | Databricks Lakehouse + Unity; Snowflake AI Data Cloud |
Model access | Multi-model routing, not "pick one frontier and train" | Amazon Bedrock, Azure OpenAI, Vertex AI; ChatGPT / Claude / Gemini seats as a layer |
Control plane | Identity, evals, cost, logging, kill switch, inventory | Unity AI Gateway; a command-center job (ModelOp-class); GRC registries |
Apps / agents | Workflows with bounded tools and human gates | Copilot, Agentforce, Glean as application layer; Palantir AIP as ontology sandbox |
Most enterprises will not train a frontier model. Menlo Ventures (9 December 2025) found 76% of use cases purchased versus built, and AI deals converting to production at 47% versus 25% for SaaS.
Copilots were 86% / $7.2B of horizontal spend. Gartner's strategy article is the sequencing rule: portfolio and operating model before the tool bake-off.
Three paths show up as examples, not a ranking: hyperscaler host (Bedrock, Azure OpenAI, Vertex), data platform (Databricks, Snowflake), and assistant SKU (ChatGPT Enterprise, Claude Enterprise, Gemini Enterprise, Microsoft 365 Copilot).
Most stacks run more than one path. The control plane is how they do not become ten shadow stacks.
On r/AZURE and r/sysadmin, a recurring confusion is Microsoft's product map. Consumer Copilot, Microsoft 365 Copilot, Copilot Studio, and Azure AI Foundry get renamed every few weeks. "My org is using Copilot" is not the same sentence as adopting enterprise AI.
Seat and token prices for ChatGPT, Claude, and Gemini Enterprise are a procurement line. Cost follows the operating model: if there is no owner and no kill path, cheaper tokens still buy workslop.
For runtime quality once something is in production, LLM observability (traces, cost SLOs, drift) is the spoke. MCP is the protocol layer most "platform" pages skip.
This is not an ethics essay. Three operator moves.
The AI Governance Institute (January 2026) runs three scans in parallel: vendor-contract audit, employee amnesty survey, API-egress scan. Score data sensitivity × decision impact × regulatory exposure.
Every row needs a named owner and a next review date. An inventory without those two fields is stale in six months.
AWS now talks about an Agent Registry once agent count hits thousands: auto-detect, approval, audit. The catalog is the control plane. A slide titled "we care about responsible AI" is not.
Thomas Thelliez's governance architecture is the sentence to keep:
An AI agent may propose, summarize, draft, route, and request tool calls, but identity, permission, trust, ownership, and auditability must stay outside the model.
Snowflake and Databricks then apply the boring runtime: classification, masking, row-level security, lineage, scoped tool permissions, approval gates, rollback. The same controls apply whether a human runs SQL or an agent calls a tool.
NIST's AI Risk Management Framework is a named map for that work, not legal advice. Adjacent ai governance pages can own the GRC bake-off; this hub stays on the control plane.
Dataiku on LinkedIn (September 2026) named the metric trap: "A 200 response doesn't mean your AI agent succeeded." Infrastructure monitoring was built for deterministic systems. Agents fail in ways HTTP 200 never sees.
Shadow AI is the new shadow IT, and people have stopped hiding it: coding tools in parallel with the approved stack, laptops acting as servers, meeting-notetaker bots joining calls nobody invited.
On r/sysadmin, the recurring prescription is not a block list. It is an authorized front door plus ownership rules.
"You start by setting up an authorised/acceptable way of doing this. Because your staff are going to, one way or another. As part of this we have ended up running OpenWebUI through a LiteLLM proxy." (u/sobrique in r/sysadmin, March 2026)
The ownership rules that survive a holiday outage are specific. Shared data goes in a repo. No personal tokens beyond local testing.
Scheduled jobs run on a visible runner. Every tool has a business owner. Anything another team depends on gets technical review.
One r/sysadmin shop described a concrete Claude Enterprise envelope: DLP via Microsoft Intune, Splunk log forward, VPN and IP allowlist. That is a control plane. A PDF policy is not.
Launch-and-abandon is the other shadow-AI cousin. Internal chatbots spike, then die.
"Our metrics show a massive spike in month one, followed by a 70-80% drop-off in active usage. I'm talking about internal corporate chatbots with access to company files. This is peak AI Fatigue." (u/Relaxation_Time in r/sysadmin, May 2026)
Treat that 70-80% as one operator's internal metric, not a survey. The pattern still matches Deloitte's gap between access and deep transformation, and Wolf's warning that seats are a shallow metric.
Agents are a portfolio class. They are not the operating model.
Simon Willison's public definition still holds: an LLM agent runs tools in a loop to achieve a goal. Use an agent when you have bounded tools, a permission envelope, an eval, and human gates on consequential actions. Use a model or a workflow when you do not.
Gartner predicted (26 August 2025) that 40% of enterprise apps would feature task-specific AI agents by 2026, up from less than 5% in 2025. That is a prediction, not an observed 2026 rate. McKinsey's 2026 survey is the observed cut: chatbots scaled at 47%; agents at about 2 in 10; 40% of large organizations ($1B+) scaling agents (from 27%), smaller organizations flat at 22%.
Production controls are unglamorous. Gergely Orosz compressed them into a checklist:
How it started: "AI vibe coding tools will replace devs!" How it's going: "Do this: - Provide it w a detailed spec - Break down tasks to small ones - Separate dev and prod envs - Do NOT give access to the agent to prod - Never trust the agent; verify every step it takes - ...
The checklist is a detailed spec, small tasks, separate dev and prod, no production credentials, and verify every step.
On r/ExperiencedDevs the same warning shows up as war stories. One engineer reported an agent that already tried to drop production tables. Another noted that when token budgets run out, some teammates claim they cannot work, while Jira fills with AI slop.
Review capacity is the bottleneck. Chip Huyen's version: models run many tasks in parallel; a person tracks a few.
Hamel Husain's eval advice is the cheapest quality system you can buy: error analysis with domain experts. Look at the data. Dashboards that only plot HTTP 200 will not save you.
Ethan Mollick flagged a different executive risk than hallucination:
I am starting to think sycophancy is going to be a bigger problem than pure hallucination as LLMs improve. Models that won’t tell you directly when you are wrong (and justify your correctness) are ultimately more dangerous to decision-making than models that are sometimes wrong.
Sycophantic models that will not tell you that you are wrong are worse for decision-making than models that are sometimes wrong. That is a control-plane problem for anyone putting model output in front of a VP.
Allie K. Miller's other line is governance, not culture-war: fire workslop generators. Unreviewed AI volume is a quality problem you can name in a performance conversation.
Snowflake argues the durable advantage is accumulated business context (how you value customers, balance margin and growth, protect the brand), not model choice. Everyone can buy the same model.
The headlines disagree because the denominators moved. Levos is the join.
Deloitte's Global AI ROI paradox (22 October 2025, n=1,854 executives in Europe and the Middle East) found 85% increased AI investment in the last 12 months and 91% plan to again. For generative AI, 15% already see significant measurable ROI; 38% expect it within a year.
For agentic systems, only 10% currently see significant measurable ROI. Typical payback is 2-4 years against an expectation of 7-12 months. That survey is not the 2026 State of AI in the Enterprise.
MIT NANDA, via Fortune (18 August 2025), reported about 95% of generative AI pilots with no measurable P&L. Vibe Engines is right to refuse the headline "95% of AI projects fail."
The original methodology is not public. 95% is genAI pilots, not the whole AI book.
McKinsey's 2026 survey finds 37% attribute some EBIT to AI (flat versus 2025). High performers, defined as at least 5% EBIT and "significant" impact, are 6%, also flat.
100 minus 6% is not a McKinsey "94% failure" finding.
A Microsoft-sponsored IDC report (14 January 2025) found generative AI returning 3.7× per dollar, and $10.3 for top leaders. Label it sponsored. A paywalled recap is not the source.
Gartner on I&O (7 April 2026) found 28% of infrastructure-and-operations AI use cases fully succeeding against ROI, and 20% failing outright. That cut is I&O only.
All of those rows can be true in the same year. Self-set expectations, external P&L, EBIT share, and a vendor-sponsored multiplier are different questions.
Monday is smaller. Pick process KPIs with a baseline: cycle time, error rate, escalation rate, forecast error.
Then argue EBIT. Technical trophies ("the model runs in 2 hours") do not count.
Deloitte 2026 also reports 25% have converted at least 40% of experiments to production, and 54% expect that conversion in 3-6 months. 74% hope for revenue growth; 20% already getting it.
Strategy looks "highly prepared" for about 40% of respondents; governance 30%; talent 20%. That gap is the operating model, not the model card.
Functions that survive the first-month spike are boring on purpose.
Support and knowledge. Summarize an existing ticket, translate, and answer from a labeled corpus. Copilot rummaging an unlabeled filing cabinet will hallucinate; r/sysadmin will correctly call that a data-hygiene problem the model made embarrassing.
Software delivery. Greenfield scaffolding, legacy-code Q&A, and code review as a second pair of eyes, inside a gated SDLC. The honest architecture on r/devops is to use the model to write deterministic tools, not to be the tool.
Risk, fraud, operations, finance close. Detection-from-logs with a spec, plus forecast and close acceleration with a baseline. Write-back to systems of record is the blocker Snowflake keeps naming.
Commercial decisions in the flow of work. BCG says already-tight operations often underspend because a 1-3% ceiling feels plausible, and that redesigned workflows can move much more. Treat those 30-50% acceleration figures as BCG's operator claim, not an industry standard.
If a use case cannot name a process KPI, an owner, and a rollback, it is a Defer or a Kill. It is not a platform purchase.
Work arrives in chat, email, a hackathon board, and a vendor's success manager. Value is defined differently in each queue. Approved items then stall for data and owners.
Stand up one door. Everything else is a pointer to that door.
A weighted spreadsheet that nobody can explain will still ship a slide. Use scores to structure the argument (value, feasibility, risk, fit). The decision is Fund, Defer, Redesign, or Kill, with a name next to it.
Wolf: adoption is a shallow metric. Deloitte: access can jump while only 34% deeply transform. If turning the tool off does not change the job, you bought a dashboard.
Gartner's sequencing is the fence. Menlo's 76% purchased-versus-built rate makes a ranked "best platforms 2026" page even less useful.
Choose layers for a funded workflow. Do not rank SKUs in the abstract.
Staff will find a model. An authorized front door (SSO, DLP, logging, named owner) beats a block page they route around.
No personal tokens beyond local testing. Jobs on a visible runner.
NANDA 95%, McKinsey 6%, Deloitte 15%, and sponsored IDC 3.7× are different denominators. Instrument process KPIs, then argue EBIT.

AI governance is named owners, written policy, risk-tiered approvals, a living inventory, and a review cadence that continues after go-live.

Prompt engineering is how you spec, evaluate, and version LLM behavior in production. Covers CoT caveats, prompt injection, and when RAG or fine-tuning wins.

88% of organizations use AI, but only 12% of CEOs report real ROI. A function-by-function implementation guide covering tools, frameworks, and company-size play