The Arrival of the AI Operating Model

When I talk with CIOs these days, a subject that was mostly absent from serious enterprise discussions even a year or two ago is suddenly top of mind: What is our new operating model for AI?

It’s a revealing shift. For the first few years of generative AI, most organizations were understandably preoccupied with basic LLM access, pilots, use cases, governance, data protection, productivity, and just determining whether the technology was useful enough to warrant widespread deployment. Those questions have not entirely disappeared. But a much larger one is beginning to subsume them. AI is becoming sufficiently capable, economical, persistent, and autonomous that enterprises now have to decide how they will actually operate when machine intelligence becomes a load-bearing component of the business.

The timing is unusually fraught. The opportunity offered by AI presents itself at truly extraordinary speed, faster than enterprises have ever had to face, excepting perhaps 1H 2020. Yet so are the consequences of getting it wrong. Organizations face the most momentous window for intelligent automation since the arrival of computing itself, accompanied by the tantalizing ability to pursue entirely new categories of products, services, markets, and business models. At the same time, IT organizations face growing risks from unreliable autonomy, organizational dependency, loss of human capability, cyber exposure, regulatory intervention, vendor concentration, geopolitical fragmentation, and increasingly explicit warnings from the frontier AI laboratories themselves about the potential adverse behavior of much more advanced systems.

This is not a normal technology cycle.

Infographic titled 'The Arrival of the AI Operating Model' illustrating the economic impact and governance challenges of AI technologies. Left section highlights statistics on AI adoption, comparing performance between AI and human agents, with data on task success rates and cost reduction over time. Right section discusses the importance of operational models, governance maturity, and various risks associated with AI integration, including reliability issues and cognitive concentration risks.

The stakes are higher because both sides of the equation are becoming consequential at the same time. This means moving too quickly can create operational and even strategic dependencies on technology whose behavior cannot always be anticipated. Moving too slowly can (and will) leave an incumbent competing against a rapidly growing population of AI-native organizations with radically different cost structures, product development speeds, and capacity for experimentation.

The enterprise therefore needs something more substantial than an AI strategy or governance framework. It needs an AI operating model: The organizational and technical system through which it continuously determines what intelligence to use, where to use it, how much autonomy to grant it, how to verify its output, what risks to accept, how to preserve resilience, and how aggressively to exploit capabilities appearing at the frontier.

For many organizations, building that operating model is becoming urgent because AI itself is crossing a maturity threshold.

AI Is Becoming Load-Bearing

There is now a preponderance of evidence that frontier AI has moved well beyond the stage where it is useful primarily as an assistant. On a growing range of knowledge-work tasks, contemporary models are approaching or exceeding competent human performance, while agentic systems are increasingly able to carry work through applications and tools to an actual, reliable, usable outcome.

Stanford’s 2026 AI Index shows us just how quickly this has happened. On WebArena, which tests agents on 812 realistic, multi-step web tasks, success rates have risen from approximately 15% in 2023 to 74.3% in early 2026. This is only four percentage points below the measured human baseline of 78.2%. On OSWorld, which requires agents to operate real computer environments across applications and operating systems, the leading result reached 66.3% against a human baseline of 72.35%. SWE-bench Verified, a demanding benchmark based on real software-engineering issues, went from roughly 60% performance to near saturation in a single year.

This does not mean AI is universally reliable. Far from it. The same Stanford report illustrates what researchers call the jagged frontier: Systems capable of winning gold-medal-level mathematics competitions can still worryingly struggle with apparently elementary tasks. The leading evaluated system could correctly read an analog clock only 50.1% of the time. On Humanity’s Last Exam, performance jumped roughly 30 percentage points in a year, yet high-confidence errors remain common.

The resulting enterprise reality is important to state precisely: AI can now be better than humans at many bounded tasks while remaining unexpectedly unreliable at adjacent ones. It can be enormously capable without yet being conventionally dependable.

Yet the maturity trajectory is unmistakable. Anthropic reports that by May 2026 more than 80% of the code merged into its own codebase was authored by Claude, compared with low single digits before Claude Code entered research preview in February 2025. Its engineers were merging roughly eight times as much code per day as in 2024. More strikingly, Claude’s measured success rate on Anthropic’s most open-ended internal engineering tasks reached 76% in May 2026, a gain of roughly 50 percentage points in six months. Anthropic gives an example in which Claude diagnosed and fixed a live infrastructure problem in about two hours that would ordinarily have taken a human engineer two or three days.

For most enterprises, the implication is no longer difficult to see. Simply put, AI is becoming good enough to carry a meaningful share of important operational knowledge work.

Not all work. Not without controls. And certainly not with equivalent reliability everywhere. But sufficiently much of it that CIOs now have to design organizations on the assumption that AI will increasingly become part of the production machinery of the enterprise rather than merely a productivity tool sitting alongside it.

This is the arrival of load-bearing AI.

The Economics Make Adoption Difficult to Resist

Capability is only half of the force pushing enterprises across this threshold. The economics of machine intelligence are improving just as quickly.

The most useful metric is increasingly not price per token, but cost per successfully completed unit of work. Raw token prices obscure the extraordinary changes occurring beneath them. Models become more efficient. Smaller models inherit capabilities that previously required cutting-edge frontier systems. Reasoning techniques improve. Inference optimization, caching, quantization, specialization, and improved hardware reduce costs. Increasingly capable routing systems can select the least expensive model able to perform each task to the required standard.

The practical result is that intelligence itself has become a serious arbitrage opportunity. A routine extraction task might go to an extremely small model. A difficult analysis might go to a stronger reasoning system. A high-consequence decision might ultimately invoke several independent models, followed by deterministic verification and human review. It’s dynamic digital labor in a way that’s never been possible before. Organization’s can continuously match the cost and capability of intelligence to the economic value and risk of the work being performed.

That is a profound step change in enterprise economics. If useful cognition continues to become cheaper while its quality rises, activities that previously required too much human analysis, customization, research, coordination, or judgment can suddenly become economically viable. Enterprises will not merely automate existing work. They will begin creating work, services, and products that were previously too cognition-intensive to contemplate.

This leads directly to what may be the most important innovation discipline of the AI era.

From Moonshots to Starshots

The digital era taught large organizations to pursue moonshots: Occasional, high-ambition initiatives intended to create major new products, capabilities, or markets. They were appropriately uncommon because large-scale digital innovation was expensive, slow, organizationally demanding, and frequently dependent on substantial custom technology development.

AI easily remakes those economics sufficiently enough that the model itself needs to change.

I call the emerging equivalent a starshot: A high-ambition attempt to create fundamentally new business value specifically by exploiting capabilities near the frontier of artificial intelligence. Starshots are not simply larger AI projects. They ask what becomes possible when very large quantities of reasoning, research, software production, simulation, personalization, analysis, design, experimentation, and increasingly autonomous execution become inexpensive enough to embed into ordinary products and operations.

A service can place what once would have been thousands of hours of expert analysis behind every customer interaction. Software can increasingly be created or extensively modified on demand. Products can become individually generated. Scientific and engineering discovery loops can accelerate dramatically. Previously uneconomic micro-markets can become viable. Persistent agents can fundamentally redesign customer relationships. Entire operational functions can be reconceived around abundant machine cognition rather than scarce human attention.

This is where the competitive stakes become especially high. Established enterprises are not merely competing with one another. They increasingly face thousands of small AI-native disruptors able to make inexpensive attempts against assumptions that incumbents have treated as fixed for decades.

These challengers do not need to succeed every time. They just need one breakthrough.

Consequently, the AI-era enterprise cannot treat a moonshot-scale innovation exercise every year or two as sufficient. It needs a standing starshot capability capable of making faster and more frequent attempts, accepting intelligent failure, discovering nonlinear opportunities, and scaling the few that work. A portfolio might contain many modest experiments, several serious frontier explorations, and a small number of genuinely radical attempts to overturn some part of the organization’s own business model before an outsider does.

This requires operating deliberately close to the frontier. That is inherently uncomfortable for large enterprises because frontier technology is exactly where capability is highest and certainty is lowest.

Yet avoiding the frontier carries its own real risks that CIOs must ensure are managed closely and sustainably.

The AI operating model must therefore support both exploitation and controlled exploration. It has to make ordinary AI safe enough to become core operational infrastructure while simultaneously giving the organization a governed mechanism for flying considerably closer to the edge when the potential return justifies it.

The Door to an AI-Fluent Enterprise

Technology alone is insufficient to accomplish this. Before an organization can sustainably become highly automated, agentic, or AI-native, it has to pass through another threshold: It must become AI-fluent.

AI fluency is considerably more demanding than providing employees with basic training. It means building sufficient practical understanding throughout the organization that people can make competent decisions every day about what machines should do, what humans should retain, how the two should collaborate, where AI is trustworthy, when it should be challenged, and when a previously sensible arrangement must be redesigned.

This affects virtually every part of the enterprise. Business leaders need to distinguish incremental productivity opportunities from genuine changes in business economics. Managers must learn to redesign work around combinations of people and increasingly autonomous systems. Technology organizations must continuously evaluate a rapidly changing model landscape. Risk, cybersecurity, finance, HR, legal, architecture, procurement, and compliance must participate in decisions that increasingly combine technology, labor, capital allocation, and corporate authority.

More challenging still, organizations may have to undergo this cultural transition several times.

I argued a number of years ago that AI would not merely change tasks but alter organizational roles themselves, producing successive stages in which responsibilities shift from execution toward managing, interpreting, and innovating with AI. I also observed that transformation at this scale would require moving beyond a traditional centralized Center of Excellence toward a Network of Excellence: A distributed structure that connects central expertise with practitioners, champions, managers, and local leaders across the enterprise.

That concept is becoming considerably more important now. A central AI organization can establish platforms, policies, architecture, governance, standards, and expertise, but it cannot possibly absorb and direct every change required when AI begins affecting most functions and potentially most workers. The organization needs a network capable of propagating new practices, identifying failures, gathering lessons from the edge, educating employees, spreading successful patterns, and repeatedly helping local teams reorganize work as the capabilities change.

The Network of Excellence becomes the human adaptation layer of the AI operating model.

This is the door organizations must pass through. On one side are companies using AI tools. On the other are organizations capable of repeatedly reorganizing themselves around changing machine intelligence.

Those unable to cross it risk becoming marginal organizations even if they make substantial investments in AI. They may own the tools without developing the adaptive capacity required to exploit them.

The Central Tension: Bounded Acceleration

This brings the opportunity and danger together. The AI operating model has two mandates that are simultaneously essential and increasingly difficult to reconcile.

It must help the enterprise consume rapidly improving intelligence as aggressively as competitive conditions demand, including creating a repeatable starshot discipline that intentionally explores capabilities near the frontier. At the same time, it must prevent the enterprise from accumulating unacceptable operational, human, security, financial, regulatory, supplier, and systemic risks as that intelligence becomes embedded more deeply into the organization.

An operating model optimized primarily for control will probably move too slowly, and worse, it’s likely to fail. One optimized primarily for acceleration may eventually create an enterprise whose operations exceed management’s ability to understand or reliably govern them.

The objective is therefore bounded acceleration: Move as quickly as capability, economics, reversibility, observability, and verification permit, while tightening controls as autonomy, consequence, connectedness, and uncertainty increase.

This becomes especially important as AI shifts from advising people to acting for them. An incorrect answer from a chatbot is generally an information-quality problem. An incorrect decision by an agent possessing credentials, financial authority, communications access, corporate data, and permission to invoke other systems is an operational event. Once agents begin interacting with other agents, the relevant object of governance becomes even larger: Models, prompts, memory, enterprise data, tools, permissions, external services, humans, and automated feedback loops form a single socio-technical system.

The enterprise cannot make probabilistic intelligence deterministic. It can, however, surround probabilistic intelligence with deterministic controls wherever possible.

That means authoritative retrieval, independent verification, deterministic validation, policy constraints, permission boundaries, action limits, anomaly detection, observability, human escalation, rollback, audit trails, and kill mechanisms. At higher levels of consequence, several of these mechanisms should operate simultaneously.

This is not mere governance added after automation. It is what allows consequential automation to occur at all.

The Downside Is Becoming Material

Reliability is the most obvious risk, but it is far from the only one. The jagged capability frontier means that even extremely sophisticated systems can fail unpredictably, and increasing capability does not necessarily eliminate this problem. More capable models will simply be entrusted with harder work, perpetually moving the enterprise back toward the edge of what machines can reliably accomplish.

Autonomy quickly magnifies the consequences of those errors because failure moves from generating incorrect information to taking incorrect action. Cybersecurity risk rises as agents receive credentials, APIs, data access, tools, and the ability to communicate externally. A compromised or misdirected machine worker potentially operates at machine speed and scale.

There is also an underappreciated human risk. Successful automation can progressively remove not merely jobs but organizational knowledge. Manual procedures disappear. People who understood why processes worked leave. Entry-level roles through which future experts acquired experience can vanish. Human situational awareness declines as machines perform more of the intermediate work.

The result can be an extraordinarily efficient organization that is surprisingly unable to function without its AI.

I believe enterprises should begin treating this explicitly as cognitive concentration risk. We already manage concentration risk in cloud infrastructure, telecommunications, financial services, semiconductor supply chains, and critical vendors. Increasingly, we will have to manage dependence on externally supplied cognition in much the same way.

Provider concentration compounds the problem. A relatively small group of companies produces much of the world’s frontier intelligence, while the physical infrastructure beneath it has its own geographic and supplier concentrations. Stanford notes that nearly all leading-edge AI chips still depend upon a single Taiwanese foundry, illustrating how seemingly abstract machine intelligence ultimately depends on very tangible industrial infrastructure.

A major model can therefore become unavailable or unsuitable for reasons having little to do with enterprise architecture: Provider distress, pricing changes, model retirement, compute shortages, cyber incidents, geopolitical intervention, regulatory decisions, litigation, safety restrictions, or infrastructure failures. Once AI becomes load-bearing, these become business-continuity scenarios.

There is also the problem of governance lag. Technical capability is changing faster than organizational controls, management practices, workforce skills, laws, and social expectations can comfortably absorb. That gap will generate recurring friction. Regulatory environments will differ across jurisdictions. Certain models, data uses, autonomous actions, or decisions may be permissible in one country but prohibited in another. Enterprise AI routing consequently becomes not merely an economic mechanism but potentially an increasingly important compliance mechanism.

Perhaps most difficult is the possibility of organizational exhaustion. Enterprises are accustomed to large transformation programs separated by periods of relative stability. AI may instead require significant changes in roles, processes, controls, incentives, skills, and operating structures repeatedly over a relatively short period. Without distributed change capacity, enterprises may simply lose the ability to assimilate what technology makes possible.

And then there is the frontier itself.

The Coming AI Event Horizon

The leading AI laboratories are now discussing risks that would have seemed extraordinary in an enterprise technology article only a few years ago. Anthropic, OpenAI, and Google DeepMind all explicitly examine scenarios involving advanced autonomy, loss of control, autonomous AI research, or systems materially accelerating the development of more capable AI.

Anthropic’s internal experience is particularly instructive. More than 80% of the code it merged by May 2026 was AI-authored, while the typical engineer was merging roughly eight times as much code per day as in 2024. Anthropic emphasizes that this is not yet recursive self-improvement and that such an outcome is not inevitable. But it also states plainly that extending the trend could eventually produce an AI capable of autonomously designing and developing its successor.

That possibility changes long-range planning because the variables begin to interact. Better AI can improve AI research. Improved research creates better AI, which can then contribute still more effectively to subsequent research. Compute, energy, semiconductor fabrication, physical experimentation, capital, and other real-world constraints may keep this feedback loop bounded. We simply do not yet know.

For CIOs, however, the practical concept matters even before anything resembling runaway recursive improvement occurs.

There is an AI event horizon when the rate at which useful machine intelligence changes becomes faster than the organization’s ability to understand, govern, and adapt to it.

An enterprise does not need to encounter superintelligence to cross this threshold. If relevant capabilities double or transform faster than architecture cycles, procurement processes, workforce adaptation, regulation, and management practices can respond, conventional three- or five-year AI roadmaps become increasingly speculative.

This is another reason the operating model itself must be dynamic.

Four Hard Choices

The job of the AI operating model is ultimately to help leadership navigate four increasingly stark tradeoffs rather than pretending that they can be eliminated.

The first is Transformation or Marginalization. Enterprises can use AI principally to make the existing organization faster and cheaper, or they can continually use frontier capability to challenge their own products, services, processes, and economics. Productivity matters enormously, but organizations that never build a serious starshot capability increasingly risk being attacked by companies that have. This is not a call for indiscriminate disruption. It is recognition that the cost and speed of attempted disruption are falling sharply, which means incumbents must increase their own rate of meaningful experimentation.

The second is Delegation or Sovereignty. Organizations can give machines progressively more responsibility for execution and decision-making, capturing enormous economic benefit, or maintain meaningful human authority over consequential activities. Too little delegation eventually leaves much of AI’s value unrealized. Too much creates a business whose executives remain accountable for systems they no longer genuinely understand or control. The objective should be maximum economically useful autonomy consistent with verified control.

The third is Dependence or Optionality. Tight integration with a small number of leading AI providers may offer excellent capability, economics, and simplicity. Preserving multiple providers, model portability, open systems, fallback procedures, and human expertise adds cost and complexity. But once intelligence becomes operational infrastructure, optionality becomes a form of resilience. The inefficiency looks unnecessary until the primary system is unavailable.

The fourth is Acceleration or Survival. There will be moments when a new capability justifies moving surprisingly quickly and occasions when an organization should deliberately stop. A security incident, reliability regression, regulatory action, geopolitical event, provider failure, or unexpected autonomous behavior could require an immediate reduction in AI authority. A mature enterprise needs the operational ability to accelerate toward opportunity and decelerate away from danger without rebuilding its entire architecture each time.

In other words, the enterprise needs both an accelerator and a brake, and increasingly sophisticated judgment about when to use each.

The Architecture of the AI Operating Model

The resulting operating model is not another governance committee. It is a persistent enterprise capability for controlling and exploiting machine intelligence.

At its core is an intelligence control plane. Significant AI workloads need an explicit business purpose, approved model set, required capability level, acceptable cost, data-access envelope, tool permissions, autonomy ceiling, verification requirements, human escalation path, fallback model, and continuity plan. These parameters should increasingly be adjustable as models, economics, regulations, and risk conditions change.

An assurance layer surrounds consequential AI with independent verification, deterministic validation, authoritative data, monitoring, constraints, action controls, and escalation. A resilience layer ensures that critical functions can degrade gracefully if a model or provider disappears and that enough human knowledge survives to maintain organizational sovereignty.

The Network of Excellence provides the human adaptation layer, distributing learning and organizational change at a speed no centralized AI group can achieve on its own. The FinOps and routing layer continuously arbitrages between models and methods to obtain the least expensive intelligence that can produce the required verified result.

And the operating model needs one more first-class function: A starshot portfolio.

Infographic titled 'The Arrival of the AI Operating Model', depicting a dynamic enterprise system for AI with a focus on governance, risk, and autonomy.

This should not sit on the periphery of the AI program. It should be an explicit mechanism through which the enterprise continuously tests whether new frontier capability has invalidated an existing constraint, opened a new market, enabled a new product, or made an apparently impossible operating model economically feasible.

Governance and starshots belong inside the same system precisely because the organization must take more risk in some places in order to remain conservative in others. A sandboxed starshot with carefully bounded data, capital, authority, customers, and blast radius can explore the frontier aggressively without granting experimental technology equivalent authority over core production operations.

This is how a large enterprise can learn to fly close to the edge without betting the company every time.

A Different Operating Cadence

Traditional enterprise technology was built around periods of stability. Select a platform, implement it, standardize it, optimize it, operate it for several years, and eventually replace it.

AI is increasingly incompatible with that cadence. The emerging loop is continuous: Sense new capabilities, evaluate them, experiment, route work to appropriate intelligence, execute, verify, observe results, adjust authority, distribute learning, retire obsolete assumptions, and repeat.

The starshot portfolio runs beside this cycle asking one especially important question:

What has become possible now that was impossible six months ago?

That deserves to become a standing executive question because annual strategy cycles will increasingly miss meaningful portions of the frontier.

The human organization must operate at a similar rhythm. People learn new capabilities, redesign work, discover local practices, share them through the network, adapt roles and controls, and then prepare to do it again. AI fluency is therefore not an educational endpoint. It is the institutional ability to keep learning as the underlying intelligence changes.

There May Be No Final AI Transformation

One of the hardest ideas for enterprises to absorb may be that there is no stable future state waiting at the end of an AI transformation program.

Transformation is traditionally imagined as a bridge between current and future operations. AI increasingly resembles a changing environment instead. Organizations may have to substantially redesign how work is allocated among people and machines several times as models become more capable, persistent, inexpensive, autonomous, and connected.

This is why all of these threads belong in one operating model. AI fluency allows the enterprise to understand what is changing. A Network of Excellence allows it to adapt at organizational scale. Dynamic routing allows it to exploit changing capability and economics. Assurance makes increasingly consequential probabilistic systems usable. Resilience prevents successful automation from becoming dangerous dependence. Starshots ensure that safety and scale do not turn into strategic timidity. Frontier governance determines how much authority the organization is prepared to grant as the technology approaches increasingly uncertain territory.

Together they create something more important than an AI platform. They create organizational adaptive capacity.

The Path Forward

I remain optimistic about our ability to build this. Human beings have repeatedly created institutions capable of operating technologies and systems vastly more complex than any individual can understand. Aviation, electrical grids, global financial systems, telecommunications, supply chains, hyperscale computing, and the Internet became dependable not because uncertainty disappeared, but because we surrounded them with professional disciplines, redundancy, monitoring, controls, standards, training, and cultures capable of managing their risks.

AI can be treated with the same seriousness. The difference is that we may have to build those mechanisms while the underlying technology continues changing at exceptional speed.

The enterprises that succeed will therefore not necessarily be those with today’s best model, the largest AI budget, or the highest percentage of automated tasks. They will be organizations that can repeatedly absorb new intelligence without surrendering judgment; automate aggressively without becoming brittle; preserve human and architectural optionality; distribute AI fluency throughout the enterprise; recognize when risk requires a brake; and continuously make enough ambitious starshot attempts to discover what the new frontier makes possible before competitors do.

That is a demanding operating model, and the stakes are higher than in most previous technology transitions. An enterprise can now plausibly move too quickly and create profound new vulnerabilities. It can also move too slowly and find that its economics, products, or even its reason for existing have been overtaken by organizations built around a fundamentally different abundance of intelligence.

The goal is therefore neither unrestrained acceleration nor defensive caution. It is to create an enterprise capable of repeatedly approaching the frontier, extracting disproportionate value from it, and returning safely enough to do it again.

Many organizations will make it through this door. Some will become radically more capable than they are today, combining human judgment with machine intelligence at a scale that was previously impossible. They will automate much of the ordinary work of the enterprise while redirecting considerably more human energy toward invention, relationships, judgment, leadership, and ambitious new outcomes.

But some organizations will not make the transition. They will adopt AI without becoming AI-fluent, automate without building resilience, govern without innovating, experiment without scaling, or protect today’s business so thoroughly that they leave tomorrow’s business to somebody else.

For CIOs, this is why the AI operating model has moved so quickly to the center of the agenda.

The downside of moving too quickly is increasingly real, yet the downside of moving too slowly may ultimately be larger. The task now is to build an organization capable of knowing the difference—and changing its answer continuously as the frontier continues to move swiftly.

Leave a comment