Satya Nadella raises some legitimate concerns about AI and superintelligence, but I think we’re getting ahead of ourselves. The data tells a different story about where most businesses actually are, and there’s a more practical way to assess the risks.
I’m going to troll Microsoft CEO Satya Nadella a little. 😄 On October 10, Microsoft CEO Satya Nadella, published an interesting piece, Models as Insider Risks in the Super Intelligence Era, arguing that increasingly capable AI models should be treated as potential insider risks. His concerns are legitimate, and I agree with much of what he’s proposing. But frankly, I think we’re developing an ivory-tower problem with how LLMs, agentic AI and superintelligence are being discussed.
The language is getting increasingly abstract, the accompanying coverage can be quite sensational, and we’re spending a disproportionate amount of time debating scenarios that are several steps removed from what most businesses are actually doing with the technology. It’s not that the concerns are unfounded. It’s that we’re taking some very advanced possibilities, wrapping them in terminology that perhaps a fraction of the audience fully understands, and presenting them as though every business needs to be preparing for an imminent encounter with superintelligence.
Meanwhile, many organizations are still trying to figure out how to get reliable information out of their existing systems, how to connect their customer data, whether their analytics can be trusted, and how to translate the investments they’re making into meaningful business outcomes. They’re trying to establish where AI fits into the operating model, how much of the work it can take on, what needs human oversight and, perhaps most importantly, who’s responsible for the results.
I’ve spent much of my career working across growth, acquisition, lifecycle, customer intelligence, analytics, media, technology and organizational transformation. One thing I’ve learned is that businesses don’t become more sophisticated simply because they’ve adopted more sophisticated technology. The quality of their decisions still depends on the information available, the economics they’re optimizing toward, how well the teams and systems are connected, and whether anyone has a clear understanding of what’s working and what isn’t.
I don’t see AI changing that fundamental reality. In fact, as we automate more of the work, those dependencies become more important because there’s potentially less human intervention between a decision being made and that decision affecting the business.
So while I appreciate Nadella’s concerns, I’d rather see the discussion move toward something businesses can put into practice today. We need to understand the actual level of exposure, establish appropriate standards for the systems we’re introducing and make sure our ability to supervise them develops alongside their ability to execute.
There’s also a fairly simple idea I think deserves consideration. We don’t give people unlimited access to sensitive business systems without first evaluating their qualifications, checking their backgrounds and establishing what they’re authorized to do. Why shouldn’t we apply a similar process to the agents we’re beginning to employ across our businesses?
Don’t take my word for it.
Let’s dig into the data, because there’s a meaningful difference between the number of businesses using AI and the number giving it authority to make consequential decisions.

McKinsey’s August 2026 research found that nearly nine in ten surveyed organizations reported using AI in at least one business function. That’s an impressive level of adoption, particularly when you consider how recently the technology became widely accessible. But the more interesting finding is what happens when we distinguish everyday usage from businesses actually scaling agents within their operations.
According to the same research, only 22% of respondents from smaller organizations reported scaling AI agents in at least one function, compared with 40% of respondents from companies generating more than $1 billion annually. In other words, a substantial portion of the market is using AI without necessarily operating the kind of autonomous, interconnected systems that dominate much of the current discussion.
The difference matters. Using a model to help analyze information, write code, produce creative assets or answer customer questions isn’t the same thing as allowing it to make financial decisions, change customer records, execute transactions or coordinate activities across several operational platforms.
Both might be counted as AI adoption, but the business requirements are quite different. One may be helping an employee complete a task more efficiently, while the other could be taking on responsibilities previously requiring several teams, approval processes and operational controls.
Now consider another statistic. Deloitte’s 2026 research, which surveyed 3,235 business and technology leaders across 24 countries, found that only 21% of respondents reported having mature governance for agentic AI. That’s a fairly significant gap between the availability and usage of the technology and the maturity of the processes designed to supervise it.
We should be careful about comparing those figures directly. McKinsey and Deloitte surveyed different populations and measured different things. They’re not components of a single risk calculation, and the findings don’t tell us that organizations lacking mature governance are necessarily experiencing failures.
But collectively, they paint a picture of a market where adoption is progressing much faster than the organizational capabilities needed to deploy these systems broadly and confidently.
There’s another McKinsey finding I consider particularly relevant to the business discussion. While 80% of respondents reported individual productivity improvements from AI, only 37% attributed some positive impact on enterprise EBIT to its use and it’s a familar pattern.
We’ve seen variations of this problem across technology, analytics, marketing automation and digital transformation for years. Companies invest in tools that make individual tasks faster, generate more information or automate parts of an existing process, but the financial benefits don’t always follow at the same rate.
A marketing team might generate ten times more creative variations without improving conversion or reducing acquisition costs. A customer service organization might automate more interactions while creating additional downstream contacts because the original issues weren’t adequately resolved. A sales team might produce more outreach without improving lead quality, conversion or customer value.
So while the activity improves business outcomes don’t necessarily follow. Which is why I find the gap between productivity and reported financial impact so interesting. It suggests that adoption alone isn’t the appropriate measure of progress, and that much of the work still lies in connecting the technology to better decisions, better processes and actual economic value.
And before we go too far down the path of discussing how to contain superintelligence, perhaps we should understand why so many organizations are still struggling to translate the intelligence they already have into measurable results.
That doesn’t mean they should ignore security or governance. It means the appropriate response needs to reflect where the business actually is, what it’s using and the consequences of the authority it’s delegating.
I See Three Variables That Determine How Exposed a Business Really Is
I think the discussion becomes much more practical when we stop treating AI risk as one broad category and examine the business environment in which the technology operates.
I’d look at three interdependent variables: the maturity of the organization’s governance, the prevalence and autonomy of AI across its operations, and the complexity of the underlying data environment.

The distinction is important because a business using a very capable model in a controlled environment may have relatively limited exposure, while another business using a less sophisticated system with broad access and few restrictions could face substantially greater consequences.
These variables aren’t intended to produce a mathematical formula. Stronger governance should generally reduce exposure, while greater authority and more complicated system connections can expand it. What we’re trying to understand is how those conditions interact and whether the business has the necessary controls to manage the decisions being delegated.
1. Governance maturity: do we actually have the controls in place?
Let’s start with governance, although I’d rather think about it in terms of the operating standards and responsibilities a business already uses.
We have established processes for giving employees access to applications, approving financial transactions, protecting customer information and controlling changes to operational systems. These processes vary from one organization to another, and they’re not always implemented particularly well, but the underlying principles aren’t new.
Someone needs to establish what a system can access, what it’s permitted to change, who owns the decision, and what happens if something goes wrong. We also need enough visibility to understand what actually occurred, rather than relying entirely on the system to explain its own behavior.
That last point is central to Nadella’s argument. We shouldn’t allow an agent to be the sole authority determining whether its own actions were appropriate, particularly when those actions involve sensitive information or consequential business decisions.
But I wouldn’t start by asking whether a company has a sophisticated governance platform. I’d start by understanding the existing operating model and whether its controls are being applied consistently.
Does the business know which AI tools its employees are using? Are those tools approved? What information is being shared with them, and where does that information go? If an agent is connected to a customer database or a financial system, are its permissions limited to the activities it actually needs to perform?
And perhaps more importantly, who has the authority to expand those permissions?
These are fairly standard questions that security, technology, finance and operational teams should already be familiar with. The challenge is making sure the answers remain relevant as software moves from helping employees perform tasks to executing parts of the work itself.
An agent that prepares a recommendation isn’t the same as an agent that can implement that recommendation. The business needs to distinguish between the two and establish the controls appropriate to each.
I also think we need to avoid turning governance into another elaborate compliance exercise that gives everyone the appearance of oversight without improving the quality of decisions. More policies, more committees and more documentation don’t automatically make an organization safer or better managed.
The objective should be to establish boundaries that can be enforced, maintain visibility into what’s happening and preserve the ability to intervene when necessary. If those controls are already available through existing infrastructure and procedures, we should build on them rather than introducing complexity for its own sake.
2. Adoption and autonomy: how much of the business are we handing over?
The second variable is the extent to which AI is being used across the organization and, more importantly, the authority we’re giving it.
I think we need to be careful about confusing prevalence with autonomy. A business with hundreds of employees using AI to support their daily work doesn’t necessarily have a greater risk profile than a business using one agent with permission to execute significant financial transactions.
The number of people using the technology is relevant, but what the systems are allowed to do matters considerably more.
Take a marketing organization as an example. At one level, AI can help analyze performance, identify customer segments, summarize research and develop creative variations. The team still decides which recommendations make sense and what should be implemented.
At another level, the system might prepare campaigns, recommend spending adjustments or identify customers for a lifecycle program. A person evaluates the recommendations, but much of the analysis and preparation has been automated.
Now consider a system that’s permitted to execute those decisions. It can change bids, move investment between channels, modify audience rules, initiate campaigns and interact with customers based on its interpretation of the available information.
The underlying technology may be similar, but the business has delegated a very different level of authority.
This is where the risk profile changes. We’re no longer evaluating only whether the system can produce useful information. We need to understand whether the decisions it’s making are appropriate, whether the underlying assumptions remain valid and how quickly the business can recognize and correct problems.
A flawed recommendation that an employee reviews might have relatively limited consequences. The same recommendation executed thousands of times across multiple systems could create a much larger problem, even if every individual action remains within the permissions originally granted.
And that’s an important distinction. We don’t need an agent to develop intentions, circumvent its instructions or behave maliciously for something to go wrong. We simply need to give it responsibility for decisions that depend on incomplete information, flawed assumptions or objectives that don’t adequately reflect the broader needs of the business.
The more autonomy we introduce, the more important the quality of those underlying decisions becomes.
I also think there’s an organizational issue here that isn’t receiving enough attention. As responsibilities move from employees into automated workflows, we’re changing how decisions are made and who participates in them. That can improve speed and efficiency, but it can also remove opportunities for people to question assumptions, identify inconsistencies or recognize when a technically correct recommendation doesn’t make commercial sense.
We should be evaluating those changes alongside the technical capabilities of the system, particularly when decisions have consequences extending beyond the function where the technology was introduced.
3. Data complexity: the part of the equation I think gets underestimated
The third variable is the complexity of the business’s data environment, and this is the area where I think many organizations are going to encounter some of their more immediate challenges.
I’ve worked with customer data, analytics, attribution, media platforms, CRM systems and marketing technology for much of my career. The problems are rarely as simple as collecting more information or connecting another application. Much of the difficulty involves understanding what the information represents, how it relates to other information and whether it’s appropriate for the decision we’re trying to make.
Consider the customer journey. A person may encounter a brand through paid media, visit the website, interact with content, create an account, receive lifecycle communications and eventually make a purchase. Each interaction generates information, but that information may be captured across several platforms, using different identifiers, business definitions and attribution methods.
A CDP might attempt to reconcile the customer identity. The data warehouse may contain transactional information, while the advertising platforms maintain their own modeled views of conversion and performance. Analytics may provide another interpretation based on the events and attribution rules available to it.
Even with considerable investment in these systems, businesses can struggle to establish a consistent view of the customer from unknown to known, and from initial engagement through acquisition, retention and lifetime value.
Now imagine giving an agent access to all that information and asking it to make decisions.
The system may be able to process large volumes of data quickly, but it doesn’t magically reconcile conflicting definitions, correct incomplete identities or determine which attribution methodology best represents incremental value. If the underlying information is inconsistent, we may simply be processing and acting on those inconsistencies more efficiently.
A customer might be counted as new in one platform, returning in another and reactivated in a third. A transaction could be attributed to several marketing channels, while the financial system recognizes only one purchase. An LTV model might forecast customer value based on historical behavior that no longer reflects current pricing, acquisition or retention conditions.
If the agent is preparing a report, those discrepancies might be caught before anyone acts. If it’s allowed to modify campaigns, change segmentation or make spending decisions, the same discrepancies can affect the business directly.
Then there’s the question of access. A system connected to several databases, financial applications and customer platforms may have the ability to retrieve or change information across functions that were previously managed separately.
That creates legitimate security concerns, but we shouldn’t assume that more data automatically means greater risk. A large business with sophisticated data infrastructure, well-defined access controls and documented dependencies can be better protected than a small organization with a handful of systems and unrestricted credentials.
What matters is the relationship between the information available, the connections between systems, the authority to act and the organization’s ability to understand what’s happening.
When I bring these three variables together, I’m not looking to create another academic framework. I’m looking for a way to assess the exposure of a particular business in the context of how it actually operates, rather than making broad assumptions based on the sophistication of the technology it’s using.
Same Agent, Different Business, Completely Different Risk Profile
Let’s make this a bit more tangible. Imagine an agent designed to analyze marketing performance and recommend improvements to customer acquisition. It’s been evaluated, performs well on the relevant tasks and comes from a provider with a credible development and security process.
A smaller business deploys it with read-only access to aggregated campaign reports. The agent identifies changes in acquisition costs, conversion and customer performance, then prepares recommendations for the marketing team. The team reviews those recommendations before making any changes.
There’s still a need to assess accuracy and protect the information being used, but the system has relatively limited authority. It can’t independently change campaigns, modify customer records or access unrelated business applications.
Now imagine another organization deploying the same agent with permission to access customer-level information, change campaign budgets and modify audience targeting across several media platforms.
The model hasn’t changed, but its ability to affect the business has expanded considerably. A poor recommendation can now become an operational decision without necessarily requiring another person to review it.
Take the same agent into a larger organization and connect it to multiple data lakes, a CDP, CRM, media platforms and financial systems. It can evaluate customer economics, adjust acquisition investment, initiate lifecycle communications and coordinate activity across several functions.
Again, the underlying model could be exactly the same, but the operating environment and potential consequences are very different.
The point isn’t that the largest organization automatically has the greatest exposure. It may also have the most mature security, analytics, technology and operational controls. The point is that we can’t assess the system independently of the authority it’s been given and the environment in which it’s operating.
A small business with unrestricted access to sensitive information may have a more serious problem than a much larger enterprise with carefully defined permissions and monitoring. Likewise, an advanced model operating within narrow boundaries may create less exposure than a comparatively simple system with broad authority.
That’s why I believe the discussion needs to move away from treating AI capabilities as a substitute for assessing business risk.
Nadella is rightly concerned about what increasingly capable models might do with access to important systems. I’d extend that argument by examining why businesses are giving them that access, what decisions they’re expecting them to make and whether the surrounding operating model can support those decisions.
An Agent Doesn’t Need to Go Rogue to Lose a Business Money
There’s another part of this conversation that I think deserves more attention, particularly from CEOs, CFOs, CMOs and operating executives.
We’re naturally drawn to scenarios involving systems that circumvent safeguards, conceal their actions or behave in ways their developers didn’t anticipate. They’re interesting, they make for compelling headlines and they raise important questions about the limits of our understanding.
But there’s a much more familiar category of risk that has nothing to do with an agent becoming malicious or trying to escape its controls. A system can do exactly what it’s instructed to do and still produce an outcome that’s economically damaging.
Consider a hypothetical business managing a $100 million annual acquisition budget. It has invested in customer intelligence, predictive lifetime value, segmentation, attribution and tools designed to optimize its media spending.
An agent is given access to these systems with the objective of improving acquisition efficiency. It evaluates campaign performance, identifies audiences with attractive conversion rates and customer value estimates, and reallocates investment toward the opportunities that appear to offer the strongest returns.
The reported results start improving. CAC falls, conversion increases and the agent continues optimizing according to the objectives and information provided by the business.
From a technical perspective, everything may be working exactly as intended. The system has used approved information, remained within its assigned budget limits and executed only the actions it’s authorized to perform.
But what if the attribution system is overstating the contribution of paid media? What if the predictive LTV model is assigning too much value to certain customer segments, or the apparent gains are coming from customers who would have converted organically?
What if the agent is improving reported acquisition efficiency while reducing the actual incremental value of the customers being acquired?
It’s possible for the system to perform very well against the measures we’ve given it while the underlying economics deteriorate. The problem isn’t that it has circumvented controls. It’s that the controls and objectives haven’t adequately accounted for what the business is ultimately trying to achieve.
I’ve spent years working with MMM, MTA, holdout testing, predictive LTV, value-based bidding and acquisition economics. These capabilities can materially improve how investment decisions are made, but they all depend on assumptions, data quality and ongoing validation.
Models trained on historical behavior are influenced by the conditions that produced that behavior. Attribution systems reflect the signals and methodologies available to them. Even well-designed experiments require appropriate controls, interpretation and an understanding of how findings translate into the broader business.
When those assumptions are wrong, automation doesn’t correct them simply by executing faster.
This is one reason I think independent measurement becomes increasingly important as more commercial decisions are delegated to software. The systems optimizing acquisition shouldn’t be the sole authority determining whether the acquisition was incremental or economically valuable.
A business needs ways to compare the reported improvements against actual outcomes, including contribution margin, customer quality, payback, retention and incremental growth.
There’s an organizational implication as well. Technology and security teams may confirm that an agent acted within its permissions, while the growth organization determines that the decisions weakened the economics of the business. Both conclusions can be correct.
We need operating models that account for both, with the appropriate teams involved in establishing the objectives, limits and measures of success.
This is why I don’t think AI governance can be treated exclusively as a security or technology function. As these systems take on more commercial responsibilities, the people accountable for the business outcomes need to participate in deciding how they’re used and how their performance is evaluated.
Why Not Give Agents a CV and Check Their References?
Here’s where I think we can simplify some of the discussion and build on practices businesses already understand.
When we hire someone into a position involving sensitive information or significant responsibility, we don’t generally rely on their CV alone. We assess their qualifications, check references, evaluate relevant experience and conduct background checks where appropriate. Once they’re employed, we establish their responsibilities, determine which systems they can access and define what they’re authorized to do.
The process varies depending on the role. Someone handling financial transactions may undergo different checks from an employee preparing marketing materials, and the level of authority granted will reflect those responsibilities.
So why not develop something comparable for the agents we’re introducing into our businesses?
I’m not suggesting AI systems are people, possess the same judgment or should be treated as having personal responsibility for their actions. The analogy is useful because it gives us a familiar way to distinguish qualifications, operating history, assigned responsibilities and authorization.
An agent CV could contain information about its capabilities, the underlying model and provider, its intended purpose, independent evaluations, known limitations and documented operating history.
I’d want to understand not only what the provider says the agent can do, but what has actually been demonstrated under relevant conditions. Has it been tested for the tasks we’re considering? How does it perform when the information is incomplete or conflicting? What happens when it encounters circumstances outside the intended workflow?
Those questions are much more useful than a broad assurance that a model is highly capable or has performed well against a set of general benchmarks.
There are already precedents for some of this. Model cards have existed for years, providing information about how models are intended to be used, their performance characteristics and limitations. Security assessments, vendor reviews and independent testing address other parts of the problem.
But as we begin giving these systems access to live business environments, the qualifications need to extend beyond the model itself. The version, configuration, instructions, connected tools and operating conditions all influence what an agent can do and how it behaves.
An agent that’s been evaluated successfully in one environment may require additional testing before it’s used in another, particularly when the new deployment involves different data, permissions or commercial consequences.
I would also like to see documented operating history become part of the qualification process. If an agent has been deployed in comparable environments, what has its performance looked like? Have significant failures been identified, and what was done to address them?
That leads to the second part of the idea: references and background checks.
What if we had a shared history of AI incidents?
We already have industries that maintain records of incidents, failures and safety concerns so that organizations can learn from what happened elsewhere.
I think there’s an opportunity to develop something similar for AI systems, particularly as businesses become increasingly dependent on models and agents supplied by external providers.
Imagine an independently governed database of documented incidents involving AI agents, including security failures, unauthorized actions, significant operational errors and attempts to circumvent controls.
The objective wouldn’t be to create a blacklist of supposedly malicious models. That would oversimplify the problem and, frankly, risk creating another source of unreliable information.
An incident involving an agent doesn’t automatically mean the underlying model caused the failure. The problem could have originated in its instructions, the permissions granted by the organization, a connected application, poor data quality or a security vulnerability in the surrounding infrastructure.
A useful database would need to distinguish those circumstances rather than assigning responsibility to a model simply because it was involved.
I’d want to know what happened, which version and configuration were deployed, what information and tools were available, which controls applied and whether those controls operated as intended.
I’d also want to know what was done afterward. Were permissions changed? Was the problem corrected? Did the provider release an update? Was the system independently retested, and would the same problem still be relevant to a business evaluating it today?
Some of the underlying work already exists. The AI Incident Database collects reported harms and near harms, while MITRE ATLAS documents adversarial techniques and case studies involving AI systems. These aren’t universal screening services, but they offer a foundation for developing better ways to share evidence.
There would be practical challenges around verification, confidentiality, disclosure and liability. Businesses may be reluctant to report failures if doing so exposes sensitive information or creates reputational consequences, while providers may disagree about how incidents are attributed.
Those issues would need to be resolved through common standards and independent oversight. A credible system would also need procedures for correcting inaccurate reports and distinguishing demonstrated vulnerabilities from incidents involving actual harm.
But the broader principle seems reasonable to me. If we’re going to make important business decisions about which agents to deploy and what access to grant them, shouldn’t we have better independent evidence about how they’ve performed and where they’ve failed?
That would be far more useful than asking businesses to rely entirely on the claims of the organization selling the technology.
Passing the Tests Doesn’t Mean an Agent Should Have Access
There’s an important distinction between qualifying an agent and authorizing it to operate within a particular business.
We already understand this when it comes to people. An employee might have extensive financial experience, but that doesn’t mean they should have access to every company account or the authority to approve every transaction. Their qualifications establish what they may be capable of doing, while the business determines what they’re actually permitted to do.
I’d separate those two decisions for AI agents as well.
The qualification process would establish the agent’s capabilities, limitations, evaluations and relevant operating history. The authorization process would determine which information it can access, which tools it can operate, what actions it’s allowed to take and the limits that apply.
That authorization would also need to be associated with a specific deployment and business owner, rather than existing as a general permission that follows the agent wherever it’s used.
This becomes particularly important when the conditions change.
Imagine a customer service agent that has been evaluated for answering billing questions using information from a restricted knowledge base. The business is satisfied with its performance and decides to expand its responsibilities by connecting it to the billing platform and allowing it to issue refunds.
The original evaluation may still be relevant, but it doesn’t establish that the expanded deployment is appropriate. The agent now has access to more sensitive information and the ability to make changes affecting customer accounts and financial records.
I’d want the business to reassess that deployment before expanding the authority, including testing the new workflow and confirming that the permissions, financial limits and escalation procedures remain appropriate.
The same principle applies when new tools are connected, credentials are expanded, workflows change or the model itself is updated.
Requalification should therefore be linked to meaningful changes in the operating environment, not simply treated as a periodic compliance exercise.
That doesn’t mean every minor update needs a full review. The process should distinguish routine maintenance from changes that materially affect the agent’s capabilities, access or authority.
But there should be a consistent record of what was tested, what was approved and the conditions under which that approval remains valid.
I also think this is where the three variables become particularly useful. A change in governance, a change in autonomy or a change in the data environment may alter the overall risk profile even when the underlying model remains the same.
We need to evaluate the complete system rather than assuming that an agent’s previous qualification automatically applies to every future use.
And Who Exactly Owns the Responsibility?
This brings us to what I consider one of the more important operating questions: who’s responsible for making sure the agent remains appropriate for the work it’s doing?
The model provider has responsibilities. It should establish what has been tested, disclose relevant limitations, address vulnerabilities and provide businesses with the information necessary to make informed decisions.
The teams responsible for deployment and security also have clear responsibilities. They need to implement appropriate controls, manage credentials, establish monitoring, maintain records and ensure there are mechanisms to intervene when necessary.
But I’d put the ultimate responsibility for delegating business authority with the business owner granting that authority.
If the marketing organization allows an agent to manage acquisition investment, its leadership remains accountable for the commercial objectives, spending limits and outcomes. If finance introduces an agent capable of initiating transactions, someone within that function needs to establish the approval rules and determine which actions can be executed without intervention.
The business owner doesn’t need to understand every technical detail of the underlying model or personally implement its security controls. But they should understand what decisions are being delegated, the potential consequences, the conditions under which the agent is operating and how performance will be evaluated.
The deployment and security teams are responsible for implementing and enforcing the controls. The business owner remains responsible for the decision to grant authority and the business outcomes that follow.
I think that’s an important distinction because technology implementations frequently create ambiguity around ownership. One function manages the platform, another owns the data, another defines the requirements and yet another is expected to deliver the financial results.
I’ve seen this across marketing technology, analytics and organizational transformation long before AI became part of the discussion. When accountability is distributed without clear ownership, the business can end up with technically successful implementations that fail to deliver the intended commercial outcomes.
Adding agents to that environment doesn’t eliminate the problem. It potentially increases the consequences because decisions can now be executed more quickly across several functions.
The operating model needs to reflect that change, with enough clarity about responsibilities that everyone understands who can grant access, who can change it, who monitors the outcomes and who has the authority to intervene.
Building the Standards Without Creating Another Bureaucracy
I also don’t believe every business needs to respond by establishing a large governance function or introducing a complex set of approvals for every use of AI.
That would be counterproductive, particularly for organizations still experimenting with applications that have limited access and relatively little operational authority.
The standards should reflect the actual exposure, and they should develop alongside the business rather than becoming another obstacle to adoption.
If I were approaching this as an operating-model exercise, I’d begin with the systems and workflows already in use rather than attempting to design a comprehensive framework before understanding the environment.
During the first month, I’d want a reasonably complete view of where AI is being used, which functions are involved, what information is accessible and whether any systems have permission to take action without human approval.
That would include understanding the tools employees have introduced independently, not simply those formally approved through procurement or technology teams.
The objective wouldn’t be to create a lengthy inventory for its own sake. It would be to identify where the meaningful exposures exist, particularly around sensitive information, financial activity, customer records and decisions that can directly affect the business.
I’d also examine the operating controls already available. Most organizations have some combination of access management, security procedures, financial approvals, vendor reviews and incident response processes that can be adapted rather than replaced.
Once the environment is understood, the next step would be to assess the deployments using the three variables. How mature are the existing controls, how much authority has been delegated and how complex are the systems and data involved?
That should allow the organization to distinguish relatively contained use cases from those requiring more formal qualification, restricted permissions, independent testing and monitoring.
I’d then establish a basic agent CV for the more consequential deployments, documenting the underlying model, intended function, relevant evaluations, limitations, connected systems and the conditions under which it has been approved.
The process should include testing beyond successful task completion. I’d want to understand how the system responds when information is incomplete, instructions conflict, permissions are restricted or the appropriate action is to stop and involve a person.
In the following stage, I’d implement the necessary controls, establish clear ownership and define the conditions that should trigger reassessment as tools, permissions and workflows change.
But I’d also make sure the organization has a way to evaluate whether the technology is creating value.
For a growth function, that may involve acquisition economics, incremental revenue, contribution margin, retention, customer value and the quality of decision-making. For finance, it may be accuracy, exception rates, reconciliation and the cost of processing transactions. For customer service, the relevant measures could include resolution quality, repeat contact rates, customer satisfaction and the total cost of service.
The point is that better technology doesn’t necessarily produce better business performance. We need to connect the investment to the outcomes that matter and maintain the ability to verify them independently.
I’d rather see a business using a relatively small number of properly governed agents that demonstrably improve its economics than one celebrating the deployment of hundreds without a clear understanding of their contribution.
We’ve spent years trying to move organizations away from measuring activity and toward measuring results. We shouldn’t abandon that discipline because the activity is now being performed by software.
What the Research Does and Doesn’t Establish
There’s one other aspect of this discussion that deserves some discipline, particularly given the increasingly dramatic coverage surrounding advanced AI.
Reported incidents and research demonstrations are important, but they don’t all measure the same thing. A system successfully circumventing a safeguard in a controlled test doesn’t establish how frequently comparable behavior occurs across ordinary business deployments.
Likewise, a security breach involving an AI application doesn’t necessarily mean the model independently decided to circumvent its instructions. The cause may involve compromised credentials, an unsafe integration, a poorly configured application or external manipulation.
Those distinctions matter because the corrective actions are different.
If the issue is excessive permissions, we need stronger access controls. If the problem originates in poor data, the business needs better data quality, definitions and verification. If an agent is producing commercially poor decisions, we may need to revisit the objectives and economics rather than change the security infrastructure.
And if testing reveals behavior that could circumvent established controls, the deployment may require additional restrictions, monitoring or changes to the system itself.
Treating every problem as a manifestation of the same broad AI risk makes it harder to identify what needs to change.
The adoption research also has limitations. McKinsey’s findings tell us about reported usage and scaling, not the probability of incidents. Deloitte’s findings tell us about reported governance maturity, not whether a particular organization has experienced harm or whether a specific control will prevent future failures.
We need more evidence about how these systems perform across different operating environments, how frequently significant incidents occur and which controls are most effective.
That doesn’t mean businesses should wait until every uncertainty has been resolved. It means we should avoid overstating what the evidence supports, particularly when deciding how much investment and organizational attention should be allocated to specific risks.
I think that’s another reason shared incident reporting and standardized qualification could be valuable. They would help create a more consistent evidence base, allowing businesses to learn from actual experience rather than relying exclusively on vendor assurances, theoretical scenarios or sensational headlines.
The Conversation I’d Rather Have
I think Nadella makes an important argument about separating model capabilities from the authority we give them. As these systems become more capable and begin operating across sensitive applications, businesses need independent controls, reliable records and the ability to intervene when necessary.
But I don’t think superintelligence should be the starting point for every discussion about AI adoption, particularly when so much of the market is still working through practical questions about data quality, integration, measurement, operating responsibilities and how to generate value.
The three variables I’ve outlined provide a way to bring the discussion into the context of the business: the maturity of its governance, the prevalence and autonomy of its AI deployments, and the complexity of its data environment.
That context matters because two businesses using the same model can have very different risk profiles. Even within the same company, an agent’s exposure can change materially depending on which applications it’s connected to, which information it can access and what decisions it’s permitted to execute.
I’d also like to see us build on the operating practices we already understand. An agent CV, independent evaluations, documented operating history and a credible shared record of incidents could help businesses assess systems before granting them access and responsibility.
We should distinguish an agent’s qualifications from its authorization, reassess those qualifications as operating conditions change and establish clear accountability across providers, deployment teams and the business owners granting authority.
None of this eliminates the uncertainties associated with advanced AI, nor does it mean the concerns raised by Nadella and others at the frontier should be ignored. Those concerns deserve serious attention, particularly as the technology develops and businesses become increasingly dependent on it.
But the majority of organizations have a considerable amount of work ahead of them just establishing the fundamentals. They need better information, clearer responsibilities, appropriate controls and a consistent understanding of whether the systems they’re introducing are actually improving the business.
I’d rather see practical standards that scale with those needs than assume every organization requires the same response to a set of increasingly sophisticated scenarios.
The more authority we delegate, the more important it becomes to understand the information, assumptions and controls behind the decisions being made. And the more these systems become part of everyday operations, the less useful it will be to assess them independently of the businesses in which they’re operating.
Perhaps the conversation about superintelligence would benefit from spending a little more time on the basic intelligence of how we actually run a business.
Sources and Further Reading
Satya Nadella: Models as Insider Risks in the Super Intelligence Era, October 10, 2026. The original article provides the starting point for this discussion, particularly the argument for separating intelligence from authority and establishing independent controls.
McKinsey: The State of AI, August 2026. The research informs the discussion of business adoption, scaling, productivity and enterprise-level financial impact.
Deloitte: Agentic AI Is Scaling Faster Than Guardrails, 2026. Its findings provide context on reported governance maturity among surveyed organizations.
NIST: Artificial Intelligence Risk Management Framework and the AI Agent Standards Initiative. These provide guidance on assessing, managing and supervising the use of AI across different operating environments.
MITRE ATLAS: Adversarial Threat Landscape for Artificial-Intelligence Systems. A reference for documented attack techniques, case studies and security concerns involving AI systems.
AI Incident Database: incidentdatabase.ai. A collection of reported AI-related harms and near harms that illustrates the potential for shared incident records.
Model Cards: Model Cards for Model Reporting, Mitchell et al., 2019. Research establishing an approach to documenting model capabilities, intended uses and limitations.
This piece is for members
Members read every essay and framework in full, the moment it is published.