Enterprise AI has a bit of a “she’ll be right, mate” problem. And as an Australian, I feel uniquely qualified to say this:

We have become extraordinarily good at proving what AI can do. But that isn’t the question that matters most anymore. The question is whether we can depend on it.
Because AI is moving out of the demo environment and into the machinery of the enterprise. Into workflows. Into decisions. Into customer interactions. Into security operations. And now, with agents, into systems that can access tools, consume resources and take action.
So the stakes have changed. The challenge is no longer simply building more capable AI. It is building organizations capable of operating that AI reliably, economically, securely — and knowing what to do when it doesn’t behave as expected.
And this is where I think the “she’ll be right” mentality starts to catch up with us. Right now, enterprise AI has no shortage of capability. More models. More copilots. More agents. More GPUs. More investment. What we still don’t have enough of are organizations that know how to run it properly. Because the demo is the easy bit. The demo doesn’t have to survive your data architecture. It doesn’t have to justify its token bill. It doesn’t have to navigate identity, permissions, legacy systems, governance, security — or the 47 people who somehow need to approve it. Production does.
That was the rabbit hole John Napoli and I went down during our recent You & AI conversation. And rather than spending 30 minutes talking about how amazing the latest model is, we focused on something I find much more interesting: what actually has to happen after the demo?
- How do you stop successful pilots from dying before production?
- How do you decide which AI use cases are actually worth pursuing?
- How do you fix the data bottlenecks underneath them?
- How do you choose between traditional ML, GenAI and agents instead of defaulting to the newest thing?
- How do you design for GPU constraints, model portability and changing economics?
- How do you govern a technology where every additional unit of “intelligence” can have a cost?
- And, perhaps most importantly: what changes when AI stops generating answers and starts making decisions and taking actions?
Because at that point, we are no longer simply deploying a model. We are giving software access, autonomy and the ability to act inside the enterprise. And that fundamentally changes what “production ready” means. It changes architecture. It changes economics. It changes governance. It changes security. And it changes the amount of evidence we need before we can say: yes, we trust this system to operate here.
That is why my biggest takeaway from the conversation wasn’t about a model, an agent or even a particular technology. It was this:
The next enterprise AI challenge isn’t proving AI is capable. We’ve done a fairly spectacular job of that already. The challenge is building the operating system around that capability.
And that is where my conversation with John really begins.
First, why John Napoli?
John was the opening episode of this You & AI series because he brings something I think the AI conversation needs more of: the perspective of someone who has actually had to scale AI solutions inside large, complex financial institutions.
John has spent about 30 years in financial-services technology, and his AI journey began around 2016 while he was COO of Technology at JPMorgan Chase. Today he is a technology executive at DTCC, the financial-market infrastructure company involved in clearing and settling trades. He described DTCC during our conversation as “probably the most important company nobody knows unless you’re really in financial services,” and said it processes about $4.7 quadrillion in trades annually. He has also written and continued updating a book on AI, an Amazon bestseller: Achieving Impact: Implementing Enterprise AI Via an AI Factory.
That background matters. Because our conversation wasn’t really about imagining what enterprises might someday do with AI. It was about the much less glamorous — and much more important — question of what happens when you actually try to deploy it.
A successful AI pilot proves surprisingly little
John put it this way.
A pilot really proves that AI can work. Production really proves that business can kind of work differently because of AI.
That deserves more attention. Because enterprise AI has become very good at producing moments of possibility. A model works. A demo impresses the executive team. A prototype answers the right questions. A pilot shows measurable improvement. Everyone gets excited. And then the organization tries to put it into production. That is when the actual work begins.
John described the model itself as often being the easy part. The harder work is integrating AI with enterprise data, security, risk controls, workflows, applications, people and operating processes — then keeping it reliable as the business changes.
This is why the phrase “AI pilot graveyard” resonated so much during our conversation. John’s argument was not that the graveyard is full of AI that couldn’t perform. It is full of technology that never became an operating capability.
That changes the diagnosis completely. If a technically successful pilot dies before production, the problem may not be the AI. The problem may be everything surrounding it.

The answer may be an AI factory, not more AI projects
John’s framework for this is what he calls the AI factory. When he explained why he wrote his book, he said enterprise AI shouldn’t be about “chasing kind of like the shiny object.” It requires a framework and structure. His AI factory concept brings together the intake process, platform, governance, data, cultural change and digital ethics into an overall enterprise approach.
This was one of the parts of our conversation I found particularly useful because it reframes scale. Most organizations don’t have an idea shortage. They have the opposite problem. Everybody has an AI idea. Every business unit can produce a list of use cases. Every executive has seen a demo. Every team wants a pilot.
The real enterprise capability is not the ability to start them. It is the ability to repeatedly select, build, govern, deploy and operate the right ones.
Don’t build a 100 AI pilots. Build a factory that can turn a 100 pilots into a 100 production systems.
That is a very different definition of AI maturity.

AI strategy is as much about saying “no” as saying “yes”
This became even clearer when I asked John how executives should prioritize the enormous lists of AI use cases now circulating inside companies. His advice was wonderfully unsexy: think big, start small, be ruthless about what you stop.
John said one of the biggest portfolio mistakes is treating every AI idea as equally important. His preferred starting point is a simple value-versus-effort analysis: move quickly on high-value, lower-effort opportunities; deliberately roadmap the harder, high-value bets; selectively automate low-value, low-effort work; and question why you are investing in low-value, high-effort ideas at all.
Then he went one step further. Before funding something, he said leaders should ask what business metric will move, by how much and by when. And he made a point I think a lot of AI strategies still miss:
An AI strategy isn’t just like a list of everything we could do. It’s a decision on what we’re gonna do and what specifically we’re not gonna do.
That might be one of the most practical pieces of advice from the entire episode. A long AI use-case inventory is not necessarily an AI strategy. A strategy makes choices.

Stop building AI because of FOMO
This flowed naturally into the next question: how do leaders separate AI hype from actual impact? John’s answer was direct:
Don’t build AI because you’re afraid of missing out.
Instead of starting with “What can we put in a chatbot?”, he said successful organizations are more pragmatic. Where are we losing time? Money? Customers? Competitive advantage? What does the process cost today? How long does it take? What exactly should improve?
John described a previous global-bank environment where approximately $500 million of AI investment was associated with around $1.7 billion in expected benefits. His point, he stressed, wasn’t the dramatic size of the numbers. It was that each major investment was connected to an economic outcome.
That distinction matters. AI value isn’t the existence of the AI. It is the change in the business produced by it.

Sometimes the AI bottleneck isn’t AI at all
Then we got into data. I have been talking about data as the lifeblood of AI for years, and enterprise adoption keeps proving the point.
John gave a striking example from one of his earlier AI environments. He said data scientists could wait up to nine months to obtain the data they needed because of restrictive data-use processes. The organization redesigned the governance model, pushing a number of detailed approvals closer to production while maintaining appropriate controls. According to John, the approval cycle went from roughly nine months to less than one month. And then he said something important:
It wasn’t a technology improvement. It’s really an operating model improvement.
This is an underappreciated part of enterprise AI transformation. Sometimes you don’t need a better model. You need a better process for approving data. Or clearer ownership. Or better identity architecture. Or a different operating model. Or faster governance.
John referenced the FAIR framework — Findable, Accessible, Interoperable and Reusable — and argued that the goal isn’t to eliminate data governance but to place it where it protects the organization without “locking it up or paralyzing it.”
That is what good governance should do: enable movement with control. Not control by preventing movement.

GenAI did not make traditional machine learning obsolete
I was glad we spent time here because this is something I believe the market is currently underestimating. The explosion of generative AI has created a strange tendency to treat every AI problem as a generative AI problem. John strongly disagreed.
He pointed out that traditional machine learning remains powerful for structured prediction, fraud detection, credit risk, forecasting, pricing, recommendation systems, demand planning and anomaly detection. GenAI changes the equation because it can reason across unstructured information, generate content and interact with tools and workflows. But those are different capabilities for different problems.
Don’t build an architecture around today’s favorite model.
Instead, John argued for architectures that allow organizations to change models as capabilities, economics and risks change. And then came what may be one of the best enterprise AI procurement rules I’ve heard recently:
Use the smallest, cheapest, and safest model that solves a problem.
Traditional machine learning where it is superior. GenAI where its flexibility creates value. Agents where the system needs to take action, not merely generate an answer.
The enterprise advantage, in other words, may not come from selecting one winning model. It may come from becoming exceptionally good at selecting the right intelligence for each job. John described the winning enterprise as the one able to continuously choose the best model for each task.

AI handles volume. Humans handle judgment.
When the conversation moved into high-stakes operational environments, John reduced a much larger debate about automation to one very useful sentence:
AI really handles volume, and the goal of humans is to really handle judgment.
That is a more interesting framing of the future of work than simply asking which jobs disappear. John was talking specifically about environments with large numbers of operational alerts. The opportunity isn’t just saving a few seconds of employee time. It is changing the signal-to-noise ratio so humans stop spending their days investigating alerts that lead nowhere and can concentrate on genuine risk.
This is where AI transformation becomes operational redesign. Machines absorb volume. Humans move toward judgment. And the workflow itself changes.
AI may look digital, but its foundations are physical
Then the conversation moved down another layer of the stack. GPUs. Memory. Networking. Data centers. Electricity. Cooling. Supply chains.
John’s point was a useful reminder that enterprises can talk endlessly about models while ignoring the physical infrastructure underneath them. He warned executives not to confuse access to a model with access to durable AI capability.
If an organization’s architecture depends entirely on one accelerator, one cloud provider or one model vendor, John argued, that creates a new concentration risk. That means model portability and cloud interoperability are not merely engineering preferences. They are becoming questions of enterprise resilience.
The objective, as John described it, isn’t to eliminate every vendor dependency. It is to make sure the business — rather than its infrastructure provider — retains options.

Tokenomics is turning AI into an enterprise resource-allocation problem
This may have been the most forward-looking part of our conversation. Traditional enterprise software gave leaders relatively understandable economics. Buy a number of licenses. Estimate the annual cost. Budget accordingly. AI changes the unit of consumption.
The AI cost model is really changing from buy the software to pay for the work the software does.
That is a profound change. An employee — and increasingly an agent — can continuously consume models, inference, tools and tokens. So the question cannot simply be: how much did we spend? It becomes: what did that intelligence produce?
John said knowing that an organization spent a million dollars on tokens isn’t enough. Leaders need to know which use cases consumed them, which business processes they supported, which outcomes were produced and whether the economics made sense. He compared the concept to “FinOps for the cloud but for intelligence.”
And that led us into AI unit economics. Cost per customer interaction. Cost per resolved case. Cost per decision. Hours saved. The interesting point here isn’t simply controlling token spend. It is connecting consumption to value.
John argued that the organizations that win won’t necessarily be those using the cheapest tokens, but those able to prove they are buying the right amount of intelligence for the value they create. And once agents become autonomous, economics itself starts to become a control surface. John gave an example of an agent being allowed to spend $5 on inference for a task but requiring escalation before spending $5,500.
Tokenomics ultimately turns AI from an IT procurement issue into, like, an enterprise resource allocation issue.
CFO. CIO. CISO. AI leadership. Same table.

Governance has to move with the AI
About halfway through our conversation, John turned the tables and asked me about AI governance and the growing risk around agentic systems. This is where I think enterprise thinking has to change significantly.
I told John that one of the biggest mistakes I see is organizations treating AI governance like paperwork — something you assess, approve, file away and move past. But AI doesn’t stay in the state in which it was approved. AI is in motion. The model changes. The data changes. The context changes. The tools change. The permissions change. And increasingly, the system itself is making decisions and taking actions.
So testing, evaluation, verification and validation have to move with it. That is the shift I described as: from governance as a gate to governance as a control plane.
And agentic AI makes that shift much more urgent. The question is no longer simply: was this model approved? Now we need to understand the whole system in motion. What can it access? What tools can it call? What permissions does it have? How independently can it decide? What can it actually do? And if something starts to go wrong, can we see it, explain it and stop it?
At Optica Labs, we created the 4As as one way to frame that risk:
Access
What can the system reach?
Autonomy
What can it decide without a human?
Action
What can it actually do in the world?
Accountability
Can we trace it, explain it, intervene and ultimately own the outcome?

But the part I think is especially important is that the 4As don’t stay static. They change as the mode of agentic AI changes.
I also walked John through the progression we use to help leaders understand that journey. At the lowest end, you may have copilots and personal-productivity systems. Then you move into task agents, workflow agents, multi-agent systems and ultimately semi-autonomous and autonomous operations. I think of that as a progression through the six modes of agentic AI, with M0 as the baseline before meaningful agentic autonomy begins:
- M0Copilot / Assistive AIAI helps a human, but the human remains firmly in control.
- M1Task AgentAI can complete a bounded task with limited access and tightly defined actions.
- M2Connected Task AgentAI begins using tools, systems and enterprise data to complete more complex tasks.
- M3Workflow AgentAI can coordinate multiple steps across a business process, with greater autonomy and more opportunities for something to go wrong.
- M4Multi-Agent SystemMultiple agents interact, delegate, coordinate and potentially influence one another across systems.
- M5–M6Semi-Autonomous to Autonomous OperationsAI can increasingly plan, decide, act and operate with significantly less human intervention.

And this is the part executives need to understand: every time you move up a mode, the risk profile changes. The agent generally needs more access to data, tools, identities and systems. It receives more autonomy to make decisions without asking a human every time. It gains more ability to act — whether that means sending an email, changing a record, calling an API, moving money, writing code or triggering another system. And because access, autonomy and action increase, accountability has to increase with them.
That means the governance model for a copilot should not look the same as the governance model for a multi-agent system operating across critical workflows.
M0 does not require M6 controls. But M6 absolutely cannot survive on M0 governance.
That is where I think a lot of organizations are going to get caught. They are upgrading the capability of the AI without upgrading the controls around it. More autonomy without stronger oversight. More access without tighter identity and permissions. More action without better monitoring. More agents without understanding how risk compounds when those agents interact.
The higher the mode, the more governance has to become continuous, technical and operational — not just a policy document sitting somewhere in SharePoint. You need visibility into identity, permissions, tool use, data access, decisions, actions, costs, system-to-system interactions and human intervention points. And critically: you need the ability to stop the system.
Because the real governance question in an agentic enterprise is no longer simply “Was this AI approved?” It is: “What is this AI allowed to do right now — and do we still trust it to keep doing it?” That, to me, is what governance in motion really means.
Agentic AI changes the cybersecurity model
We ended the episode in perhaps the most important place of all: security. John made the point that AI is no longer simply another application that needs to be secured. AI is entering the environment on multiple sides. It can be used by defenders. It can be used by attackers. And AI systems themselves can now become part of the attack surface.
Agents can read emails, browse the internet, write code, call APIs, access databases and take actions on behalf of employees. In John’s words:
We moved from software that kind of waits around for instructions to software that can make decisions and act.
That requires a different security architecture. John offered three particularly practical principles:
- 01
Treat every AI agent as a privileged digital employee
Give it identity, least-privilege access, clear entitlements and continuous monitoring. John even suggested the concept of KYA — Know Your Agent.
- 02
Assume the model can be manipulated
Security controls should not depend solely on the model enforcing its own rules; controls need to exist around identity, APIs, networks, data and tools.
- 03
Red-team the entire system, not just the model
Test what happens when an agent receives malicious email, encounters poisoned content, calls a compromised tool or attempts to move beyond its boundaries.
That last point is something I care deeply about, get to build products around and work on every day at Optica Labs. Because testing a model’s outputs is no longer enough when the system surrounding the model has identities, memory, tools, permissions, APIs, workflows and the ability to act. We have to test the ecosystem, not just the model.

The question I would take into every boardroom
Near the end of the interview, John gave perhaps the cleanest executive-level security question of the entire conversation. He said leaders shouldn’t only ask what an AI can do. They should ask:
What could this AI do if it was compromised, and what do we need to do to stop it?
I think that question extends well beyond cybersecurity.
- What can the AI access?
- What decisions can it make?
- What actions can it take?
- What resources can it consume?
- Who owns the outcome?
- Can we see what it is doing?
- Can we intervene?
- Can we prove what happened afterwards?
Those questions are going to become increasingly important as enterprises move from AI that assists to AI that acts.
And that was ultimately where John and I landed. Enterprise AI is no longer just a model problem. It is a data problem. An operating-model problem. An economics problem. An architecture problem. A governance problem. A security problem. And increasingly, an assurance problem.

At the end of the conversation, I described where I think this is going:
Test, evaluate, verify, and validate systems — and do that in motion.
Because as AI becomes more dynamic, our controls have to become dynamic too. The enterprises that win this next phase of AI may not simply be the organizations that deploy the most AI. They may be the ones that become exceptionally good at determining where AI creates value, which intelligence to use, what authority to give it, how much it should cost, how to continuously test it — and when to stop it.
That is a much harder problem than building a chatbot. It is also a much more interesting one. And it gets us closer to the question at the center of my new series on You & AI: how do we build AI we can trust?
Until the next rabbit hole finds us.
Tiarne “T” Hawkins

