Moving Enterprise AI From Pilot to Production: A Four-Stage Framework

Updated: 6 days ago
Title: Moving Enterprise AI From Pilot to Production: A Four-Stage Framework
Date: 12 June 2026
Type: Paper
Author: SAASiQ (contact@saasiq.ai)
Word count: 2690 words
Reading time: 10 min
Published: 12-06-2026
Oracle told investors on 10 June that its customers have moved past the experiment stage with AI, and that 33 of them had pre-purchased token bundles for extra agentic capacity in the fourth quarter. Survey evidence puts most organisations some way behind that: in Deloitte's State of AI in the Enterprise report, only 25 per cent had moved 40 per cent or more of their AI pilots into production. This paper sets out four stages that take a working pilot into operation, using the evaluation and control tools that Oracle, Microsoft and others released in the first half of 2026.
What the surveys measure
Deloitte's report, published on 21 January, drew on 3,235 business and IT leaders in 24 countries, surveyed in August and September 2025. Alongside the 25 per cent that had moved at least 40 per cent of their pilots into production, another 54 per cent expected to reach that level within three to six months. Nearly three quarters planned to deploy agentic AI within two years, and 21 per cent said they had a mature model for governing agents.
McKinsey's State of AI survey, published in November 2025 from 1,993 participants in 105 countries, found 88 per cent of organisations using AI in at least one business function, but only about a third had begun to scale it across the enterprise. Sixty-two per cent were at least experimenting with AI agents and 23 per cent were scaling an agentic system somewhere in the business. Thirty-nine per cent reported any impact on EBIT from AI.
PwC's 29th Global CEO Survey, published at Davos on 19 January, asked 4,454 chief executives in 95 countries and territories what AI had done for them over the previous 12 months. Fifty-six per cent had seen neither a revenue nor a cost benefit, and 12 per cent had seen both. KPMG's first Global AI Pulse, published on 31 March from 2,110 senior leaders in 20 markets surveyed between 17 February and 17 March, put 11 per cent of organisations in its most advanced group, deploying and scaling agents in ways that produce results across the whole business.
The figure most often quoted, that 88 per cent of pilots never reach production, comes from IDC research with Lenovo reported by CIO.com in March 2025: for every 33 AI proofs of concept a company launched, four reached production. The coverage did not give the sample size. MIT's Project NANDA report of August 2025, which reviewed more than 300 publicly disclosed AI initiatives and added interviews and a survey of senior leaders, found that about 5 per cent of generative AI pilots were achieving rapid revenue acceleration.
These studies count different things (pilots converted, pilots with a financial return, organisations operating at scale), so their numbers cannot be set side by side. All of them find that most organisations have more pilots than production systems.
Why pilots stall
Gartner predicted in June 2025 that over 40 per cent of agentic AI projects will be cancelled by the end of 2027, because of escalating costs, unclear business value or inadequate risk controls. Anushree Verma, a senior director analyst there, described most current agentic projects as 'early stage experiments or proof of concepts that are mostly driven by hype'. In a Gartner poll of 3,412 webinar attendees in January 2025, 19 per cent said their organisation had made significant investments in agentic AI and 42 per cent conservative ones.
IDC's authors put the low conversion rate down to low organisational readiness in data, processes and IT infrastructure. Ashish Nadkarni, a group vice president at IDC, told CIO.com that many generative AI initiatives start at board level and that the proofs of concept are often underfunded or not funded at all.
Deloitte's respondents named insufficient worker skills as the biggest barrier to fitting AI into existing workflows, and 84 per cent had not redesigned jobs or the nature of work around AI. McKinsey's high performers, the roughly 6 per cent reporting an EBIT impact of 5 per cent or more from AI, were nearly three times as likely as others to have fundamentally redesigned their workflows. They were also more likely to have defined processes for deciding when a model's output needs human validation, and to have senior leaders who own the work.
MIT's researchers found that AI tools bought from specialised vendors succeeded about 67 per cent of the time, and internal builds about a third as often. Most of the reasons in these studies are organisational. The framework below therefore deals with definitions, testing, controls and ownership, and says little about which model to use.
Before starting
The framework assumes three things: at least one pilot that works in a controlled setting, a system of record it has to connect to (an ERP, HR or case management system), and a team that could run it once it is live. An organisation with no working pilot yet needs to choose and scope a first use case before any of this applies.
It also needs a list of the AI already running, because a good deal of it arrives in software updates. Oracle said on its March earnings call that it had delivered well over 1,000 AI agents inside its applications, at no additional cost. On 9 June Google switched Gemini 3.5 Flash on by default in Gemini Enterprise in all regions and removed the option to turn it off. Features like these reach production without passing through a pilot at all, and they need the same definition, testing and ownership as anything built in-house.
Stage one: define production before building
Each system needs a written production definition, agreed with the people who will approve it, before build work starts. It should state the business measure the system has to move, the accuracy it needs to reach, the failures that are unacceptable and the cost per transaction above which it stops being worth running. A pilot shows that a model can do the task. The definition says what the organisation needs to see before it relies on the system.
The business measure should be something an operations manager already tracks, such as the time taken to clear an invoice exception, with a limit on payment errors attached. A model accuracy figure on its own does not tell an approver whether the process will improve.
The cost ceiling needs working out at full volume, because AI pricing is moving towards consumption. Oracle said on 10 June that much of the AI in its core applications stays included at no extra charge, while token bundles buy additional agentic capacity and access to advanced reasoning models. ERP Today reported the next day that Oracle is also extending outcome-based pricing, with interview agents priced by the number of candidates screened. AI Agent Studio has shown token consumption for premium models since its October 2025 release. An agent that calls a model several times per transaction can look cheap in a pilot of a few hundred cases and expensive at a few hundred thousand.
Unacceptable failures should be named specifically. A misrouted service ticket can be corrected later. A payment approved that should have been flagged, or a candidate wrongly screened out, cannot, and some of these uses carry legal duties. Under the provisional Digital Omnibus agreement reached on 7 May, the EU AI Act's obligations for high-risk systems in Annex III, which covers employment, education, biometrics and critical infrastructure among others, now apply from 2 December 2027, subject to formal adoption. UK organisations are outside the Act's direct reach unless they operate in the EU.
Stage two: build the evaluation alongside the system
The evaluation is a deliverable in its own right. It needs a test set of real cases with agreed answers, including the awkward ones such as poor scans and duplicate supplier records, scoring that maps to the production definition, and a routine that runs it after every change. With it, the conversation with risk and audit teams is about results across a known set of cases rather than a demonstration.
Oracle has put this into AI Agent Studio for Fusion Applications. Its documentation tells customers to evaluate agents before deploying them, scoring each answer from 0 to 1 against a reference answer in an evaluation set and reporting median correctness, median and 99th percentile response times, and token counts. The 26A update added three measures for agents that retrieve documents (groundedness, answer relevance and context relevance), scored by a model acting as judge. The same documentation says to rerun evaluations after any change to an agent or after a model update, and a Monitoring view tracks response times, token counts and errors once the agent is in production.
Microsoft announced similar tools at Build. On 3 June it said tracing and evaluations in Microsoft Foundry were generally available, with previews extending them to agents built on LangChain, LangGraph, the OpenAI SDK and other frameworks through OpenTelemetry, along with multi-turn evaluation, a rubric evaluator and simulated users for testing edge cases. On 2 June Microsoft's Sarah Bird released ASSERT as open source. It turns an organisation's written policies into targeted test cases, and it comes with a proposed open standard, the Agent Control Specification, that defines controls at five points in an agent's work: input, the model call, state, tool execution and output.
The test sets and the harness should belong to the organisation, whichever tool runs them. On 3 June OpenAI deprecated its hosted Evals platform, which becomes read-only on 31 October and shuts on 30 November, and its migration guide points customers to Promptfoo.
Models also change underneath a live system. On 5 June Anthropic told developers that Claude Opus 4.1 would be retired from the Claude API on 5 August, with Opus 4.8 as the recommended replacement, and requests to a retired model fail. Each retirement notice or default change is a reason to rerun the test set before the switch happens.
Stage three: connect to the system of record with an off switch
A pilot that runs against a sandbox or a data extract has not yet met the hard part, which is working against live data inside the system's security and approval rules. The integration should be written down: which systems the AI reads from and writes to, through which interfaces, with which permissions. Every action it takes against a system of record should be authorised and logged, and reversible where possible. Where an action cannot be undone, such as an external payment or a message sent to a customer, it should go through a person.
Deloitte gives the same advice for agents: start with lower-risk uses, set clear limits on what an agent may decide without a person's approval, and keep audit trails that capture the full chain of its actions.
In Fusion, much of this comes from the role design. Oracle's 26B readiness notes say that to use agents on a Fusion page, a user's job role has to contain the Fai Genai Agent Runtime Duty, with permission groups enabled on that role in the Security Console, and the role must already give access to the pages where the agents run. An agent works within the access its user already has, so the roles set up in an implementation are also the security design for its agents.
The notes also show the off switch. The Cost Accounting Close Workspace stays off until an administrator sets the profile option ORA_CST_PERIOD_CLOSE_AGENTIC_APP_ENABLED to Yes at site level, and setting it back withdraws the application. The note tells users to check AI-recommended actions against their own priorities before applying them. For agents working directly against the database, Deep Data Security in Oracle AI Database 26ai, available since 1 May, enforces row, column and cell rules for the end user inside the database and blocks an update that touches a value outside that user's rights. It is a 26ai feature, so databases still on 19c do not have it.
When the AI is withdrawn, the process also has to carry on, manually or on another model that has passed the same tests. At 5:21pm Eastern time on 12 June, the date of this paper, Anthropic received a US Commerce Department directive suspending access to Claude Fable 5 and Mythos 5 by foreign nationals, and because it could not check nationality in real time it disabled both models for every customer. Fable 5 had been released three days earlier. Anthropic's other models, including Opus 4.8, were not affected.
Stage four: name an owner and run it
Every system in production needs a named owner, a person or team accountable for its behaviour, performance and running cost for as long as it is live. The owner reruns the evaluation on a schedule, watches cost against the ceiling set in stage one, responds when results drift, and decides when the system is adjusted or withdrawn. That needs authority over the budget and the off switch as well as the accountability.
IBM's 2026 CEO Study, published on 4 May from 2,000 chief executives in 33 geographies surveyed with Oxford Economics between February and April, found 76 per cent had a chief AI officer, up from 26 per cent a year earlier. The same study found that only 25 per cent of the workforce used AI regularly, and 83 per cent of the CEOs said AI success depends more on people's adoption than on the technology. A chief AI officer does not replace an owner for each system, sitting in the business function the system serves.
Deloitte found that companies whose senior leadership actively shapes AI governance get significantly more business value than those that leave the work to technical teams. McKinsey's finding on senior ownership among its high performers points the same way.
The vendors now supply reporting for the owner. Oracle added an Agent ROI dashboard to AI Agent Studio on 24 March, reporting time saved, cost savings and productivity gains for each agent by workflow, team and function. Microsoft's ROI view for agents in Foundry, in private preview since 3 June, shows task completion rates, time saved and cost efficiency together, and Agent 365, generally available since 1 May at $15 per user per month, keeps a registry of the agents running in a tenant. These report the vendor's own measures, so the owner should check them against the business measure in the production definition.
Each change to a live system, whether a new prompt, a new model or a new permission, should be recorded with the evaluation result that justified it, so it can be reversed. A withdrawal handled that way is a normal operating event, and the record also answers an auditor who asks how a particular output was produced.
What it costs and where it breaks
The framework has running costs. Test sets need real cases and agreed answers, which means time from the people who do the work now, and Deloitte's respondents already name skills as their biggest constraint. The integration and fallback design is software work that someone has to build, secure and maintain.
Evaluation by a model acting as judge, as in Oracle's retrieval measures and Microsoft's rubric evaluator, is itself a model and should be checked against a sample of cases marked by people. Vendor claims need the same checking: Gartner estimated in June 2025 that only about 130 of the thousands of vendors claiming agentic AI offer real agentic features, a practice it calls 'agent washing'.
This framework covers the move from pilot to production. It does not cover choosing the first use case, retrieval design, or the contract terms on data handling and availability, all of which change the calculation for particular systems.
SAASiQ's view is that for Fusion customers the cheapest place to start is the agents already in the licence, since the role security, the profile-option off switch and the evaluation screens are already there.
Claude Opus 4.1 is due to be retired from the Claude API on 5 August 2026, OpenAI's hosted Evals platform becomes read-only on 31 October 2026, and the EU's Annex III high-risk obligations are due from 2 December 2027, once the Omnibus is formally adopted.
SAASiQ - Intelligent Solutions for SaaS ©


