Scaling AI Agents With Controls: A Five-Step Framework

Updated: 6 days ago
Title: Scaling AI Agents With Controls: A Five-Step Framework
Date: 10 April 2026
Type: Paper
Author: SAASiQ (contact@saasiq.ai)
Word count: 2961 words
Reading time: 11 min
Published: 10-04-2026
Tags: #AI #Enterprise #Governance #Agentic #Oracle
Oracle announced another set of Fusion Agentic Applications at its AI World Tour in New York on 9 April, 12 for finance and supply chain, eight for HR and five for customer experience, each built to carry routine work forward inside Fusion's security framework and pass exceptions to a person. Two days earlier WRITER published a survey in which 36 per cent of executives said they had no formal plan for supervising AI agents. This paper sets out how an organisation can take AI from pilots into production with the controls in place before the volume arrives.
How far adoption has got
McKinsey's State of AI survey, published in November 2025 from 1,993 respondents in 105 countries, found 88 per cent of organisations using AI regularly in at least one business function, up from 78 per cent a year earlier. About a third had begun to scale AI across the enterprise. Thirty-nine per cent could attribute any effect on enterprise-wide EBIT to AI, and most of those put it below 5 per cent. On agents, 62 per cent were at least experimenting and 23 per cent said they were scaling an agentic system somewhere in the business.
Deloitte's State of AI in the Enterprise report, published on 21 January from a survey of 3,235 business and IT leaders, found that the share of workers with access to sanctioned AI tools had gone from under 40 per cent to about 60 per cent in a year, and that fewer than 60 per cent of those with access used the tools in their daily work. Only 25 per cent of organisations had moved 40 per cent or more of their AI pilots into production. Deloitte's explanation is that a pilot can run with a small team on cleaned data in an isolated environment, while production needs integration with existing systems, security reviews, compliance checks and monitoring, and use cases estimated at three months can stretch to 18.
PwC's Global CEO Survey, published at Davos on 19 January, found that 56 per cent of 4,454 chief executives had seen neither a revenue nor a cost benefit from AI over the previous year, and 12 per cent had seen both. Those whose companies had what PwC calls strong foundations, such as a responsible AI framework and a technology estate that lets AI be used across the whole business, were three times more likely to report meaningful financial returns.
What changes when agents reach production
An assistant that drafts text leaves a person to act on it. An agent takes the action itself, whether that means posting a journal, sending a message or changing a record, and Deloitte treats governing agents as a separate problem for that reason. Nearly three quarters of its respondents planned to deploy agentic AI within two years, and 21 per cent said they had a mature model for governing agents. Data privacy and security topped the list of AI risks at 73 per cent, followed by legal, intellectual property and regulatory compliance at 50 per cent and governance capabilities and oversight at 46 per cent. In Deloitte's interviews, one AI leader found there was no clear inventory of the AI tools and models active in the organisation, because development had happened without central tracking.
WRITER's survey of 2,400 people in the US, UK and four other European markets, published on 7 April, found that 67 per cent of executives believed their company had suffered a data leak or security breach because an employee used an unapproved AI tool, and 35 per cent said they could not immediately 'pull the plug' on a rogue agent. WRITER sells an enterprise AI platform, and its figures record belief. IBM's Cost of a Data Breach Report, published in July 2025, measured something close to it: of the organisations that had an AI-related security incident, 97 per cent lacked proper AI access controls, and 63 per cent of the 600 organisations studied had no AI governance policy.
In July 2025 Jason Lemkin, the founder of SaaStr, told Replit's coding agent that a code and action freeze was in place. According to Fortune, the agent ran commands anyway and deleted a production database holding data on more than 1,200 executives and 1,190 companies. The fixes Replit announced were automatic separation of development and production databases, better rollback and a planning-only mode, and none of them relies on the model following an instruction. Gartner said in June 2025 that it expected more than 40 per cent of agentic AI projects to be cancelled by the end of 2027, and it listed inadequate risk controls alongside cost and unclear business value as the reasons.
What Oracle released on 9 April
Oracle's New York announcements named 12 applications across Fusion Cloud ERP and Supply Chain and Manufacturing, among them the Claims Settlement Workspace, the Collectors Workspace and the Maintenance Operations Workspace. The eight for HR include a Workforce Operations Command Center and a Hiring Workspace for Store Managers, and the five for customer experience include a Sales Command Center. They follow the 22 Fusion Agentic Applications Oracle announced in London on 24 March.
The wording on control is the same in all three releases. Working inside the existing Fusion security framework, the applications move routine work forward within set guardrails and surface the exceptions, trade-offs and decisions where a person's judgement changes the outcome. Oracle says they draw on the data, workflows, policies, approval hierarchies and permissions already held in Fusion, and that AI Agent Studio supplies built-in observability, ROI measurement and safety controls.
For a Fusion customer, most of the controls therefore come from configuration it already owns. An agent acting for a user reaches what that user's roles allow, so the role design done at implementation now decides what agents can see and do. On 24 March Oracle added an Agent ROI dashboard to AI Agent Studio, and its 26B readiness notes, updated on 27 March, show how it works: an administrator enters the time and cost saved by each run of an agent team, and the dashboard multiplies those figures by the number of runs. Oracle says AI Agent Studio comes at no additional cost, and SiliconANGLE reported on 24 March that basic agents on the built-in models are included while premium language models are charged by usage.
The frameworks already published
NIST's AI Risk Management Framework, version 1.0 from January 2023, organises the work into four functions: Govern, Map, Measure and Manage. ISO/IEC 42001, published in December 2023, describes an AI management system that an organisation can be certified against, with the policies, roles, risk assessments and records an auditor expects to find.
The first framework written for agents came from Singapore. Its Infocomm Media Development Authority launched the Model AI Governance Framework for Agentic AI at the World Economic Forum on 22 January. It is voluntary and has four parts. Risks are assessed and bounded upfront, by restricting an agent's tools, permissions and operating environment. Humans are made meaningfully accountable, with oversight that can override, intercept or review an agent's actions. Technical controls run through the lifecycle, from least-privilege access and testing before deployment to progressive rollout and real-time monitoring. And end users are given the transparency and training to use agents responsibly.
On security, the OWASP GenAI Security Project published its Top 10 for Agentic Applications on 9 December 2025. Its first three entries are an attacker changing an agent's goals through content it reads, an agent using a legitimate tool in an unsafe way, and an agent acting with more authority than it should have. In the UK, the Information Commissioner's Office published an early report on agentic AI on 8 January. It is not guidance, but it flags purposes set too broadly for open-ended tasks, agents processing more personal data than an instruction needs and wider use of automated decisions with significant effects, and it says organisations remain responsible for their use of agentic AI systems.
The framework in outline
The framework below turns these into five steps: inventory and tier every use, approve at a small number of gates, bound each agent's identity and access, enforce policy at runtime, and measure and review on a schedule. The first two are done before anything scales. The last three keep control once it has.
Three things need to be in place first. There should be a named owner for AI with the authority to approve or stop a use, a data classification that says which data may go to which model and on what terms, and money for the controls inside the rollout budget, since the security reviews, compliance checks and monitoring Deloitte lists are the work that separates a pilot from production.
Step one: inventory and tier every use
The inventory covers AI in production, AI in pilots and AI in use without approval, including features that arrived switched on in a quarterly update and tools reached through personal accounts. Netskope's Cloud and Threat Report in January found that 47 per cent of people using generative AI at work did so through personal accounts, down from 78 per cent a year earlier. For each use, record the business owner, the model and where it runs, the data it reads, the systems it can write to and who approved it.
Then tier each use by what it can do. A workable scale has four levels: it reads and summarises; it drafts something for a person to send or post; it writes to a system of record within set limits; or it takes actions outside the organisation, or actions that cannot be reversed, such as payments and messages to customers. Controls rise with the tier.
Deloitte found that the companies doing best started with lower-risk uses, built governance capability and then scaled deliberately, and Singapore's framework asks for progressive rollout. The tiers give that sequence a shape: move uses up a tier only when the controls for the next one are working.
Step two: approve at a small number of gates
Each tier needs a few gates, each with a named decision maker and a stated body of evidence. At use-case approval the evidence is the business owner, the tier and the data classification. Before production it is test results on the organisation's own cases, and Singapore's framework lists what to test: task execution, policy compliance and accuracy in using tools, across varied data. At go-live it is a rollout plan that starts with a limited group of users and a date for review.
Deloitte recommends cross-functional governance that brings IT, legal, compliance and business unit leaders together to set policies, monitor performance and manage escalations, and says it should sit within existing risk and oversight structures, with no parallel function beside them. Each gate should have a published turnaround time, so that teams can plan around it.
Step three: bound identity and access
Every agent needs an identity that ties its actions to the agent and to the person or process that gave it authority, and its access should be the minimum its tier requires. Microsoft's Agent 365, announced on 9 March and due for general availability on 1 May at $15 per user per month, keeps a registry of agents and gives each one an identity in Microsoft Entra, so conditional access and identity governance apply to it as they do to a user. NIST's National Cybersecurity Center of Excellence published a concept paper on agent identity and authorisation on 5 February, with comments open until 2 April, which asks whether an agent's identity should persist or last only for a single task, and how a compromised agent's credentials are revoked.
In Oracle Fusion the agent works through the user's roles. Oracle's 26A readiness notes show that a user needs a job role with the Fai Genai Agent Runtime Duty to use an agent, and that the same role must already hold the privileges for whatever the agent does. Any separation of duties conflict inside a role therefore applies to the agent as well. Oracle's Risk and Security Snapshot report shows conflicts within a single role and those that arise from the combination of roles one person holds. SAASiQ's view is that the snapshot should be run on the roles that will carry the runtime duty before any agent is switched on.
Step four: enforce policy at runtime
Approval at a gate says what an agent may do. Runtime enforcement checks each action as it happens, outside the model and before the tool runs. The UK's National Cyber Security Centre set out the principle on 8 December 2025. Design on the assumption that the model will sometimes be fooled, drop its privileges to those of whoever supplied the content it is processing, limit its actions with safeguards that do not depend on the model, and log inputs, outputs, tool use and API calls. Dave Chismon, the NCSC's technical director for platforms research, described language models as 'inherently confusable'.
AWS made Policy in Amazon Bedrock AgentCore generally available on 3 March in 13 regions, including London, Frankfurt and Ireland. A policy engine attached to AgentCore Gateway intercepts traffic between agents and tools and evaluates each request before allowing or refusing it. Policies written in plain English are converted to Cedar, AWS's open-source policy language, and the agent's code does not change. Inside Fusion, the 26A update added a human approval node to workflow agents in AI Agent Studio, which pauses a workflow until a person approves, rejects or asks for a change. It is a node the builder adds, so an approval exists only where someone has designed one in.
Every tier above read-only also needs a stop, a way to halt an agent and revoke its credentials quickly, which is the ability 35 per cent of WRITER's executives said they did not have.
Step five: measure and review
The records runtime enforcement produces are also the evidence that the controls work. Each should show what the agent was asked to do, what it did, which policy allowed it, who approved it if anyone did, and the outcome, with one identifier per agent session so that a request can be followed across the gateway, the identity service and the application. AI Agent Studio has had a monitoring dashboard since October 2025 showing sessions, latency, error rates and token usage, and 26A added scoring of answers drawn from documents for groundedness and relevance.
Value belongs in the same review. PwC's link between strong foundations and returns assumes returns are measured, and Oracle's dashboard totals are multiples of the per-run figures an administrator enters, so they are estimates the customer sets.
The review runs on a fixed cycle: incidents and blocked actions monthly, policies and tiers quarterly, and the whole inventory once a year, with an extra review whenever a vendor changes a model or switches on a new feature. NIST's Manage function and ISO/IEC 42001 both expect evidence of this kind, showing that controls operate and are revised when they fail.
The regulation as it stood on 10 April
The EU AI Act entered into force on 1 August 2024, and its bans on prohibited practices and the AI literacy duty in Article 4 have applied since 2 February 2025. Under the text as it stands, the obligations for high-risk systems apply from 2 August 2026. The Commission's Digital Omnibus on AI proposed deferring them. The Council agreed its position on 13 March and the European Parliament adopted its mandate on 26 March by 569 votes to 45, and both back 2 December 2027 for the stand-alone high-risk systems in Annex III and 2 August 2028 for AI in products covered by Annex I. Trilogue talks are under way, and A&O Shearman reported on 9 April that a political agreement was expected at the next trilogue on 28 April. Until an amendment is adopted, 2 August 2026 remains the legal date.
When the high-risk obligations do apply, they match steps four and five closely. Article 12 requires automatic recording of events, Article 14 requires that the people overseeing a system can override it or stop it in a safe state, and Article 26 requires deployers to assign oversight to competent people and keep logs for at least six months. Annex III covers recruitment, promotion, task allocation and credit scoring, so some HR and finance agents may fall within it. UK organisations are outside the Act's direct reach unless they operate in the EU. The UK's voluntary AI Cyber Security Code of Practice, published on 31 January 2025, sets out 13 principles running from secure design to end of life.
What it costs and where it breaks
The framework has running costs. The inventory and the test sets need time from the people who do the work, the policy engine and the log store are systems someone has to build and run, and per-user charges such as Agent 365's come on top of model usage, which grows with the number of agent runs.
It also has limits. A policy engine checks only the actions put to it, so an agent hijacked through a malicious document can still misuse the tools it is allowed, within the limits set, and the NCSC's view is that prompt injection may never be fully mitigated. Several of the survey figures above come from companies that sell AI or security products, and all of them record what respondents say about their own organisations. The framework does not cover model selection, fairness testing or contract terms, each of which needs work of its own.
The next trilogue on the Digital Omnibus is scheduled for 28 April, Agent 365 becomes generally available on 1 May, and Oracle's 26B update reaches the first customers' test environments on 1 May and production on 15 May.
SAASiQ - Intelligent Solutions for SaaS ©


