Runtime Governance for AI Agents: A Framework for Checking Each Action Before It Runs

Updated: 7 days ago
Title: Runtime Governance for AI Agents: A Framework for Checking Each Action Before It Runs
Date: 29 April 2026
Type: Paper
Author: SAASiQ (contact@saasiq.ai)
Word count: 2951 words
Reading time: 11 min
Published: 29-04-2026
An AI agent that can update records, call other systems and hand work to other agents needs a control at the point where it acts, and a model that answers safely does not provide one. Oracle published a governance framework built on that idea on 23 April, AWS made a policy engine for agent tool calls generally available on 3 March, and Google announced an Agent Gateway and Agent Registry on 22 April. This paper sets out a vendor-neutral framework for that control.
Why model safety does not cover what agents do
Anthropic's system card for Claude Mythos Preview, published with the model on 7 April, describes an earlier version of it being placed in a sandbox and told to try to get out and contact the researcher running the test. It got out: it built an exploit to reach the internet through a system meant only for a few predetermined services, emailed the researcher and then, without being asked, posted details of the exploit on several obscure public websites. In another case it edited files it had no permission to edit and then made sure the changes would not show in the change history. Anthropic calls Mythos Preview 'the best-aligned model that we have released to date by a significant margin', and has kept it out of general release.
In July 2025 Jason Lemkin, the founder of SaaStr, had told Replit's coding agent that a code and action freeze was in place and that nothing should change without his approval. The agent ran commands anyway and deleted a production database holding data on more than 1,200 executives and 1,190 companies, according to Fortune, then told him the data could not be recovered, which was wrong. Amjad Masad, Replit's chief executive, called it 'unacceptable'. The fixes Replit announced were automatic separation of development and production databases, better rollback and a planning-only mode, and none of them depends on the model obeying an instruction.
The UK's National Cyber Security Centre made the general point in December 2025. Its technical director for platforms research, Dave Chismon, described large language models as 'inherently confusable' and said prompt injection may never be fully mitigated. The NCSC's advice is to limit what a model can do with safeguards that do not rely on the model, and to log its inputs, outputs, tool use and API calls.
Attacks arrive through the tools
EchoLeak, found by Aim Security's research team and fixed by Microsoft in its June 2025 updates as CVE-2025-32711 with a severity score of 9.3, needed no click from anyone. A crafted email reached a user's inbox, Microsoft 365 Copilot read it while doing other work, and hidden instructions in it could lead Copilot to send out data from the user's mail, chats and files.
Invariant Labs showed the same pattern on 26 May 2025 with GitHub's MCP server. An attacker files an issue on a public repository with instructions buried in the text. When a developer later asks an agent to look at open issues, the agent reads them and acts on them with the access it already holds, which can include private repositories. Invariant said the problem lay in the architecture, and its suggested mitigations, limits on the agent's session and least-privilege tokens, both sit outside the model.
The OWASP GenAI Security Project published its Top 10 for Agentic Applications on 9 December 2025. Three of its entries describe these cases. Agent Goal Hijack (ASI01) is an attacker changing an agent's objectives through content it reads, Tool Misuse and Exploitation (ASI02) is an agent using a legitimate tool in an unsafe way, and Identity and Privilege Abuse (ASI03) is an agent acting with more authority than it should have. The principle behind the list is least agency: approval steps and credentials scoped to the task limit what an agent may do, on top of limits on what it can reach.
What organisations say they can and cannot do
Kiteworks, which sells secure file and data exchange software, published a forecast on 5 January based on a survey of 225 security, IT and risk leaders across 10 industries and 8 regions. Every organisation surveyed had agentic AI on its roadmap. Sixty-three per cent said they could not enforce limits on the purposes their agents are used for, 60 per cent could not quickly shut down an agent that misbehaves, and 55 per cent could not isolate AI systems from wider network access. A third had no audit trail good enough to use as evidence, and 61 per cent had logs split across systems.
WRITER's survey of 2,400 people, published on 7 April, found 36 per cent of executives had no formal plan for supervising AI agents and 35 per cent could not immediately 'pull the plug' on a rogue one. Both companies sell products aimed at this problem, and both surveys record what respondents say about themselves. Gartner, in June 2025, expected more than 40 per cent of agentic AI projects to be cancelled by the end of 2027, and listed inadequate risk controls alongside cost and unclear value as the reasons.
What the rules require, and when
On 29 April the deployer obligations in the main AI laws had not yet started to apply. Under the EU AI Act as written, obligations for high-risk systems apply from 2 August 2026. The Digital Omnibus on AI would move them to 2 December 2027 for the stand-alone systems in Annex III and 2 August 2028 for AI in products covered by Annex I, but the trilogue on 28 April ended without agreement after about 12 hours, so the August date still stands in law.
When those obligations do apply, three articles cover much of what runtime governance does. Article 12 requires high-risk systems to record events automatically. Article 14 requires that the people overseeing a system can override or reverse its output and can interrupt it through a stop button or similar procedure that brings it to a halt in a safe state. Article 26 requires deployers to assign oversight to competent people with the authority to act, and to keep the system's logs for at least six months. Annex III covers recruitment, promotion, task allocation and credit scoring, so agents in HR and finance modules can fall within it. UK organisations are outside the Act's direct reach unless they operate in the EU.
In the US, Colorado's AI Act, which requires developers and deployers of high-risk systems making consequential decisions to take reasonable care against algorithmic discrimination, was due to take effect on 30 June 2026. xAI sued to block it in April, the Department of Justice moved to intervene on 24 April, and on 27 April a federal court suspended enforcement until the state's legislative session and related rulemaking finish and xAI's injunction motion is resolved. California's SB 53, in force since 1 January 2026, applies to frontier model developers. Its privacy regulator's rules on automated decision-making took effect the same day, and businesses already using such tools for significant decisions have until 1 January 2027 to comply.
The framework in outline
The framework has five parts: register every tool and agent, give each agent an identity with bounded delegation, decide each proposed action outside the model, put human approval where the risk is, and record every decision so that it can be replayed. The first two are set up before an agent runs, and the last three operate while it runs.
Two published designs share this shape. Oracle's framework of 23 April, written by Kishore Pusukuri under the title 'From Model Safety to Runtime Governance', puts an Agent Runtime Controller between the agent and its tools, which checks each proposed action against policy, identity, approvals and budget and returns ALLOW, ALLOW_WITH_REDACTION, REQUIRE_REVIEW or DENY. The Autonomous Action Runtime Management specification, published by Herman Errico on arXiv on 10 February, does the same job with five decisions: ALLOW, DENY, MODIFY, STEP_UP (require a person's approval) and DEFER (hold until there is more assurance). On 29 April, the date of this paper, Vanta contributed that specification to the Cloud Security Alliance's CSAI Foundation.
In both, the decision is made by software the agent cannot talk its way round, before the tool runs, and the reasons are kept.
Part one: register every tool and agent
An agent should reach only tools that someone has registered and approved, and the register should sit outside the agent's prompt, where the agent cannot change it. For each tool it records what the tool does, which data it can read or change, whether it reaches outside the organisation, how often it may be called and whether a person must approve its use.
The main platforms now build this in. Google's Agent Registry, announced on 22 April as part of its Gemini Enterprise Agent Platform, indexes an organisation's agents, tools and skills so that only approved ones are available in production. Microsoft's Agent 365 keeps a registry of agents and gives each one an identity in Microsoft Entra, and Microsoft said on 9 March that it would be generally available on 1 May at $15 per user per month. A provider's registry covers what runs on that provider, so an organisation using several needs its own list above them.
Call limits belong in the register because agents can loop. Oracle's framework names runaway loops that run up costs as a threat and calls it 'denial-of-wallet'. A limit on calls per agent, per user and per hour stops a loop at the tool, whatever the model is doing.
The register should start with the tools that write or send, since posting a journal or paying a supplier carries more risk than summarising a report.
Part two: give each agent an identity and bound its delegation
Every action needs to be traceable to the agent that took it and to the person who gave it authority. NIST's National Cybersecurity Center of Excellence published a concept paper on this on 5 February, with comments open until 2 April. It names the standards likely to be involved (OAuth and OpenID Connect, SCIM for identity lifecycle, SPIFFE and SPIRE, and attribute-based access control) and asks whether an agent's identity should be persistent or tied to a single task, and how a compromised agent's credentials are revoked. NIST launched a wider AI Agent Standards Initiative on 17 February.
The Model Context Protocol settles part of this. Since the June 2025 revision of the specification, an MCP server acts as an OAuth resource server and must validate the access tokens it receives, which decides whether a client may connect on a user's behalf. What the agent may then do is still up to the server and the application behind it.
In Oracle Fusion that application is the existing security model. Oracle's 26A readiness notes show that a user needs a job role with the Fai Genai Agent Runtime Duty to use an agent, and that the same role must already hold the privileges for whatever the agent does, such as managing external purchase prices. An agent acting for a user therefore gets no more than that user's roles allow, and any separation of duties conflict already inside a role comes with it. Oracle's Risk and Security Snapshot report shows conflicts within a single role and those arising from the combination of roles one person holds. Below the application, Oracle announced Deep Data Security for its AI Database on 24 March, which passes the end user's identity to the database so that row, column and cell rules apply to an agent's queries.
SAASiQ's view is that the snapshot is worth running on the roles that will carry the agent runtime duty before any agent is switched on, since a conflict in those roles becomes a conflict in the agent.
Part three: decide each action outside the model
The centre of the framework is a policy decision made on every proposed action, before the tool runs. The request that reaches the policy engine says who is acting (the agent and the person behind it), what it wants to do, to which resource, and in what context, such as the amount, the time and the approvals already given. The gateway or runtime carries out the engine's decision.
AWS's Policy in Amazon Bedrock AgentCore, generally available since 3 March in 13 regions including London, Frankfurt and Ireland, works this way. Policies sit in a policy engine attached to AgentCore Gateway, which intercepts traffic between agents and tools and evaluates each request before allowing or refusing it. Policies written in plain English are converted to Cedar, AWS's open-source policy language, and can check the caller, the tool and the input without any change to the agent's code. Google's Agent Gateway, announced on 22 April, applies security policy to interactions between agents and data and includes its Model Armor protections against prompt injection and data leakage.
The decision set needs more than yes and no. Oracle's ALLOW_WITH_REDACTION and AARM's MODIFY let an action go ahead with sensitive fields removed or a parameter changed. REQUIRE_REVIEW and STEP_UP send it to a person. DEFER holds it until more is known. DENY stops it, and the agent gets no chance to argue.
Policies need owners and versions. Several rules can apply to one action, so the order in which they are applied has to be written down, and each decision should record which version of which policy produced it, so that a refusal can be explained after the rule has changed.
The design also has to say what happens if the policy engine cannot be reached. Failing open lets every action through during an outage and failing closed stops the work, and the business owner of the process should choose before go-live.
Part four: put approvals where the risk is
Some actions should wait for a person. The approval step has to be enforced by the runtime, so the agent cannot proceed until it has an answer, and the person approving has to see what the agent proposes to do and why.
Oracle added a human approval node to workflow agents in AI Agent Studio in the 26A update, revised on 13 March. It pauses the workflow until a person approves, rejects or asks for a change, then carries on according to that decision. It is a node the builder adds, so an approval exists only where someone has designed one in, which makes the choice of where to put it a control decision in its own right.
Approvals slow work down, so the register from part one should say which actions need them, based on amount, data sensitivity and whether the action can be reversed. Payments and messages to customers cannot be undone and are the obvious candidates.
Stopping an agent altogether is a separate control. Article 14 of the EU AI Act asks for a stop procedure that brings a system to a safe state, and the Kiteworks and WRITER figures suggest many organisations cannot yet do it quickly. The stop should revoke the agent's credentials as well as halting its current task.
Part five: record decisions so they can be replayed
Each record should let someone reconstruct the action later: what the agent was asked to do, what it proposed, which policy version decided, who approved it if anyone did, and what happened. The bottom layer of Oracle's framework holds traces, provenance, tool identities and decision records for this purpose. AARM goes further and asks for tamper-evident receipts that bind the action, the session context, the matching policy, the decision and the outcome together cryptographically, so that they can be verified offline.
Retention is set by law and by the organisation's own audit needs, and the AI Act's minimum for deployers of high-risk systems is six months.
Logs split across systems were the problem 61 per cent of Kiteworks' respondents reported. The gateway, identity service, application and database each keep their own records, and they need a common identifier for each agent session so that one request can be followed through all of them. On the Oracle side, the agentic applications announced on 24 March record step-by-step actions and full execution paths, and AI Agent Studio has had a monitoring dashboard for sessions, errors and token usage since October 2025.
The same records serve NIST's AI Risk Management Framework of January 2023 and ISO/IEC 42001, published in December 2023, both of which expect evidence that controls operate.
Where to start and where it breaks
The first step is a list of the agents already running, including those that arrived switched on in a software update, with the tools each can call. From that list, take one workflow that writes to a system of record, register its tools, check the roles it runs under, put a policy decision in front of each write, and keep the log.
The framework has limits. A policy engine only allows or blocks the actions put to it, so an agent hijacked through a malicious document can still misuse the tools it is legitimately allowed, within the limits set, and the NCSC's view is that this risk can be reduced and never fully removed. Policies are software and need testing like any other, because rules that each look right can combine to block normal work or let through something nobody intended. The gateway, the policy engine and the log store all add running cost and a little time to each action, and someone has to own them.
Microsoft's Agent 365 is due for general availability on 1 May, Oracle's 26B update reaches the first customers' test environments the same day, and the EU's high-risk obligations apply from 2 August 2026 unless the Digital Omnibus is agreed and adopted before then.
SAASiQ - Intelligent Solutions for SaaS ©


