Provider, Use-Case and Agent Risk: A Framework for Buying Enterprise AI

Updated: 6 days ago
Title: Provider, Use-Case and Agent Risk: A Framework for Buying Enterprise AI
Date: 30 June 2026
Type: Paper
Author: SAASiQ (contact@saasiq.ai)
Word count: 2955 words
Reading time: 11 min
Published: 30-06-2026
The US Department of Commerce lifted its export controls on Anthropic's Claude Fable 5 on 30 June, 18 days after they led Anthropic to switch the model off for every customer. In the same fortnight OpenAI opened its GPT-5.6 models only to partners whose names it had shared with the US government, Google DeepMind and its partners offered up to $10 million for research into what happens when large numbers of AI agents deal with each other, and the EU completed its votes on delaying parts of the AI Act. This paper sets out a framework for assessing three risks together: dependence on a provider, the consequences of AI in each area of the business, and the authority given to agents.
Two frontier models held back in June
Anthropic released Fable 5 on 9 June. At 5:21pm Eastern on 12 June it received a directive from the Commerce Department suspending all access to Fable 5, and to Mythos 5, a more restricted model released alongside it, by foreign nationals inside or outside the United States. Anthropic had no reliable way to check a user's nationality in real time, so it disabled both models for all customers. Its other models, including Claude Opus 4.8, were not affected.
The control applied to the model itself. Customers who bought Claude through a cloud provider lost Fable 5 along with those who bought it direct, so the purchasing route made no difference. Fortune and The Wall Street Journal reported that the concern reached the administration from Amazon, whose researchers had prompted the model into identifying software vulnerabilities. Anthropic described the finding as 'a narrow potential jailbreak'. On 30 June it said it had received notice that the controls on both models had been lifted and that it would begin restoring access the following day.
OpenAI announced its GPT-5.6 family on 26 June, in three models called Sol, Terra and Luna, and at the government's request limited the rollout to 'a small group of trusted partners' whose participation had been shared with the government. The request followed an executive order signed on 2 June that asks developers of advanced models to share them with the government before a wider release. OpenAI called the limit a short-term step and said it was working to make the models generally available in the coming weeks, and TechCrunch reported the company's view that restrictions of this kind should not become the norm.
Neither event came from a technical fault or a commercial dispute. In both, a US government decision set when customers could use the newest model from one of the two largest US labs: Fable 5 was unavailable for 18 days, and on 30 June GPT-5.6 was still limited to approved partners.
Where the models and the chips come from
OpenAI and Broadcom unveiled Jalapeño on 24 June, the first chip OpenAI has designed itself, built to run its models once they are trained, with initial deployment planned by the end of 2026. Anthropic does not design chips. In early June Apollo and Blackstone completed a $35 billion private credit deal to fund Google-designed Tensor Processing Units which, Bloomberg reported, Anthropic then leases. Broadcom builds both OpenAI's chip and the Google chips.
A buyer's contract with a model vendor sits on top of the vendor's own supply contracts. SpaceX's registration statement, reported by TechCrunch on 20 May, showed Anthropic paying SpaceX $1.25 billion a month through May 2029 for compute at Colossus 1, the data centre near Memphis that xAI built, under a deal either side can end on 90 days' notice. Anthropic confidentially submitted its own draft S-1 on 1 June, and OpenAI said on 8 June that it had done the same. Once those filings are public, they have to disclose material risks and commitments of this kind.
The alternative to calling a vendor's hosted model is to run an open-weight model, one whose weights are published, on infrastructure the organisation controls. Nvidia released Nemotron 3 Ultra on 4 June, and Artificial Analysis scored it 47.7 on its Intelligence Index, against 39.2 for Google's Gemma 4 31B and 33.3 for OpenAI's gpt-oss-120b, with Moonshot AI's Kimi K2.6, from China, on 53.9. OpenRouter's review of the open-weight models that mattered in June, published on 27 June, picked four. Three came from Chinese labs (DeepSeek V4 Flash, Z.ai's GLM 5.2 and MiniMax M3) and the fourth was Nemotron 3 Ultra. An organisation whose procurement rules exclude models of Chinese origin has a shorter list to choose from, and on these benchmarks a weaker one.
AI in work where errors cost more
OpenAI said on 18 June that GPT-5.5 Instant, the default model in ChatGPT, now performs on a par with its Thinking models on health questions, and that more than 230 million people a week bring health and wellness questions to ChatGPT. It reported a 71 per cent fall over two months in the share of health responses flagged for factual problems. Dataconomy, reporting the update on 19 June, noted that the evaluation results had not been released for outside review, so the figures are OpenAI's own.
Some of this capability arrives in software an organisation already runs. Google switched Gemini 3.5 Flash on by default in Gemini Enterprise in all regions on 9 June and removed the option to turn it off. Microsoft's MAI-Transcribe-1.5, one of the in-house models it announced at Build on 2 June, is built into Teams. In both cases the vendor chose the model behind the feature, and a later product update can change it.
The European Parliament adopted the Digital Omnibus on AI on 16 June by 423 votes to 57, with 174 abstentions, and the Council gave final approval on 29 June. Once published in the Official Journal it moves the obligations for high-risk systems listed in Annex III, which include AI used in recruitment and credit decisions, to 2 December 2027, and those for AI in products regulated under Annex I to 2 August 2028. On 30 June it had not been published, so the original date of 2 August 2026 still stood.
Whatever the date, a deployer of a high-risk system has to assign human oversight and keep the logs the system generates, under Article 26. The Article 50 transparency rules, which require people to be told when they are dealing with an AI system, apply from 2 August 2026. UK organisations are outside the Act's direct reach unless they operate in the EU.
Agents that deal with other agents
On 11 June Google DeepMind, Schmidt Sciences, the Cooperative AI Foundation and the UK's Advanced Research and Invention Agency, with support from Google.org, opened a call for up to $10 million of research into multi-agent AI safety. The premise is that millions of agents built by different organisations will communicate, negotiate and transact with one another. The announcement points out that most safety evaluations analyse models in isolation. It funds work in four areas: test environments, the science of how agent networks behave, infrastructure including protocols for agent identity and reputation, and oversight of agents once deployed. Applications close on 8 August.
The OWASP Top 10 for Agentic Applications, published in December 2025 by more than 100 contributors, includes Identity and Privilege Abuse (ASI03), where an agent acts with more authority than it should have or on stale credentials, Insecure Inter-Agent Communication (ASI07), where agents exchange messages without adequate authentication or policy checks, and Cascading Failures (ASI08), where one error spreads through connected systems.
NIST's National Cybersecurity Center of Excellence published a concept paper on agent identity and authorisation in February, and as Biometric Update reported it, the paper wants every agent action traceable to the non-human identity that performed it and to the person who delegated the permissions. In a survey for Okta published on 27 May, of 292 executives and 492 knowledge workers in seven countries including the UK, 58 per cent of executives said their company had had an AI-related security issue or close call in the previous 12 months, and only 34 per cent of organisations applied the same security controls to agents as to human staff.
Before starting
The framework needs three things in place. There has to be an owner for AI with the authority to approve or refuse a use, and a data classification that says what may leave the organisation and on what terms. There also has to be a list of where AI is already running, including features that arrived switched on in software the organisation already licenses.
None of this needs machine learning expertise. For an Oracle Fusion estate it does need a working knowledge of the role model and of how identities are provisioned, because Fusion's agents run inside that model, as step three sets out.
Step one: map what each AI use depends on
For each AI use, record the provider, the model and its version, the route by which it is bought (direct, through a cloud marketplace, or embedded in a SaaS product), where the data goes and under what contract. Record also who chooses the model. For embedded features that is usually the vendor, as with Gemini in Gemini Enterprise and the MAI models in Microsoft 365.
The Fable 5 suspension shows why the model belongs in the record and not only the channel. Customers buying through cloud providers lost access along with those buying direct, because the control attached to the model. A register that listed only suppliers would have shown several vendors and missed that they all depended on one model.
Smaller interruptions belong in the record too. On 19 June Anthropic said about 3 per cent of Claude Code Max and Pro users had hit a bug that showed an incorrect weekly usage limit and in some cases stopped them sending messages, until it was fixed.
The output of this step is a list that answers, for each use, what stops if the model becomes unavailable and what the fallback is. Where the answer is 'nothing has been tested', that is the first gap to close.
Step two: grade each use by the consequence of an error
Grade each use by what happens when the AI is wrong, whatever the technology behind it. A low grade covers drafting internal text that a person reads before use. A middle grade covers recommendations that a person acts on, such as a suggested match between an invoice and a purchase order. A high grade covers uses that affect a person's employment, credit or health, and anything that moves money or changes records without review.
The grade belongs to the use. GPT-5.5 Instant, ChatGPT's default model, drafts emails and also answers the health questions that, by OpenAI's count, more than 230 million people a week bring to ChatGPT. So a model approved for one use is not approved for another by default, and a single model may sit in all three grades.
High-grade uses need a named person with the authority to override the output and the time to do it, and logs kept for as long as the organisation may have to explain a decision. Where the organisation operates in the EU, the high grade overlaps with Annex III for uses such as recruitment and credit, and Article 26 already requires both of those things from deployers of high-risk systems.
Grades change when a feature's scope changes, for example when it moves from suggesting an action to taking it. Vendors' quarterly updates are one route for such changes. Oracle applied its 26B update to Fusion test environments on 1 May and production on 15 May, and some of the new agentic features arrived switched off: the Cost Accounting Close Workspace, for example, stays hidden until an administrator sets a profile option. Each update's new AI features can then be checked against the grading before anyone switches them on.
Step three: give each agent an identity and a ceiling
An agent with access to systems is an identity and needs the same controls as a person: a named owner, access set to what the task needs, and credentials that can be revoked. It needs them more than a person does, because it runs continuously and acts faster than anyone can review.
Oracle Fusion sets the ceiling through the user. Oracle's 26B readiness notes say that to use agents on a Fusion page, a user's job role must contain the Fai Genai Agent Runtime Duty and must already give access to the pages where the agents run. An agent working for a user gets no more access than the user has. So separation of duties for Fusion agents rests mainly on the roles of the people they run for, which existing separation of duties analysis already covers.
Below the application, Oracle's Deep Data Security, available in Oracle AI Database 26ai since 1 May, passes the identity of the user and the agent to the database at runtime, and SQL policies decide which rows, columns and cells come back. Oracle's managed MCP servers, announced on 12 May, come with three built-in roles. Oracle suggests limiting MCP_User to named, fixed reports, while MCP_Operator can also run ad hoc SQL through a run-sql tool.
Chains of agents need their own limits. An agent that can call another can reach whatever the second agent can reach, which is the ground ASI07 and ASI08 cover. The controls are an allowlist of which agents may call which, authentication on messages between them, and a record of each step in the chain that reaches the systems the security team already watches.
Monitoring an agent does not constrain it. Oracle's AI Agent Studio has a token usage view that measures consumption of premium models, which helps with cost but does not stop an agent doing anything. Microsoft said in May that runtime blocking of agent behaviour in Agent 365 would reach public preview in June. A dashboard shows what an agent did after the event, and only a block or a permission stops it.
Step four: keep the model replaceable
Applications should call an internal routing layer and not a specific model's endpoint, so that changing model is a configuration change followed by a test run. Prompts, test sets and integration logic should belong to the organisation and not sit only in one vendor's tools.
In June the models left untouched were the previous generation. Opus 4.8 was outside the directive, and OpenAI's existing models were outside the GPT-5.6 limits, so a second provider's newest model was not a sure fallback either. A fallback only helps if it has been run against the organisation's own tasks, with prompts kept current.
For data that must stay inside a boundary, open-weight models hosted by the organisation or its cloud provider are sometimes the only option. OCI Generative AI has let customers import their own models since 12 November 2025, and in June Oracle added Nemotron 3 Ultra (5 June) and DeepSeek V4 Flash and V4 Pro (15 June) to the list of models that can be imported. A release note dated 30 June adds private endpoints for imported models, so their traffic can stay on a private network. For UK public bodies the service has run in Oracle's UK Gov South region in London since 21 August 2025.
On the applications side, AI Agent Studio for Fusion Applications has supported models from OpenAI, Anthropic, Cohere, Google, Meta and xAI since October 2025, so a Fusion agent can be pointed at a different provider's model. The fallback still has to meet the use's data rules, and the easiest model to switch to is not always one with a compliant deployment.
Step five: put it in the contract and review it
Anthropic's statement on 12 June told customers it was working to restore access and gave no other instructions. Buyers can ask their AI vendors, and the cloud providers that resell them, how customers are told when a model is suspended for legal or regulatory reasons, and whether fees are credited for the time it is unavailable. They can also ask for notice periods on retirement: Anthropic commits to at least 60 days for publicly released models, and OpenAI to at least six months for generally available ones unless safety or compliance requires faster.
The mapping, the grades and the fallbacks should be reviewed each quarter, and again whenever a provider withdraws, restricts or reprices a model the organisation uses. Each decision should be recorded with its date and the test results behind it, so that an auditor asking how an output was produced gets an answer.
In SAASiQ's view, long fixed commitments to a single model are worth avoiding while events like June's can take a model away for weeks at a time.
What it costs and where it breaks
The framework takes time from people who have other jobs. The inventory needs input from every team that uses AI, the test sets need real cases with agreed answers, and the routing layer is software someone has to build and maintain. Giving each agent its own identity adds work for the identity team.
It also has limits. Features embedded in a vendor's product cannot be rerouted, so for those the controls are the grade, the switch-on decision and the contract. Grades go stale as uses change, and a register nobody maintains is soon wrong. This paper does not cover contract drafting or the detailed control mapping a particular regulator will expect.
Anthropic said it would begin restoring Fable 5 on 1 July, OpenAI has said GPT-5.6 will be generally available in the coming weeks, the Omnibus has to be published in the Official Journal before 2 August to take effect in time, and applications to the multi-agent safety fund close on 8 August.
SAASiQ - Intelligent Solutions for SaaS ©


