Government AI Rules as a Governance Framework for Enterprise Deployments

Updated: 6 days ago
Title: Government AI Rules as a Governance Framework for Enterprise Deployments
Date: 14 May 2026
Type: Paper
Author: SAASiQ (contact@saasiq.ai)
Word count: 2929 words
Reading time: 11 min
Published: 14-05-2026
The Government Digital Service launched GOV.UK Chat in the GOV.UK app on 14 May, after safety checks with the AI Security Institute and a soft launch in which more than 7,800 people asked over 15,000 questions. It was built under the AI guidance the UK government wrote for its own use in February 2025. The US federal government and the EU have written their own rules, and many of their controls are the same. This paper sets out the controls they share as a framework an enterprise can apply to its own AI deployments.
The UK guidance and GOV.UK Chat
GDS published its AI guidance for government on 10 February 2025. It runs to 118 pages, replaced the Generative AI Framework for HM Government of January 2024, and widened the scope from generative AI to AI in general. Contributors included central departments, GCHQ, the ICO and the NHS, the Alan Turing Institute, and Amazon Web Services, Google, IBM and Microsoft.
It rests on ten principles. Civil servants should know what AI is and what its limitations are, use it lawfully, ethically and responsibly, know how to use it securely, keep meaningful human control at the right stages, and understand how to manage the full AI life cycle. The other five cover using the right tool for the job, being open and collaborative, working with commercial colleagues from the start, having the skills to implement and use AI, and using the principles alongside the organisation's own policies with the right assurance in place.
GOV.UK Chat is the guidance's main case study. GDS conceived it in July 2023 as a retrieval system over GOV.UK content, calling OpenAI's models through an API and hosted mainly on Google Cloud, with business guidance as the first subject because it cuts across so many departments. Hundreds of users tested it through private links. Nearly 70 per cent found the answers useful and about 65 per cent were satisfied, but answer accuracy was 80 per cent and the system occasionally produced 'hallucinated' responses, so GDS concluded that accuracy had to improve before the tool went on the site.
The version launched on 14 May followed a soft launch on 26 March with no publicity. GDS says it tells users not to share personal information and filters out any that is provided, signposts when users should check the original guidance, and does not attempt to give advice. GDS describes the route from prototype to launch as the Scan, Pilot, Scale approach in the government's AI Opportunities Action Plan, and it took a little under three years.
The US memos and the NIST framework
In the US, the Office of Management and Budget issued two memoranda on 3 April 2025. M-25-21 covers how agencies use AI and replaced the Biden administration's M-24-10 of March 2024. M-25-22 covers how they buy it. Each agency had to name a Chief AI Officer within 60 days, and the larger agencies covered by the CFO Act had to convene an AI governance board within 90, chaired at Deputy Secretary level with the Chief AI Officer as vice-chair. Policies on the acceptable use of generative AI were due within 270 days.
M-25-21 defines 'high-impact' AI as AI whose output serves as a principal basis for decisions or actions with legal, material, binding or significant effect on a person's rights or privacy, their access to education, housing, insurance, credit, employment or government services, human health and safety, critical infrastructure, or strategic assets. High-impact uses must meet a set of minimum practices, and agencies had 365 days, to early April 2026, to document that they do. A use that does not comply must be discontinued safely.
OMB published the consolidated 2025 inventory on 14 April: 3,611 use cases across 56 agencies, up from 1,757 in 2024, of which 445 were classed as high-impact. According to Nextgov, the Department of Veterans Affairs accounted for 215 of the high-impact uses, Microsoft Copilot appeared in 102 use cases and Anthropic's Claude in 25. Defence and intelligence uses are exempt from reporting.
A third memo, M-26-04 of 11 December 2025, implements the July 2025 executive order on what the administration calls unbiased AI principles, truth-seeking and ideological neutrality. In any solicitation for a large language model, agencies must ask the vendor for its acceptable use policy, its model, system or data cards, its end-user resources and a way to report problem outputs, and they must ask for the same information when a model comes built into another software product. Agencies had until 11 March to update their procurement policies.
Underneath the memos sits NIST's AI Risk Management Framework, released on 26 January 2023 and organised around four functions: govern, map, measure and manage. Its Generative AI Profile, NIST AI 600-1, followed on 26 July 2024, and a preliminary draft Cyber AI Profile, NIST IR 8596, on 16 December 2025. The White House's AI Action Plan of July 2025 called for the framework to be revised, and NIST lists that revision as under way.
The EU position on 14 May
The EU AI Act is law, and it reaches organisations that use AI as well as those that build it. Its prohibitions and its AI literacy duty have applied since 2 February 2025. For 'deployers' of high-risk systems, Article 26 requires use in line with the provider's instructions, human oversight by people with the competence and authority to carry it out, retention of the logs the system generates for at least six months, and notice to workers' representatives before a high-risk system is used in the workplace. Article 27 adds a fundamental rights impact assessment for public bodies, for private organisations providing public services, and for deployers using AI to score creditworthiness or price life and health insurance.
When those high-risk duties apply is not yet settled. On 7 May negotiators for the Council and the Parliament reached a provisional agreement on the Digital Omnibus on AI, which moves the obligations for stand-alone high-risk systems in Annex III, covering uses such as employment, education and biometrics, from 2 August 2026 to 2 December 2027. High-risk AI embedded in products regulated under Annex I moves to 2 August 2028. Both institutions still have to adopt the text formally, which they intend to do before 2 August, and until then the original date stands. UK organisations are outside the Act's direct reach unless they operate in the EU.
The UK has no equivalent statute. The King's Speech on 13 May announced a Regulating for Growth Bill that would give ministers powers to run regulatory sandboxes, in which AI products can be tested in real conditions under temporarily relaxed rules.
Step one: a board and a named owner
Both the UK guidance and M-25-21 start with structure. GDS lists an AI strategy and adoption plan, a governance board of senior leaders and experts to set principles and review and authorise uses of AI, a short set of principles, a register of use cases, and a sourcing strategy that says which capabilities will be built in-house. M-25-21 requires its boards to include IT, cybersecurity, data, budget, legal, privacy and civil liberties, and, where relevant, procurement and the programme offices using the AI.
In a company the equivalent is a board chaired by an executive with the authority to stop a project, with the CIO, the CISO, the data protection officer, legal, procurement, HR and the finance owner represented. Its first task is to agree the criteria a use case must meet before a proof of concept starts.
The UK guidance asks for a senior responsible owner for each AI project, accountable for its use and for acting on reports of harm throughout its life cycle. M-25-21 goes further for high-impact uses and requires a named individual to sign the risk acceptance. In an enterprise that person should be the business owner of the process the system supports, since the business owner answers for the outcome, with the technical lead reporting to them.
Step two: an inventory with risk tiers
Both governments keep a public register. US agencies update their AI use case inventories every year. In the UK, all government departments, and arm's length bodies that deliver public or frontline services, must publish a record under the Algorithmic Transparency Recording Standard for each algorithmic tool they use to support decisions.
The register needs a risk tier, and the M-25-21 definition works for a company with light editing. A system is high-impact where its output is a principal basis for a decision with a legal or significant effect on someone's employment, credit, insurance or access to a service, or where it affects health, safety or critical infrastructure. EU Annex III covers much of the same ground, including recruitment, promotion and termination, task allocation and monitoring of workers, and credit scoring. Anything in that tier gets the minimum practices in step three.
The inventory has to include AI that arrives inside software the organisation already licenses. In the federal inventory Microsoft Copilot alone accounted for 102 use cases. In an ERP estate the equivalent is the quarterly update: Oracle's 26B update for Fusion Cloud reached test environments on 1 May and is scheduled for production on 15 May, and its new agentic applications stay switched off until an administrator enables them. The register should record who switched each feature on, when, and for which users.
Step three: minimum practices before go-live
M-25-21 sets out the minimum practices for high-impact AI. They are pre-deployment testing with a risk mitigation plan, and an AI impact assessment covering the purpose and expected benefit, the quality of the data, the potential impacts, a reassessment schedule, costs, the result of an independent review and a signed risk acceptance. After go-live come ongoing monitoring, training for the people who operate the system, human oversight with a fail-safe where practicable, a route for affected people to get a timely human review and appeal, and a way for users and the public to give feedback.
Two details carry straight over to software bought as a service. Where an agency cannot see the source code, model or data, it must test by other means, for instance by querying the service and observing its outputs, or by giving the vendor evaluation data and getting the results back. And the independent reviewer must be someone in the agency who was not involved in developing the system.
The UK guidance covers oversight as well. Humans should validate any high-risk decisions influenced by AI, and where review in real time is not possible, as with a chatbot, human control has to be exercised at other stages. GOV.UK Chat applies that by not giving advice and by sending users back to the source guidance.
M-25-21 exempts pilots from the minimum practices if they are limited in scale and duration, certified by the Chief AI Officer and tracked centrally, and where possible let people opt in or out. An enterprise can use the same rule, with every pilot on the register with an end date, so that nothing reaches production without passing the minimum practices. For systems that fall within Annex III, the EU's human oversight and log retention duties in Article 26 sit on top, whichever date is finally adopted.
Step four: security controls
The NCSC and the US Cybersecurity and Infrastructure Security Agency published Guidelines for secure AI system development on 26 November 2023, co-sealed by 23 agencies from around the world and organised in four stages: secure design, development, deployment, and operation and maintenance. DSIT's AI Cyber Security Code of Practice followed on 31 January 2025 with 13 principles, running from awareness of AI threats through securing the supply chain to the proper disposal of data and models. ETSI turned it into a technical specification, TS 104 223, in April 2025, covering five life cycle phases that end with end of life.
On 8 December 2025 the NCSC published a blog arguing that prompt injection may never be mitigated in the way SQL injection can be, because a large language model does not separate the instructions it is given from the data it reads. Its advice was to concentrate on reducing the impact of a successful attack. For an agent with access to finance or HR systems, that means credentials scoped to the task, approval steps before anything that moves money or changes a record of employment, and logs that show what the agent did.
NIST has two projects on agents under way. Its National Cybersecurity Center of Excellence published a concept paper on agent identity and authorisation on 5 February, with comments closing on 2 April. Its Center for AI Standards and Innovation launched an AI Agent Standards Initiative on 17 February.
Governments are also testing models before release. On 5 May the Center for AI Standards and Innovation announced that Microsoft, Google and xAI would give it access to new models before release for national security testing, extending agreements it signed with OpenAI and Anthropic in 2024. In the UK, GDS carried out safety checks on GOV.UK Chat with the AI Security Institute before launch.
Step five: contract terms
M-25-22 applies to contracts awarded from solicitations issued 180 days or more after 3 April 2025, and to options exercised on existing contracts after that point. It requires terms that set out the ownership and IP rights of each party, and that permanently prohibit the vendor from using non-public agency inputs and outputs to train publicly or commercially available AI without explicit consent. Against lock-in it lists knowledge transfer, data and model portability, rights to code and models produced under the contract, and transparent licensing and pricing.
M-25-22 is specific about testing. Agencies must be able to evaluate performance regularly, quarterly or twice a year for example, using their own test data that the vendor cannot see, and contracts must not stop them disclosing internally how the vendor tests. Agencies are encouraged to require vendors to meet performance standards before deploying a new version, to roll back if a new version falls short, and to give notice before new AI features are added to a service under contract. SAASiQ's view is that the notice and roll-back terms are the ones most missing from enterprise SaaS contracts, because quarterly updates now add AI features to systems that are already in production.
The UK guidance calls AI an emerging market from a commercial perspective and asks buyers to consider how to avoid lock-in, who holds IP rights in anything developed, and what liabilities are acceptable. The M-26-04 list, starting with the vendor's acceptable use policy, belongs in the same contract file.
The US government's dispute with Anthropic began with an acceptable use policy. Anthropic's policy ruled out mass domestic surveillance and fully autonomous weapons, and when it would not agree to use of Claude for 'all lawful purposes', President Trump directed federal agencies on 27 February to stop using its technology, and formal supply-chain risk designations followed on 3 March. Judge Rita Lin blocked the directive and one of the designations on 26 March, but on 8 April the D.C. Circuit refused to pause the other while the case is heard. Inside the Pentagon, Claude is being removed on a 180-day timetable, while other agencies and private customers are unaffected.
Step six: change, monitoring and retirement
The UK guidance asks teams to know how to update a system and how to close it down securely at the end of its useful life, to monitor for drift, bias and hallucination, and to keep fallback processes so that critical services continue if a change has to be reverted or an AI system terminated. M-25-21 requires monitoring designed to detect changes to the system after deployment and changes to the context of use or the data it relies on.
M-25-22 asks agencies to set criteria for retiring a system, such as changes in cost, in the agency's needs, in the vendor's requirements or in model performance. When a contract is not renewed, the agency and the vendor are to agree the format and usability of the data, anything that could end access to it, and a plan for transferring it. The DSIT code's final principle covers the disposal of data and models at the same point.
The Pentagon's removal of Claude is a case of an organisation replacing a model to a deadline it did not set. An enterprise fallback plan should name the replacement model or manual process for each high-impact system, tested in advance.
GDS ran GOV.UK Chat without publicity for seven weeks and says it performed at the same level as in the most recent pilot before it announced the launch.
What it costs and where it breaks
The framework has running costs. The board needs senior time, every high-impact system needs an impact assessment and an independent reviewer, and testing needs real cases with agreed answers. GDS recorded that its manual quality checks for GOV.UK Chat would not scale, and proposed a bank of quality-assessed questions for semi-automated testing.
Government rules also change with governments. M-24-10 lasted about a year before M-25-21 replaced it, M-26-04 lapses after two years unless extended, the NIST framework is being revised, and the EU dates moved on 7 May. The six steps above use only controls that appear in at least two of the three regimes, so a change to any one text leaves most of the framework standing.
The D.C. Circuit hears argument on Anthropic's designation on 19 May, and the Digital Omnibus goes to the Parliament and the Council for formal adoption, which both intend to complete before 2 August.
SAASiQ - Intelligent Solutions for SaaS ©


