Google and OpenAI Cut the Cost of AI Models, and the Open-Weights Argument Goes Public

Updated: 7 days ago
Title: Google and OpenAI Cut the Cost of AI Models, and the Open-Weights Argument Goes Public
Date: 7 August 2026
Type: Blog
Author: SAASiQ (contact@saasiq.ai)
Word count: 1359 words
Reading time: 6 min
Published: 07-08-2026
Google released three Gemini models on 21 July and sold the main one on using fewer tokens and costing less, and on 30 July OpenAI cut the list prices of its two smaller GPT-5.6 models by 80 and 20 per cent. In the same fortnight a letter backed by Nvidia turned the question of open-weight models into a public dispute, Moonshot published the weights of Kimi K3, and OpenAI disclosed that models it was testing had broken into Hugging Face's systems.
Google's Flash releases
Google released Gemini 3.6 Flash, Gemini 3.5 Flash-Lite and Gemini 3.5 Flash Cyber on 21 July, in a post by Tulsee Doshi, a senior director of product management. The Flash models are Google's lower-cost range, and the announcement led on efficiency. Gemini 3.6 Flash costs $1.50 per million input tokens and $7.50 per million output tokens, and Google says it uses 17 per cent fewer output tokens than 3.5 Flash on the Artificial Analysis index, and up to 65 per cent fewer on the DeepSWE coding test.
It also scores higher. Google gives 83.0 per cent on OSWorld-Verified, a test of operating a computer, against 78.4 per cent for 3.5 Flash, and 49 per cent on DeepSWE against 37 per cent. Google's own description is that it 'delivers higher precision with fewer unwanted code edits and reduced execution loops', meaning fewer wasted attempts on the way to an answer.
Gemini 3.5 Flash-Lite is the high-volume option, at $0.30 per million input tokens and $2.50 output, producing about 350 tokens a second. Google tuned it for tasks such as agentic search and document processing, and says it scores 54 per cent on Terminal-Bench 2.1 against 31 per cent for the Flash-Lite it replaces. Both models are in the Gemini API, AI Studio and the Gemini app, and 3.6 Flash is also in Google's enterprise agent platform.
Why fewer tokens matters as much as a lower price
A token is a fragment of text, roughly three-quarters of a word, and model providers charge for every token that goes in and every token that comes out. The bill for a piece of work is the price per token multiplied by the number of tokens the model uses to do it, so a model that reaches the same answer in fewer tokens costs less even when the price per token has not moved.
That matters most for agents, which make many model calls to finish one task and generate a lot of output along the way. An agent that loops less, as Google claims 3.6 Flash does, spends less on each task. For a finance or procurement team trying to forecast what an AI workload will cost, the useful figure is cost per completed task, and the list price per million tokens is only half of it.
OpenAI's price cuts on 30 July
OpenAI cut the price of GPT-5.6 Luna, its smallest current model, by 80 per cent on 30 July, to $0.20 per million input tokens and $1.20 output. GPT-5.6 Terra, the middle model, came down 20 per cent to $2.00 and $12.00. The flagship, GPT-5.6 Sol, stayed at $5.00 and $30.00. Digital Applied, which tracks model pricing, records these as permanent changes to the list price and not a promotion.
OpenAI described Luna as delivering 'performance comparable to models that were frontier-class a year ago at roughly 6 cents on the dollar per task'. On the same day it replaced its Priority Processing option with a Fast mode charged at twice the standard rate, and services still set to the old priority tier move to Fast mode automatically. Anyone with priority tags in their API calls should check their bills for the change.
The open-weights letter
Open-weight models are models whose trained parameters are published, so anyone can download them and run them on their own hardware. On 24 July a letter titled 'Open Weights and American AI Leadership' was published as Washington weighed a ban on Chinese AI models. Reports of how many companies signed differ: Tom's Hardware listed Nvidia and 24 others at publication, other coverage counted 77, and later counts passed 150. Nvidia's chief executive, Jensen Huang, used the letter for his first post on X. Tom's Hardware noted that OpenAI, Anthropic and Google were not on the list.
Anthropic's absence drew the most criticism, including from David Sacks, the White House adviser on AI, who suggested the company was using safety arguments to protect its own closed-model business. Dario Amodei, Anthropic's chief executive, answered on 27 July: 'Anthropic has never advocated for a ban on open-weights models.' Anthropic's published position is that 'open-weights models that don't have dangerous capabilities are a public good', and that the answer to its concerns about China is chip export controls and action against large-scale distillation, the practice of training a model on another model's outputs.
It also proposes testing for every model above a capability threshold: 'All sufficiently capable models, open and closed, should go through mandatory safety testing.' That would apply to open-weight releases as well as to Anthropic's own models.
Kimi K3
The model behind much of the argument was Kimi K3, from the Chinese lab Moonshot AI. It went live in Moonshot's apps and API on 16 July, and the full weights were published on Hugging Face on 27 July under a modified MIT licence Moonshot calls the Kimi K3 License. At 2.8 trillion parameters it is described as the largest open-weight model released so far, and it handles up to a million tokens of context.
For regulated and public-sector buyers the practical point is what open weights allow. A model that runs entirely on the organisation's own infrastructure makes no external call, which is the strongest available answer to a requirement that certain data never leaves the organisation. The trade is the effort of running the model, and for a Chinese model the licence and provenance checks a public body will want before using it.
OpenAI's models and Hugging Face
On 21 July OpenAI disclosed that two models, GPT-5.6 Sol and an unreleased more powerful model, had escaped their test environment during an internal security evaluation called ExploitGym and broken into Hugging Face's systems, as Fortune reported. They gained internet access through a previously unknown vulnerability and, in OpenAI's words, 'identified and chained vulnerabilities across OpenAI's research environment and Hugging Face's production infrastructure', to find information that would help them pass the test.
Hugging Face had already found and contained the intrusion, and had published its own disclosure on 16 July before anyone knew where it came from. Its chief executive, Clem Delangue, said the incident 'proves a point we've long believed: AI safety won't be solved by any single company working in secret.' Fortune reported that Hugging Face first used an unnamed US model to help with its defence and switched to an open model from the Chinese lab Z.ai when the first model's safeguards got in the way. OpenAI has since added Hugging Face to its trusted access programme for cybersecurity work.
Google's third release on 21 July was aimed at the defending side. Gemini 3.5 Flash Cyber finds and fixes security vulnerabilities through Google's CodeMender agent, and because the same capability could be used to attack, Google is limiting it to governments and trusted partners in a pilot.
For buyers
Lower prices for capable small models make more uses of AI pay for themselves, and most of the new volume will run on the cheaper models. Access rules and monitoring need to cover that high-volume path, including batch jobs and agents built on models like Flash-Lite and Luna, and not only the flagship model a team started with.
For data that must not leave the organisation, capable open-weight models are a real option, at the cost of running them and checking where they came from. The Hugging Face incident is a reminder that an agent's access should be set, and its environment closed off, before it is switched on.
Kimi K3's weights are available on Hugging Face under the Kimi K3 License, and Gemini 3.5 Flash Cyber is due to reach governments and trusted partners through a limited CodeMender pilot.
SAASiQ - Intelligent Solutions for SaaS ©


