This Week's Releases: GPT-6 Sol and Luna, Claude Opus 5.5, Grok 4.7 and Gemini 3.8 Voices

Updated: 6 days ago
Title: This Week's Releases: GPT-6 Sol and Luna, Claude Opus 5.5, Grok 4.7 and Gemini 3.8 Voices
Date: 24 September 2026
Type: Blog
Author: SAASiQ (contact@saasiq.ai)
Word count: 1585 words
Reading time: 6 min
Published: 24-09-2026
Four AI labs released new models on 21 and 22 September. OpenAI brought out the mid and small GPT-6 models at half the price of the ones they replace, Anthropic released Claude Opus 5.5, xAI shipped Grok 4.7 and Xiaomi published open-weight models that anyone can download. Google added new voice models and made its video avatars generally available to businesses, and Oracle set out planned AI features for healthcare at its summit in Orlando.
GPT-6 Sol and GPT-6 Luna
OpenAI released GPT-6 Sol and GPT-6 Luna on 22 September, 19 days after its flagship GPT-6 Astra. Sol is the mid-sized model and Luna the small one. Sol costs $2 per million input tokens and $10 per million output tokens, half the $4 and $20 charged for GPT-5.6 Sol. Luna costs $0.10 and $0.50, down from $0.20 and $1.20. Astra, for comparison, is $10 and $50.
A token is a small piece of text, roughly three-quarters of a word, so a million tokens is about 750,000 words. Input is what is sent to the model, such as a question and any documents attached to it, and output is what the model writes back. Material sent over and over, such as a standing set of instructions, can be cached, and cached input on the new models is 90 per cent cheaper.
Published figures for Sol include 33.2 per cent on AutomationBench 1.0.6, a test of multi-step business tasks, at a cost of $0.27 per task, and 68.8 per cent on DeepSWE 1.1, a software engineering test.
Both models are in the API now as gpt-6-sol and gpt-6-luna, and are rolling out gradually in the paid ChatGPT plans. Luna is also available to Free and Go users through the desktop app. For a finance or HR team pushing large volumes of invoices or case notes through an API, the same workload now costs about half as much to run.
Claude Opus 5.5
Anthropic released Claude Opus 5.5 on 22 September. It costs $4 per million input tokens and $20 per million output tokens. Storing material in the cache costs $5 per million tokens and reading it back costs $0.20. Anthropic says the model is 40 per cent cheaper than Opus 5 on typical workloads, because it uses fewer tokens to do the same job, and that it produces output more than 30 per cent faster.
Anthropic's figures show it at 66.4 per cent on Terminal-Bench 4.0 against 52.3 per cent for Opus 5, and 81.8 per cent on OSWorld 2.0, which tests a model operating a computer, against 74.0. In one example Anthropic gives, when the model was asked to cut load times across every page of a web app, "Opus 5.5 succeeded 39 of 40 times".
On safety, Anthropic says Opus 5.5 is the first of its models tested by external evaluators, Frontier Design and METR, and that it scored highest on its automated behavioural audit. It reports 85 per cent fewer attempts to get around the limits set for it than Opus 5, and describes it as "much less likely than recent models to take hard-to-reverse actions". Businesses letting an agent post or delete records in a live system should test that claim in a sandbox first.
It is available on the paid Claude.ai plans, in the API as claude-opus-5-5, and through AWS, Google Cloud and Microsoft Azure.
Grok 4.7
xAI released Grok 4.7 on 21 September at the same price as Grok 4.6: $2 per million input tokens and $6 per million output tokens, with cached input at $0.50 according to The Decoder. A fast variant runs at twice the speed for twice the price. It is available through the Grok API, Cursor and cloud platforms.
xAI's own benchmarks show gains over 4.6 on every test it published, including 37.6 per cent on Terminal-Bench 4.0 against 20.3, and 71.0 per cent on DeepSWE v1.1 at high effort against 65.2. It also claims the best resistance to jailbreaks, which are attempts to talk a model out of its safety rules, with 3.3 per cent of risky dual-use prompts getting through on HackerBench v0.3.
The Decoder puts it lower, at 26 per cent on Terminal-Bench 4.0, apparently from independent testing, against 60 per cent for GPT-6 Astra and 55 per cent for Claude Fable 5.1. The Decoder also reports a score of 46 on the Artificial Analysis Intelligence Index, against 53 for the leading models.
Xiaomi MiMo-V2.6
Xiaomi released MiMo-V2.6 Pro and MiMo-V2.6 Flash on 21 September under the MIT licence, a permissive licence that allows commercial use. The weights, the trained model itself, are on Hugging Face, so an organisation can download the models and run them on its own hardware or a cloud of its choice.
Pro has 1.02 trillion parameters in total, of which 42 billion are active for any one request, and costs $0.435 per million input tokens and $0.87 per million output tokens through Xiaomi's API. Flash has 310 billion in total and 15 billion active, at $0.14 and $0.28. Parameters are the internal values a model learns in training, and a model that uses only part of them for each request is cheaper to run.
Pro scores 46 on the Artificial Analysis index, tied for the top open-weight model, level with Grok 4.7 and seven points behind the leaders. VentureBeat reports DeepSWE v1.1 scores of 71.9 for Pro and 67.9 for Flash. Suppliers are already building products on open weights from Chinese labs: Harvey, the legal AI company, has moved its flagship to an in-house model built on Moonshot's Kimi K3, Bloomberg reported on Monday. UK public-sector buyers will still want to know where a model came from before it is used on their data.
Gemini 3.8 voices and Live Avatar
Google released Gemini 3.8 Flash TTS and Flash-Lite TTS on 23 September. TTS stands for text-to-speech: the models turn written text into spoken audio. They cover more than 100 languages and dialects with over 2,000 production voices, including Scots English, Quebec French and Mexican Spanish. New voices can be designed from a written description, and copying a real person's voice needs a 30-second sample and a consent check.
All the audio carries Google's SynthID watermark and C2PA content credentials, which let other software detect that it was made by AI. The models are in the Gemini API and AI Studio now, Flash TTS is in Gemini Notebook and Flash-Lite TTS is in Google Vids. Access through the Gemini Enterprise API is listed as coming soon, and Google's announcement did not include prices.
On 24 September Google made Gemini 3.8 Live with Live Avatar generally available in Gemini Enterprise, with US and EU endpoints and provisioned throughput, meaning capacity reserved for the customer. A Live Avatar is an on-screen video character with lip-synced speech. It can hold a spoken conversation, cope with being interrupted, call other tools in the background and understand a live camera feed or a shared screen, in 97 languages. Audio and video both carry SynthID, and custom avatars are available only through an enterprise allow-list and verification process.
Google names Cox Automotive, owner of Autotrader, and Equal AI, which handles more than a million calls a day in nine Indian languages, as customers. An EU endpoint means processing can stay in the EU, which matters for European data residency. The announcement does not include a UK endpoint.
Oracle Health and Life Sciences Summit
Oracle's Health and Life Sciences Summit ran in Orlando from 22 to 24 September, and its announcements on 23 September were all of planned capabilities. The first is five AI features for the revenue cycle, the process by which a hospital is paid for the care it provides: prior authorisation, clinical document quality integrity, charge capture and integrity, professional-fee medical coding, and appeal management. Oracle says they are planned for general availability in the coming months, and they are aimed at the US.
Oracle intends to connect those reimbursement workflows to Oracle Fusion Cloud Applications, its ERP suite, for reconciliation, revenue accounting, treasury and analytics. Seema Verma, who leads Oracle Health and Life Sciences, said: "AI gives us an opportunity to prevent revenue cycle problems before they lead to denials and delayed payments."
Oracle also announced an Oncology EHR, an electronic health record for cancer care. Planned features include a single timeline of each patient's cancer care, chemotherapy and immunotherapy ordering with cumulative-dose calculations, and an AI assistant for pre-visit summaries and tumour-board preparation. Oracle has named no customers and no launch date, and HealthSystemCIO notes that the release carries a future-product disclaimer.
Before switching models
Nearly all the benchmark figures above come from the companies that make the models, and the one independent comparison this week, for Grok 4.7 on Terminal-Bench, came in well below xAI's own number. The useful test for a business is its own: run a sample of real tasks through the current model and the new one, and compare the output and the bill.
The cost of a workload depends on how many tokens it uses as well as on the price per token. Anthropic's claim for Opus 5.5 rests on it using fewer tokens for the same work, and Harvey's model bills rose after an update to its agents pushed token use up roughly twentyfold. Cost per completed task, which OpenAI quotes for Sol on AutomationBench, is the more useful figure to track.
OpenAI says the new Sol and Luna prices are permanent, and Oracle's revenue cycle features are planned for general availability in the coming months.
SAASiQ - Intelligent Solutions for SaaS ©


