This Week's Releases: GPT-6 Astra, Claude Fable 5.1, Gemini 3.8 Flash and Muse Spark 1.3

Updated: 6 days ago
Title: This Week's Releases: GPT-6 Astra, Claude Fable 5.1, Gemini 3.8 Flash and Muse Spark 1.3
Date: 3 September 2026
Type: Blog
Author: SAASiQ (contact@saasiq.ai)
Word count: 1580 words
Reading time: 6 min
Published: 03-09-2026
OpenAI released GPT-6 Astra on 3 September, the first model it has deployed broadly while rating it 'Critical' for cybersecurity risk. Earlier in the week Anthropic released Claude Fable 5.1 with a large cut to one part of its pricing, Google shipped Gemini 3.8 Flash and Meta updated its coding model. Perplexity changed how its assistant handles personal data on Macs, and GitHub changed how Copilot seats are billed from 1 October.
GPT-6 Astra
On launch day Astra went to customers of Daybreak, OpenAI's programme for security defenders, with paid ChatGPT plans, the API and Amazon Web Services following in the coming days. Reports differ on whether the ChatGPT Plus plan is included, so anyone relying on it should check OpenAI's help centre. OpenAI's Greg Brockman introduced the model with the words 'Welcome to the AGI era'.
OpenAI grades its models against what it calls its Preparedness Framework, which rates how much help a model could give someone trying to cause serious harm, with cybersecurity as one of the categories. Astra is the first model OpenAI has deployed broadly with a 'Critical' rating for cybersecurity. The public version refuses advanced offensive work, and vetted defenders get less restricted access through Daybreak.
TechCrunch calls the model controversial because of how it reasons. Models of this kind usually work through a problem in written steps, known as a chain of thought, which safety teams can read to check what the model is doing. Astra uses a looped-reasoning technique, also described as opaque recurrence, in which more of that working is harder to follow. OpenAI's Jakub Pachocki said: 'As model capabilities are increasing, monitorability is getting more challenging.' Technode reports that Astra's reasoning was less monitorable in adversarial evaluations.
Prices are quoted per million tokens, a token being a fragment of a word, with separate rates for what goes into the model and what comes out. Astra costs $10 per million input tokens and $50 per million output. Cached input, meaning material the model has already processed in an earlier request, costs $1 per million, and writing to the cache $12.50. Batch and Flex processing are half price, Fast mode is double, and requests with more than 272,000 input tokens are charged at a higher rate. These figures come from secondary sources citing OpenAI's pricing page. Fast mode is not available to customers using EU data residency.
According to CellCog's reading of OpenAI's documentation, Astra can take in about 1.05 million tokens at once, produce up to 128,000 in one answer, and has knowledge up to 30 April 2026. Axios reports it was trained on more than 100,000 GPUs at the Stargate site in Texas. OpenAI's own results include 72.6 per cent on OSWorld 2.0, a test of operating a computer, taking about 47 per cent less time per task than GPT-5.6 Sol. On Terminal-Bench 4.0, a test of work at the command line, it scores 57.9 per cent against 55.8 per cent for Claude Fable 5.1.
Claude Fable 5.1 and Mythos 5.1
Anthropic released two versions of the same model on 1 September. Claude Fable 5.1 is generally available. Claude Mythos 5.1 has different safeguards and is restricted to verification programmes in cyber defence and life sciences and, for now, to US organisations, so UK organisations, public bodies included, cannot get it at present.
The list price is unchanged at $10 per million input tokens and $50 per million output, the same as Astra. The change is to caching. When an application sends the same material to the model again and again, such as a long policy document that several questions refer to, the repeated part can be read from a cache at a lower rate. Anthropic has cut that rate by 75 per cent, from $1.00 to $0.25 per million tokens. It says typical workloads will cost about 25 per cent less, and agentic tasks, where the model works through many steps on its own, up to about 45 per cent less. Anthropic's Opus 5 costs $5 and $25.
Anthropic's benchmark figures, which VentureBeat stresses have not been independently verified, show the largest gain on Terminal-Bench-Science, a set of scientific computing tasks: 52.6 per cent, against 24.7 per cent for Fable 5 and 29.0 per cent for Opus 5. It scores 55.8 per cent on Terminal-Bench 4.0, up from 42.0 per cent, and 31.4 per cent on AutomationBench, a test of multi-step business tasks, up from 17.1 per cent.
It is available in the Claude apps and Claude Code, on the API, and through Amazon Bedrock, Google Cloud and Microsoft Foundry on Azure. Zero data retention is available now to eligible customers. Enterprise Frontier Safeguards, which run on cloud infrastructure the customer controls, roll out this autumn. Anthropic names Jane Street, Millennium and Cognition among its customers, and says Cognition moved its Opus 5 traffic to Fable 5.1 on the first day.
Gemini 3.8 Flash
Gemini 3.8 Flash, released on 2 September, is a new model and the successor to 3.7 Flash in Google's lower-cost range. The introductory price matches its predecessor's: $0.75 per million input tokens and $3.75 per million output until 31 December 2026, doubling to $1.50 and $7.50 from 1 January 2027.
A review by eesel AI says it is built on the same base as 3.7 Flash and uses more thinking tokens, the internal working a model does before it answers. Google advises staying on 3.7 Flash for work where efficiency comes first. Google's own results include 54.9 per cent on HLE-Verified and 47.2 per cent on CWE-Bench, which tests whether a model can patch security flaws, with more than 70 per cent success at finding real vulnerabilities across 20 programming languages. Wiz, a customer, reports 7.5 to 9.7 per cent higher recall on penetration-testing benchmarks, meaning it found a larger share of the vulnerabilities, at 2.3 to 5.2 times lower cost.
It is available in the Gemini app for AI Pro and Ultra subscribers, AI Studio, the Gemini API, Vertex AI, Gemini Enterprise, AI Mode in Search, Google Sheets, Android Studio, Stitch and Antigravity, and GitHub added it to Copilot on 3 September. A separate version, Gemini 3.8 Flash Cyber, is available only through Google's Fairwind Program for governments, critical infrastructure operators and software maintainers, with no published price.
Meta Muse Spark 1.3
Muse Spark 1.3 is Meta's model for agentic coding, software work the model carries out over many steps using tools, and it is rolling out in Muse Code and the Meta Model API. Meta says it makes about 20 per cent fewer tool calls and uses about 25 per cent fewer tokens than version 1.2. Mark Zuckerberg called it Meta's 'biggest jump' in coding and agentic capability, according to VentureBeat.
Prices are unchanged at $1.25 per million input tokens and $4.25 output, according to VentureBeat and DigitalApplied. A Contributor tier costs $0.10 and $0.20 in return for letting Meta train on the customer's data, and any organisation looking at it should read those terms with its data protection officer first.
The setting most developers can use, called xhigh, scores 61 on the Artificial Analysis Intelligence Index, level with GPT-5.6 Sol at its maximum setting and Grok 4.6 high. A stronger max mode, which scores 66.9 per cent on OSWorld 2.0, is still in safety testing. Bloomberg reports that Meta's Alexandr Wang called the model 'competitive' with Claude Fable 5.1 and 'better than' GPT-5.6 Sol at software development. Availability varies by region, and the regions have not been listed.
Perplexity Hybrid Compute on Mac
Perplexity's release on 1 September changes where personal data is processed. Tasks in Perplexity Computer start in the cloud. A small classifier on the user's Mac checks each step for sensitive data and sends those steps to a model running on the Mac, so that data stays on the machine. Tasks started on an iPhone or iPad can use a connected Mac.
It is included in the Pro, Max and Enterprise plans at no extra cost. It needs an Apple silicon Mac on macOS 15 or later with at least 24GB of unified memory, and Perplexity recommends 32GB. The launch models are Gemma 4 E4B, Qwen3.6 35B-A3B and a model Perplexity has trained further itself, and the one-click recommended setup is PPLX Qwen 3.8 27B.
The classifier, PII-Tracer, is published as open source on Hugging Face. It is a 0.6 billion-parameter model built on Qwen3 and detects nine types of personal data. Perplexity also released a benchmark, PII-TRACE, of 13,148 synthetic conversations in 13 languages. In SAASiQ's view the design suits work covered by UK GDPR, since personal data stays on the device and only the rest of the task goes to the cloud.
GitHub Copilot billing
GitHub is gradually reopening sign-ups for Copilot Business and Enterprise to customers who pay by card or PayPal, after strengthening its account vetting and billing. The changelog does not say when or why sign-ups were paused. From 1 October 2026 new seats have to be paid for upfront before access is granted, and all seats are charged upfront at the next billing cycle. Included usage may be prorated, overage fees apply above the allowances, and list prices are unchanged. Customers who cancel and come back may go through the new vetting.
Two smaller changes arrived on 2 September. Content exclusions, which keep chosen files out of Copilot's reach, are now generally available in the Copilot app and the command-line tool. Administrators can also now set any model as the default through enterprise-managed settings.
SAASiQ - Intelligent Solutions for SaaS ©


