This Week's Releases: Mistral Large 4, Decision Models From Cloudflare and Amazon, and Mods for Claude Code

Title: This Week's Releases: Mistral Large 4, Decision Models From Cloudflare and Amazon, and Mods for Claude Code
Date: 7 October 2026
Type: Blog
Author: SAASiQ (contact@saasiq.ai)
Word count: 1566 words
Reading time: 6 min
Published: 07-10-2026
Mistral released Mistral Large 4 on 6 October as a preview through its API. Mistral Large 4 has more than one trillion parameters. Mistral says its weights will be published at the end of October. Cloudflare and Amazon each released a small open model on 1 October that picks between set answers for AI agents. Microsoft released three speech models on 1 October. Anthropic added mods to Claude Code on 1 October. Anthropic made Claude for Government generally available on 30 September. Google said on 30 September that skills will replace Gems in the Gemini app.
Mistral Large 4
Mistral, the Paris AI company, announced Mistral Large 4 on Tuesday 6 October. Mistral calls it Le Chonk. It has 1.05 trillion parameters in total, but it is a mixture-of-experts model, so only 49 billion of them work on any one request. It reads images as well as text, through a separate image encoder of 1.6 billion parameters, and Mistral lists a context window of one million tokens.
The model is available now in Mistral's API as a preview. Mistral says the weights, the files anyone can download and run on their own servers, will follow at the end of October. The licence terms have not been fully set out yet.
Mistral's documentation lists a price of $1.36 per million input tokens and $4.18 per million output tokens. Those figures are struck through on the page and replaced with half: $0.68 and $2.09. Neither the documentation nor the announcement says how long the lower price lasts.
Mistral says Large 4 is the best open-weight model from the US or Europe on aggregated benchmarks. It also claims state-of-the-art results on cyber defence, manufacturing and finance work, and says it beats closed frontier models at visual grounding, which means locating things in an image. These are Mistral's own figures. The Next Web reports that Mistral accepts the model still trails frontier models in some areas, including coding. Mistral trained it from scratch over two months on about 4,000 Nvidia Grace Blackwell GPUs in its own European data centres.
Decision models from Cloudflare and Amazon
A decision model is a small AI model that chooses between answers set in advance. It takes some input, such as a support ticket or a step in an agent's work, plus a few typed questions. It returns a probability for each answer. A program can then route the ticket, let the agent carry on, or pass the case to a person, depending on the numbers. Several have followed since a company called TypeSafe released a decision model called Jev on 15 September.
Cloudflare released two on Thursday 1 October: Clef, with 27 billion parameters, and Clef-flash, with 9 billion. They are the first models the Workers AI team has trained itself. Clef is built on Alibaba's Qwen3.8-27B and Clef-flash on Qwen3.5-9B. Both accept text, JSON, images and video, with a context window of 64,000 tokens. They answer through the same interface as Jev, so code written for Jev can switch to them.
The weights are on Hugging Face under the Apache 2.0 licence, which allows commercial use. On Cloudflare's Workers AI service, Clef costs $0.24 per million input tokens and Clef-flash $0.09. Output is not charged. Cloudflare reports a median response time of about 39 milliseconds for Clef-flash, and says Clef leads the Jev Decision Index.
Amazon released Strands Decider 2B through Strands Labs, an AWS team, on the same day. It has about 2 billion parameters and is built on Qwen3.5-2B. The usual text-writing part of the model is replaced by a scoring layer of just over a million parameters that rates each answer option.
Strands Decider 2B scores about 72 per cent on the public JevBench test, and placed third of 33 models in its size class for accuracy and calibration combined. Reports of its speed differ slightly: a median of 106 milliseconds or about 115 milliseconds per decision on a single Nvidia RTX 3090 graphics card. It is free under Apache 2.0, and Amazon has published the weights, training data and training scripts on GitHub and Hugging Face.
Microsoft's speech models
Microsoft AI released three speech models in Microsoft Foundry on 1 October. MAI-Transcribe-2-Streaming is its first model that turns speech into text as someone talks. It covers 60 languages and detects which one is being spoken without being told. Microsoft says the first words appear just over 100 milliseconds after the audio arrives. It also says that is twice as fast as the nearest competitor. It costs $0.54 per hour of audio through 2026.
MAI-Voice-2.1 goes the other way, from text to speech. It covers 23 languages and 26 locales, and keeps one voice across all of them with a native accent in each. It can copy a voice from a few seconds of reference audio. It costs $22 per million characters. MAI-Voice-2.1-Flash is the faster, cheaper version at $15 per million characters. It produces up to 45 seconds of audio at a time, with 150 milliseconds of end-to-end delay.
Mods for Claude Code
Anthropic added mods to Claude Code, its coding assistant, on 1 October. A mod is a small program written in TypeScript that hooks into events inside Claude Code and changes what the assistant does. A mod can rewrite a prompt before it reaches the model, or block or retry a tool call. It can approve or deny a permission request and remove secrets from a tool's output.
Mods work in the Claude Code command-line tool and in the Code tab of the Claude desktop app. In the VS Code extension, the Agent SDK and cloud sessions, the hooks still run, but nothing a mod adds to the screen appears. Mods come packaged inside plugins and are installed with the /plugin command.
Anthropic's documentation says mods are not sandboxed. A mod has the same access to the computer as Claude Code itself, and Anthropic advises installing them only from trusted sources. The command claude plugin validate lists the events a mod hooks into and the calls it makes before it is installed.
Claude for Government
Anthropic made Claude for Government generally available to eligible US federal and state agencies on Wednesday 30 September. It had been in public beta since July. It runs in an environment authorised at FedRAMP High, the US government's top cloud security level for sensitive but unclassified work. Staff can use it on files held on their own machines, for work such as drafting memos and reviewing requests for proposals.
Agencies pay for usage, not per seat. They draw down a prepaid balance up to a fixed spending ceiling, and administrators can cap each user's spend in dollars and by model. Agencies can buy it from Anthropic directly, through the reseller Carahsoft or through the GSA's OneGov programme. The Pentagon's designation of Anthropic as a supply chain risk, upheld by a federal appeals court on 25 September, still bars the US military from using its models.
Skills replace Gems in Gemini
Google announced skills for the Gemini app and Google Workspace on 30 September. A skill is a saved set of instructions that guides Gemini through a type of task. Skills replace Gems, Google's earlier way of saving custom instructions. Several skills can be used together inside one chat, which Gems did not allow.
Skills began rolling out to Workspace on 5 October and reach the Gemini app from 13 October. Gems move into the app's Settings panel on 17 November. For business and enterprise customers, Gems can no longer be created, edited or used from 1 March 2027. Any Gems left then are turned into draft skills automatically. Education customers lose Gems from 1 June 2027.
ChatGPT shopping and finances
OpenAI began rolling out a Try on button in ChatGPT worldwide on 1 October. It appears on clothing and accessory listings. A user takes or uploads a selfie, and ChatGPT generates a picture of them wearing the item. It runs on ChatGPT Images 2.5, which OpenAI released on 8 September. Users can also save products into folders in their ChatGPT Library.
On 2 October OpenAI opened Finances in ChatGPT to Free and Go users in the US. It lets users connect bank and card accounts through Plaid and ask questions about their spending. It had been available to Plus and Pro users in the US since 25 June.
Before switching
Most of the performance figures above come from the companies that built the models. Mistral's half price has no end date, so a cost estimate for Large 4 is worth checking against the list price as well.
The decision models are small enough to run on a single graphics card, and their answers come as numbers a program can log. In SAASiQ's view, that makes them easier to audit than a large model's free text when an agent has to choose whether to carry on or hand a case to a person. Anyone testing one can set the threshold for passing a case to a person in advance and record every decision against it.
Claude Code mods run with full access to the machine. An organisation that allows Claude Code can decide which plugin sources are trusted before staff start installing mods.
Mistral says the Large 4 weights will be published at the end of October. Skills reach the Gemini app on 13 October, and Gems move into Settings on 17 November.
SAASiQ - Intelligent Solutions for SaaS ©


