This Week's Releases: Gemini 3.7 Flash, MAI-Thinking-1, Agent Plugins and Pixel 11

Updated: 6 days ago
Title: This Week's Releases: Gemini 3.7 Flash, MAI-Thinking-1, Agent Plugins and Pixel 11
Date: 13 August 2026
Type: Blog
Author: SAASiQ (contact@saasiq.ai)
Word count: 1159 words
Reading time: 5 min
Published: 13-08-2026
Five releases this week matter to business users. Google shipped Gemini 3.7 Flash on 13 August at an introductory price half that of the previous version, Microsoft put its own reasoning model into public preview, OpenAI previewed a much faster way to run its flagship model, and a shared format for packaging AI agent tools went live across Microsoft's developer products. Google also launched the Pixel 11 phones, with Gemini now able to act inside other apps.
Gemini 3.7 Flash
Flash is Google's lower-cost tier, the models it positions for high-volume work such as coding assistants, customer service agents and document processing. Version 3.7 arrived three weeks after 3.6. Google describes it as its most capable workhorse model for coding and agents, and the benchmarks it published show the gains concentrated there: 65.3 per cent on the DeepSWE software engineering test against 49.0 per cent for 3.6 Flash, and 30.4 per cent on AutomationBench, a test of multi-step business tasks, against 17.0 per cent.
The price is the headline for anyone running it at volume. Until 31 December 2026 it costs $0.75 per million input tokens and $3.75 per million output tokens. From 1 January 2027 that doubles to $1.50 and $7.50. On the usual rule of thumb of three-quarters of a word per token, a million tokens is roughly 750,000 words.
It is available now through Google AI Studio, the Gemini API, Android Studio and the Gemini Enterprise Agent Platform, and it powers Gemini Spark for Pro and Ultra subscribers in more than 160 countries. Google says it has updated the model's safeguards for chemical, biological and cyber misuse, and Box, one of the early testers, reported significantly better results than 3.6 Flash at lower cost.
Microsoft's MAI-Thinking-1
MAI-Thinking-1 comes from Microsoft AI, and Microsoft says it was trained entirely in-house rather than distilled from another company's model. Microsoft has relied largely on OpenAI's models for its Copilot products, so this is a reasoning model it owns outright. It is a mixture-of-experts design with about one trillion parameters in total, of which 35 billion are active for any given request, which Microsoft says gives it a smaller running footprint than larger competing models.
A mixture-of-experts model is built from many specialised sub-networks, and a routing layer sends each request to only a few of them. The model has the knowledge of a very large network but the running cost of a much smaller one, since most of it sits idle on any single request. Many of the large models released in the past two years use some version of the approach.
Microsoft's own figures put it level with Anthropic's Claude Opus 4.6 on SWE-Bench Pro, a software engineering test, and at 97.0 per cent on the AIME 2025 mathematics test and 94.5 per cent on AIME 2026. It also reports that human reviewers preferred its answers to those of Anthropic's Claude Sonnet 4.6 in a blind comparison across 1,276 tasks. It handles 256,000 tokens of context, supports function calling, and works with the standard Chat Completions API, so existing applications can switch to it with little change.
It went into public preview in Microsoft Foundry and the MAI Playground on 12 August, and Microsoft has already added it to GitHub Copilot and VS Code. Pricing was not part of the announcement.
GPT-5.6 Sol Ultrafast
OpenAI previewed a faster service tier for GPT-5.6 Sol, its flagship model, running on hardware from Cerebras. Cerebras says the Ultrafast tier produces up to 750 output tokens per second, up to 14 times the speed of the standard endpoint. On a sample of GDP-Val tasks, which are drawn from real legal, financial and engineering work, it completed the same deliverables 5.6 times faster.
According to Cerebras it is the same model, with the same weights, precision and reasoning settings, served on different chips. It launched on 13 August as a limited preview in the OpenAI API for a small group of customers, with wider access promised as capacity grows. No price has been published.
The gain is largest for agents, which make many model calls one after another, because the waiting time adds up across every call in the chain. For a single question typed into a chat window the difference is less noticeable.
Agent Plugins 1.0
Agent Plugins 1.0 is an open format for packaging the two things AI assistants need to do real work: skills, which are reusable instructions, and MCP servers, which connect an assistant to other software and data through the Model Context Protocol. A plugin built once in this format can be installed in any compatible assistant.
The standard was published on 6 August by AWS, Anysphere, Microsoft, OpenAI and Vercel, with Google joining as a core maintainer. On 12 August GitHub made it generally available on all Copilot plans, in VS Code, the Copilot CLI, the Copilot app, the Copilot SDK and the Copilot cloud agent. Existing Copilot plugins keep working.
For Copilot Business and Enterprise customers, GitHub provides managed settings to install or block plugins automatically, restrict which plugin marketplaces users can reach, and allow or block individual MCP servers by URL, command or name. Those controls are the part for IT and security teams to look at before developers start installing plugins freely.
Pixel 11
Google launched the Pixel 11 range at its Made by Google event in New York on 12 August. In the US the Pixel 11 starts at $899, the Pixel 11 Pro at $1,099 and the Pixel 11 Pro XL at $1,299, a $100 increase on its predecessor, with base storage of 256GB across the range. The Pixel 11 Pro Fold follows in October. Google says the new Tensor G6 chip improves power efficiency by 20 per cent.
Gemini on the Pixel 11 can now carry out tasks inside third-party apps, such as ordering groceries, booking rides and making reservations, and it can call businesses on the user's behalf. Other additions include translation of videos, podcasts and voice messages, dictation built on Google's Rambler product, and sign language translation to text. On the camera side, Magic Capture picks the best frames from a video clip and cleans them up automatically, and Circle to Search now works directly inside the camera app.
Before switching
The benchmark results for all of these are published by the vendors and have not been independently reproduced yet. For a business deciding whether to move a workload, the more useful test is its own: run a sample of real tasks through the current model and the new one, and compare the output and the bill.
Preview status matters for MAI-Thinking-1 and Ultrafast. Azure services in public preview are covered by Microsoft's supplemental preview terms rather than its standard service level agreements, and Ultrafast is limited to selected customers, so both suit evaluation work for now.
Google's introductory pricing for Gemini 3.7 Flash runs until 31 December 2026, and budgets for next year should use the standard rate.
SAASiQ - Intelligent Solutions for SaaS ©


