This Week's Releases: Apple's M6 Macs, Claude's Shared Memory and Perplexity's Local Agent

Updated: 6 days ago
Title: This Week's Releases: Apple's M6 Macs, Claude's Shared Memory and Perplexity's Local Agent
Date: 27 August 2026
Type: Blog
Author: SAASiQ (contact@saasiq.ai)
Word count: 1469 words
Reading time: 6 min
Published: 27-08-2026
Apple announced its M6 and M5 Ultra chips on 25 August, in a new Mac mini and Mac Studio. On the same day Anthropic gave Claude one memory across its chat and Cowork apps, and Perplexity released an AI agent that runs entirely on a desktop machine. Nvidia's Groq 3 LPX hardware went into full production, OpenAI cut GPT-5.6 Sol prices until November, and two Chinese labs released low-cost models that can read images.
Apple's M6 and M5 Ultra
The M6 is Apple's first chip made on a 2-nanometre process, and it goes into the new Mac mini. It has a 12-core CPU, a 12-core GPU with what Apple calls Neural Accelerators, and up to 32GB of memory. Apple claims up to 1.2 times the multithreaded performance of the M5 and "nearly 30 percent" more peak GPU compute for AI.
The M5 Ultra goes into the Mac Studio. It joins two M5 Max chips together, which makes it the first M-series chip built from four dies, the separate pieces of silicon a chip is assembled from. It has up to 36 CPU cores, up to 80 GPU cores and up to 512GB of unified memory, with 1.2TB a second of bandwidth. Apple claims 4.5 times the GPU AI compute of the M3 Ultra.
Unified memory is shared by the processor and the graphics cores, and AI models need a great deal of it. With 512GB, a single desktop can hold very large open models, the kind whose files are published for anyone to download, and run them locally, so the data never goes to a cloud service.
In the US the Mac mini starts at $899 with the M6 (16GB of memory, 256GB of storage), $100 more than before, and costs $1,699 with the M5 Pro. The Mac Studio costs $2,499 with the M5 Max and starts at $5,499 with the M5 Ultra (96GB, 1TB), up $200. Apple's press releases did not give UK prices. Pre-orders opened on 25 August, and the machines ship from 22 September with macOS 27.
Claude's shared memory
From 25 August, Claude's chat and Cowork share a single memory, so Cowork now remembers what a user told Claude in chat. The memory is updated as a conversation goes along, where it used to be summarised at the end. It is on by default for Free, Pro and Max users on the web, desktop and mobile apps, and mobile users need to update the app to get it. Users can see, edit and delete what Claude has stored.
Some subjects are left out by default: health, racial or ethnic origin, religious beliefs, political views and gender identity. A setting called "include sensitive topics in memory" lets users opt in, and Claude tells them when it saves something in those categories. According to TechCrunch, government ID numbers, Social Security numbers, criminal history and immigration status are never stored.
The reports from TechCrunch and 9to5Mac cover only the Free, Pro and Max plans. Organisations on Team or Enterprise plans should check Anthropic's help centre for how memory is set on their accounts.
Claude Security on Mythos 5
Since 21 August, Claude Security, Anthropic's code-scanning service, has run on Claude Mythos 5, a model that until then was available only to vetted security defenders through Anthropic's Project Glasswing. It is for Claude Enterprise customers only. An administrator switches it on at claude.ai/security, and it is billed as ordinary token usage with no add-on fee.
It scans connected GitHub repositories and reports each finding with its CWE category (the Common Weakness Enumeration, a standard list of types of software flaw), a severity and confidence rating, and a suggested fix. A person has to approve each fix. In restricted-output mode users see the scan results and no prompt box, so they cannot ask the model to write exploit code. Anthropic also announced a $35 million Defender Advantage Fund for open-source security.
Perplexity's Portable Computer
Perplexity's Portable Computer, released on 25 August, is an AI agent whose parts all run on the user's own hardware: the orchestrator that plans each step, the models, the harness and a sandbox for running code. The hardware is Nvidia's DGX Spark, a desktop machine with a GB10 chip and 128GB of memory that Nvidia lists at $4,699. A PC with an RTX graphics card carrying at least 24GB of video memory also works.
It runs Qwen 3.8 27B or PPLX 27B, a version Perplexity has trained further itself. The 27B means 27 billion parameters, the adjustable values a model learns in training. Nemotron 3.5 Lightning is listed as coming soon, and users can bring their own model.
Steps that run locally carry no per-token charge. A task can be passed up to a cloud model when needed, at about $0.415 a task against about $0.65 for Claude Opus 5 alone, according to MarkTechPost. Before anything goes to the cloud, a classifier checks it for personal data and the user approves the step. Perplexity says the remote adviser "never receives direct access to local files, tools or the conversation." Code runs in a sandbox enforced by the operating system, which limits the processes, file paths and network access it can use.
It launched on Linux, with Windows due in September and no Mac version planned. Reports differ on which subscriptions include it: MarkTechPost says Pro and Enterprise, other coverage says Pro and Max, and several outlets describe it as a free add-on. For public-sector and regulated buyers it is a working example of an agent that keeps data on a local machine and asks a person before anything leaves it.
Nvidia Groq 3 LPX
Nvidia announced on 24 August that Groq 3 LPX is in full production. It is a dedicated accelerator for inference, the stage where a trained model produces its answers, and it extends Nvidia's Vera Rubin NVL72 rack systems. It comes out of Nvidia's licensing deal with Groq, the inference chip company.
Nvidia quotes about 3,400 output tokens a second on the Gemma 4 31B model with 100,000 tokens of context. A token is a fragment of text, on the usual rule of thumb about three-quarters of a word. AI Weekly reports 3,431 tokens a second in testing by Artificial Analysis, nearly four times the fastest public endpoint. Speed of output matters most for agents, which make many model calls one after another and wait on each.
Nebius is the first cloud provider to adopt it, through its Nebius Token Factory service, and Groq also plans early adoption. Nebius has not given a launch date, regions or prices.
GPT-5.6 Sol prices cut until November
OpenAI cut prices for GPT-5.6 Sol, its flagship model, on 21 August, for three months to 21 November. Per million tokens, input falls from $5.00 to $4.00, cached input from $0.50 to $0.40 and output from $30.00 to $20.00. Cached input is text the model has already processed recently, such as a long set of standing instructions sent with every request, and it is charged at the lower rate.
The reduced prices apply to the API, Codex credits and ChatGPT Work, including Fast mode, long-context, Batch and Flex. Plus, Pro and Business subscription prices have not changed.
Low-cost image models from DeepSeek and Z.ai
DeepSeek released V4-Flash-Vision-Exp on 21 August, an experimental model that takes images and text and replies in text, with the same text ability as V4-Flash. It is priced at V4-Flash rates of $0.14 per million uncached input tokens and $0.28 per million output tokens, and each image counts as no more than 384 tokens.
Z.ai published the weights of GLM-5.3-Flash on 26 August, according to Codersera and AI Weekly, under the MIT licence, so anyone can download and run it. It has 320 billion parameters in total, of which 18 billion are active on any request, which keeps running costs down. It is the first model in the GLM-5 series built to handle images and text from the start, and it takes up to 300,000 tokens of context. Its model card reports 84.3 on Terminal Bench 2.1 and 63.4 on Deep SWE, both tests of coding and agent work.
Z.ai says it performs close to Claude Opus 4.8 on coding and agent tasks at about one-tenth the price of GLM-5.2. Bloomberg, cited by AI Weekly, confirmed that an anonymous model called "Ox Alpha" on OpenRouter was from the GLM series.
Before switching
The performance figures here come from the vendors, apart from the Artificial Analysis speed test, and none has been reproduced by SAASiQ. A business thinking of moving a workload is better served by running a sample of its own tasks through the current model and the new one and comparing the output and the bill.
Budgets for next year should use OpenAI's standard Sol rates. The reduced prices run until 21 November 2026.
SAASiQ - Intelligent Solutions for SaaS ©


