top of page

LATEST INSIGHTS

Expert insights across your SaaS environment

Blogs and papers from the SAASiQ team on enterprise SaaS, cloud and AI.

Claude Programs a Robot Dog and Grok Tops a Voice Humanness Ranking

Writer: SAASiQ.ai
SAASiQ.ai
Jun 26
6 min read

Updated: 6 days ago

Title: Claude Programs a Robot Dog and Grok Tops a Voice Humanness Ranking

Date: 26 June 2026

Type: Blog

Author: SAASiQ (contact@saasiq.ai)

Word count: 1511 words

Reading time: 6 min

Published: 26-06-2026


Anthropic reported on 18 June that Claude Opus 4.7, working with no human help, programmed a four-legged robot about 20 times faster than the best human team managed in the same test last year. The same day xAI said its text-to-speech model sounded the most human on a blind voice benchmark, and OpenAI published both a health update to ChatGPT and a research paper on keeping model behaviour steady under pressure.


What the robot dog test measured

Project Fetch is a test run by Anthropic's Frontier Red Team. In the first round, in August 2025, eight Anthropic researchers and engineers with little robotics experience were split into two teams of four and given a day to get a robot dog to fetch a beach ball. One team could use Claude (then Opus 4.1) and the other could not. Anthropic's write-up, published in November 2025, said the team with Claude finished more of the tasks and took about half the time on the ones both teams completed.


The second round, published on 18 June, gave the same tasks to Claude Opus 4.7 on its own. On the four tasks both human teams had finished, which included connecting to the robot's camera and lidar sensor and writing a manual control program, Opus 4.7 took 9 minutes 35 seconds. Team Claude had taken 181 minutes and the team without Claude 361, which makes the model 18.9 times faster than the faster human team and 37.7 times faster than the slower one. It also wrote 1,045 lines of code, against 10,309 from Team Claude.


The robot still did not fetch the ball. Anthropic said the model got the robot into position but struggled with the fine, closed-loop control needed to pick the ball up and bring it back. ForkLog, reporting the result on 19 June, noted Anthropic's point that the speed-up came from general improvements in the model and not from any robotics-specific training.


Anthropic runs the test to measure 'uplift', how much an AI model adds to what people can do, and to keep a baseline for how well models can control hardware on their own. It links that baseline to its monitoring of risks from AI speeding up research and development.


Grok's voice ranking

xAI said on 18 June that Grok TTS, its text-to-speech model, delivers the most human-like speech. At the time the model led the Humanness Index run by Vapi, a company that builds voice agents. The index clones one real speaker's voice onto every model it tests, plays listeners two unlabelled clips of the same line and asks which sounds more human. The votes are turned into a rating against a recording of the real person, who scores 100. Only models that can clone a voice are included, so that every model is compared reading in the same voice.


The Menon Lab, writing up the index on 16 June, recorded Grok TTS in first place with a score of 90 out of 100, from 721 votes across 16 models from seven providers. The streaming version of Grok TTS, which starts speaking sooner, was second on 86. Because the rating is built from votes that keep arriving, the scores and the order move as the count grows, so any ranking is a reading on the day.


xAI released standalone text-to-speech and speech-to-text APIs in April 2026. At launch it said they ran on the same infrastructure as Grok Voice, the assistant in its mobile apps, in Tesla vehicles and in Starlink's customer support.


Voice bots and the disclosure rule

The EU AI Act has a rule for voices like these. Article 50, which applies from 2 August 2026, requires providers to design AI systems that talk directly to people so that those people are told they are dealing with an AI, unless that is obvious from the context.


For an organisation putting a voice agent on a customer line in the EU, that means the caller is told at the start of the call, however human the voice sounds. The duty sits with the provider of the system, which can be the organisation itself if it builds the agent and runs it under its own name. UK organisations are outside the Act's direct reach unless they operate in the EU.


ChatGPT's health update

OpenAI said on 18 June that GPT-5.5 Instant, the default model in ChatGPT, now performs on a par with its Thinking models on health questions. It said more than 230 million people a week bring health and wellness questions to ChatGPT, and that the updated model is better at recognising when urgent care may be needed, asking for relevant context and explaining uncertainty.


According to OpenAI, the share of health responses flagged for factual problems fell by 71 per cent over two months, and a physician panel rated the model's answers above answers written by physicians on a set of real-world questions. The work draws on HealthBench, OpenAI's evaluation built with more than 260 physicians from 60 countries. Dataconomy, reporting the update on 19 June, pointed out that the evaluation results had not been released for outside review, so the figures are OpenAI's own.


OpenAI's research on steady behaviour

Later on 18 June OpenAI published a research paper, 'Reinforcement learning towards broadly and persistently beneficial models'. Its stated aim is for models taking on longer, higher-stakes tasks to carry safe behaviour into areas they were not trained on, and to keep it when someone tries to push them off it.


The researchers used reinforcement learning on realistic scenarios to reward traits including truthfulness, epistemic humility, corrigibility, risk sensitivity and concern for human welfare. They report improvements on 44 of the 53 internal and external benchmarks they used, covering deception, honesty, reward hacking and safety. Training on health conversations alone improved behaviour outside health, and the trained models were harder to steer into deceptive or harmful behaviour with adversarial prompts.


The paper says more work is needed on how long the traits last and on resistance to harmful fine-tuning. It does not say whether the method is in use in any OpenAI product.


Gemini acting across apps

Google launched Gemini 3.5 Flash at its I/O conference on 19 May, describing it as 'frontier intelligence with action', and made it available in the Gemini app, the Gemini API, Gemini Enterprise and AI Mode in Search. On 26 June Google published an example of what that looks like for a consumer: once a user gives Gemini access to Gmail and Calendar, it finds the flight details, draws up a schedule to limit jet lag and adds it to the calendar.


In a company setting, an assistant that reads mail and writes to a calendar is acting on the user's behalf in two systems, and in a corporate tenancy what it can reach is governed by the same access controls as the user's own account. SAASiQ's view is that those access rights are worth reviewing before features like this are switched on, since an assistant can act across everything the user can open.


The compute bills behind it

The money behind these models is now in public filings. SpaceX's S-1 filing with the US Securities and Exchange Commission, reported by TechCrunch on 20 May, showed Anthropic paying SpaceX $1.25 billion a month through May 2029 for all the available compute at Colossus 1, the data centre near Memphis that xAI, now part of SpaceX, built for its own work. The deal covers 300 megawatts, has a lower rate for the first two months while capacity comes up, and can be ended by either side on 90 days' notice.


On 5 June a further SEC filing showed Google agreeing to pay SpaceX $920 million a month from October 2026 to June 2029 for about 110,000 Nvidia GPUs and related hardware. Google called it bridge capacity to meet demand for Gemini Enterprise, which it said had been even higher than expected. Once Google's payments start, the two contracts together come to about $2.2 billion a month.


Customers do not see these contracts directly, but the cost of running the models sits underneath the per-seat and per-token prices they pay. The Anthropic contract carries a lower rate for its first two months, and Google's payments do not begin until October.


A usage-limit bug in Claude Code

Anthropic said on 19 June that about 3 per cent of Claude Code Max and Pro users had hit a bug that showed them an incorrect weekly usage limit and, in some cases, stopped them sending messages. It fixed the bug and reset the 5-hour and weekly limits for everyone affected, then later reset them for all users on every plan.


It came a week after Anthropic had switched off Claude Fable 5 for every customer on 12 June, following a US Commerce Department directive. For teams using these tools every day, both were interruptions they could not fix themselves and had to wait out or work around. Anthropic's other models, including Opus 4.8, were not covered by the directive, and Fable 5 was still unavailable on 26 June.

SAASiQ - Intelligent Solutions for SaaS ©

Optimise your SaaS licences and software subscriptions with SAASiQ

bottom of page