OpenAI Unveils Its First Custom Chip as Anthropic Leases Google's: How Much of the AI Stack to Own

Updated: 6 days ago
Title: OpenAI Unveils Its First Custom Chip as Anthropic Leases Google's: How Much of the AI Stack to Own
Date: 30 June 2026
Type: Blog
Author: SAASiQ (contact@saasiq.ai)
Word count: 1340 words
Reading time: 6 min
Published: 30-06-2026
OpenAI and Broadcom unveiled Jalapeño on 24 June, the first chip OpenAI has designed itself, built to run its models once they are trained. Earlier in the month Anthropic went the other way, with Apollo and Blackstone completing a $35 billion deal to buy Google-designed chips that Anthropic will lease. Two of the largest AI labs are sourcing compute in opposite ways, and organisations buying AI make a smaller version of the same decision about how much of their own setup to own.
What OpenAI announced on 24 June
Jalapeño is what OpenAI calls an Intelligence Processor, an accelerator designed around large language model inference, the stage where a trained model answers requests from users and applications. OpenAI designed it and Broadcom is building it. The two companies describe it as the first chip in a platform that will run over several generations, with initial deployment planned by the end of 2026.
Greg Brockman, OpenAI's president, told CNBC the chip was designed end to end in nine months, with help from OpenAI's own models. Hock Tan, Broadcom's chief executive, said early testing showed cost savings of roughly 50 per cent against typical AI graphics processors, and OpenAI said early results showed significantly better performance per watt than current alternatives. Both are the companies' own figures from early testing.
TechCrunch reported that the chip is meant to reduce OpenAI's reliance on Nvidia, and that the heaviest work, such as pre-training new models, will probably still run on Nvidia hardware. So OpenAI is designing the chips that run its models for users, and buying in the chips it trains them on.
The announcement follows the agreement OpenAI and Broadcom signed on 13 October 2025 to deploy 10 gigawatts of OpenAI-designed accelerators, networked with Broadcom's Ethernet equipment. Under that agreement the racks start going in during the second half of 2026 and the rollout finishes by the end of 2029.
Anthropic leases its chips instead
Anthropic does not design its own chips. In early June Apollo and Blackstone completed a $35 billion private credit deal, one of the largest ever, to fund Tensor Processing Units, the AI chips Google designs and Broadcom manufactures. Bloomberg reported that the money pays for chips which Anthropic then leases. The deal initially adds 1 gigawatt of capacity, deployed at data centres run by Fluidstack from the middle of this year, and the companies said they hope the arrangement can support more than 20 gigawatts through 2028.
It sits on top of an agreement Anthropic announced on 6 April with Google and Broadcom for several gigawatts of TPU capacity coming online from 2027, mostly in the United States. In the same announcement Anthropic said it trains Claude on Amazon's Trainium chips, Google's TPUs and Nvidia's GPUs, that Amazon remains its primary cloud and training partner, and that Claude is sold through AWS, Google Cloud and Microsoft Azure. Its chief financial officer, Krishna Rao, called the deal a continuation of a 'disciplined approach to scaling infrastructure'.
Broadcom is in both arrangements. It builds OpenAI's chip, and it manufactures the Google chips that Anthropic's lenders are paying for.
Where dependence showed in June
On 12 June Anthropic switched off Claude Fable 5 for every customer after the US Department of Commerce placed it under export controls. The directive suspended all access by foreign nationals, inside or outside the US, and Anthropic had no reliable way to check a user's nationality in real time, so it disabled the model for everyone. The control applied to the model itself, so the channel a customer bought it through made no difference. Anthropic's other models, including Claude Opus 4.8, were not affected.
On 26 June OpenAI announced its GPT-5.6 models but, at the US government's request, opened them only to a small group of trusted partners whose names it had shared with the government.
Neither event had anything to do with who designs the chips. In both cases a US government decision set when customers could use a new model. For an organisation using AI inside its finance or HR systems, the practical question is whether the model behind each feature can be swapped by changing configuration, and whether the replacement has been tested on the organisation's own work.
Open-weight models and where they come from
The alternative to calling a vendor's model through an API is running an open-weight model, one whose weights are published, on infrastructure the organisation controls. For data that cannot leave a particular environment, that is sometimes the only option, so the supply of good open-weight models matters.
Nvidia released Nemotron 3 Ultra on 4 June, a model of about 550 billion parameters with 55 billion active for any one request. Artificial Analysis called it the leading US open-weight model by far, scoring it 47.7 on its Intelligence Index against 39.2 for Google's Gemma 4 31B and 33.3 for OpenAI's gpt-oss-120b. The same index put Moonshot AI's Kimi K2.6, from China, at 53.9.
OpenRouter's review of the open-weight models that mattered in June, published on 27 June, picked four. Three came from Chinese labs (DeepSeek V4 Flash, Z.ai's GLM 5.2 and MiniMax M3) and one from the US, Nemotron 3 Ultra, which it described as the strongest US open-weight entrant. An organisation whose procurement rules exclude models of Chinese origin is left with a shorter list, and on these benchmarks a weaker one.
What Oracle customers can already do
Oracle's generative AI service on its cloud, OCI Generative AI, has allowed customers to import their own models from Hugging Face or from OCI Object Storage since 12 November 2025. Its June release notes added Nemotron 3 Ultra to the list of models that can be imported on 5 June, then Alibaba's Qwen3 Next 80B, Google's MedGemma 27B and DeepSeek V4 Flash and V4 Pro on 15 June. A note dated 30 June adds private endpoints for imported models, so traffic to them can stay on a private network, along with MiniMax M3.
For UK public bodies, the service has run in Oracle's UK Gov South region in London since 21 August 2025, through the API only, with private endpoints available there since 26 February 2026. It has been in the EU Sovereign Central region in Frankfurt since 28 August 2025, and in the commercial UK South region since June 2024.
On the applications side, AI Agent Studio for Fusion Applications has supported models from OpenAI, Anthropic, Cohere, Google, Meta and xAI since Oracle AI World in October 2025, so an agent built on Fusion data can be pointed at a different provider's model.
What buyers can take from it
The labs make their compute decision with tens of billions of dollars and multi-year contracts. An enterprise makes it on a smaller scale every time it wires an AI feature into a ledger, a payroll or a case management system, and the costs of getting it wrong are different in each direction. Custom pipelines built tightly around one model are expensive to change when that model is withdrawn or repriced. Relying on one vendor's hosted model hands that vendor control over availability and terms.
In practice the first step is a list of which model sits behind each AI feature in use, whether it can be changed by configuration, and what the fallback would be. Where data has to stay within a given boundary, the list should also show whether an open-weight model of adequate quality exists and which region it can be hosted in. Identity, access, data classification and audit around the model stay in place whichever model is plugged in, so they are worth settling first.
Our view at SAASiQ is that for most organisations the model is the part to keep replaceable, and the engineering effort belongs in those controls around it.
Jalapeño's first deployment is due by the end of 2026, the first racks under OpenAI's 10 gigawatt agreement with Broadcom from the second half of this year, and the TPUs financed for Anthropic are being installed at Fluidstack sites from the middle of this year.
SAASiQ - Intelligent Solutions for SaaS ©

