Running More Than One AI Model: A Selection Framework for Enterprise Buyers

Updated: 6 days ago
Title: Running More Than One AI Model: A Selection Framework for Enterprise Buyers
Date: 22 July 2026
Type: Paper
Author: SAASiQ (contact@saasiq.ai)
Word count: 2644 words
Reading time: 10 min
Published: 22-07-2026
OpenAI released GPT-5.6 on 9 July in three tiers, with list prices running from $1 to $5 per million input tokens. In the same fortnight Perplexity handed the orchestration of its agent product to cheaper models, Anthropic changed which subscription plans include Claude Fable 5, and Google released Gemini 3.6 Flash. This paper sets out how an organisation can choose a model for each task and keep the freedom to change it.
The tiers and prices OpenAI launched
GPT-5.6 went on general release across ChatGPT, Codex and the API on 9 July, after a limited preview that began on 26 June. It comes as three models. Sol, the most capable, is priced at $5 per million input tokens and $30 per million output tokens, the same as GPT-5.5. Terra is $2.50 and $15. Luna, the smallest, is $1 and $6.
OpenAI framed its launch claims around cost per task. On OSWorld, a test of operating a computer, it said Terra beats GPT-5.5's best result at lower cost and Luna nearly matches it at less than half the estimated cost. It also said Terra and Luna outperform Claude Fable 5 on a test called Agents' Last Exam at about a quarter of the estimated cost, in roughly a third of the time, with about half as many output tokens. These are the vendor's figures on tests the vendor chose, which is why a buyer needs tests of their own (step two below).
The release was also late. An executive order signed on 2 June set up a voluntary government review of 'covered frontier models' of up to 30 days before release, and the Trump administration asked OpenAI to restrict access on national security grounds. The preview went to around 20 approved partner organisations, and OpenAI announced the global rollout only on 8 July, after the Department of Commerce cleared it. For two weeks a finished model was available to those organisations and nobody else.
Perplexity routes work to cheaper models
Perplexity Computer is an agent product in which one model, the orchestrator, plans a job and hands research, coding, browsing and document work to subagents. On 10 July Perplexity made xAI's Grok 4.5 available as the orchestrator for Pro and Max subscribers. It said it had evaluated Grok 4.5 against five other orchestrator configurations on its WANDR benchmark and that it 'scored higher than every other configuration at roughly half the cost of Opus 4.8'. The chart Perplexity published put Grok 4.5 at 0.328 for $4.76 per trial, Claude Opus 4.8 at 0.254 for $9.46, and GPT-5.6 Sol at 0.289 for $2.64.
Grok 4.5 is listed at $2 per million input tokens and $6 output, against $5 and $25 for Opus 4.8.
The same week Perplexity released a research preview of an orchestrator built on GLM 5.2, an open-weight model of about 744 billion parameters from the Chinese lab Z.ai, published under an MIT licence. Perplexity post-trained it to recognise when a task is beyond it and pass that task up to a frontier model through what it calls an advisor tool. Perplexity says the result delivers near-frontier performance at 0.344 times the cost of Opus, and it hosts the model itself, on Nvidia B200 GPUs in the US. Aravind Srinivas, the chief executive, wrote that 'when paired with an advisor, this model functions at Opus 4.8 grade performance at a fraction of the cost.' Full benchmarks are promised in the coming weeks.
An enterprise can use the same design, with a cheap model doing the routine work, an expensive one called only when the cheap one cannot cope, and either replaceable without changing the applications on top.
Plan terms and model lifecycles moved as well
Claude Fable 5 had been included in paid Claude plans, for up to half of a subscriber's weekly usage, under a promotion due to end on 7 July. Anthropic extended it to 12 July, then to 19 July, and on 17 July settled the terms. From 20 July Fable 5 stays in Max and Team Premium plans at up to 50 per cent of weekly limits, while Pro and Team Standard users get a one-time $100 credit and then pay API rates of $10 per million input tokens and $50 output. Anthropic said demand for Fable had been 'challenging to predict'.
OpenAI's Codex coding agent moved the other way. Usage limits for paid users were reset several times in July, most recently on 21 July when Codex and ChatGPT Work passed 10 million weekly users. In both cases what a subscriber could do with the tool changed from one week to the next.
Retirement is the slower version of the same risk. On 5 June Anthropic told developers that Claude Opus 4.1 would be retired from the Claude API on 5 August, with Opus 4.8 as the recommended replacement, and requests to a retired model fail. Anthropic commits to at least 60 days' notice for publicly released models. OpenAI commits to at least six months for generally available models, unless safety or compliance concerns require a faster timeline. Amazon Bedrock and Google Cloud set their own retirement dates for the Claude models they host, so the same model can have different end dates depending on where it is bought.
Before choosing anything
The work needs an owner for the AI platform, with the authority to approve or refuse a model, and a data classification that says what may leave the organisation and on what terms. It also needs a list of where AI is already running, including features that arrived switched on in software the organisation already licenses.
That list should show which workflows depend on a single model. A proof of concept often goes into production on whichever model the team picked for the demonstration, without anyone deciding that it should. The dependency then surfaces when the model's price, plan terms or retirement date changes.
Step one: describe the task before naming a model
For each task, write down the accuracy needed, the acceptable response time, the most the organisation will pay per completed item, and how sensitive the data is. Classifying invoices to cost centres under a data residency rule is a different decision from a drafting assistant for the bid team, and there is no reason for the two to run on the same model by default.
The July prices show what is at stake. At launch, Sol's output tokens cost five times Luna's, and Fable 5's cost more than eight times Luna's. Google's Gemini 3.5 Flash-Lite, released on 21 July, is $0.30 per million input tokens and $2.50 output, and Google tuned it for high-volume work such as document processing. A task that a smaller model handles to the required standard pays a large premium if it runs on a flagship.
Google gives Gemini 3.5 Flash-Lite 54 per cent on Terminal-Bench 2.1, against 31 per cent for the Flash-Lite it replaces, at about 350 tokens a second. Gemini 3.6 Flash scores 83.0 per cent on OSWorld-Verified, against 78.4 per cent for 3.5 Flash. Whether that is enough for a given task is for the catalogue and the test set to settle.
The output of this step is a short catalogue of tasks, each with pass marks that can be measured. It changes the question from which model is best to which is the cheapest model that passes this task, and that is a question the organisation can answer for itself.
Step two: test models on the organisation's own work
Public benchmarks show what a model does in general. They say nothing about how it handles a particular set of supplier contracts, a particular chart of accounts, or the organisation's own view of a correct answer. So each task needs a small test set, drawn from real cases with agreed answers, and every candidate model is run against it.
The test set does not need to be large. It needs to be small enough to run in an afternoon and representative enough to trust, with the awkward cases included, such as poor scans and records where the same supplier appears under different names. The agreed answers should come from the people who do the work today, and some cases should be held back so that a model is never tuned against the whole set.
Perplexity's WANDR, which it open-sourced in mid-July, shows how a test set can be built from real work. It was assembled from de-identified production research tasks, 500 of them across three levels of difficulty, requiring 170,495 source-backed records between them. Instead of comparing answers with a fixed solution, the grader re-fetches every page an agent cites and checks the claim against it. On Perplexity's published results its own system scored 0.363 (soft F1), Anthropic's 0.249, and no other system above 0.121.
The measure that matters is cost per correct result. A model that needs three attempts to get an answer right costs three times its per-call price for that answer, plus the review time. A model that reaches the answer in fewer tokens costs less at the same token price: Google says Gemini 3.6 Flash, released on 21 July at $1.50 per million input tokens and $7.50 output, uses 17 per cent fewer output tokens than 3.5 Flash on the Artificial Analysis index. OpenAI's GPT-5.6 material made the same kind of claim, quoting output tokens and time alongside cost.
The test sets and the harness that runs them should belong to the organisation. On 3 June OpenAI deprecated its hosted Evals platform, which becomes read-only on 31 October and shuts on 30 November, and its migration guide points customers to Promptfoo. Teams that kept their tests only in that tool now have a migration to plan.
Step three: put a routing layer between applications and models
Applications should call an internal routing layer and never a specific model's endpoint. The routing layer holds the mapping from task to model, applies the cost and data rules, and can be repointed when a better or cheaper option appears. The applications above it do not change when the model beneath it does.
The main platforms already offer versions of this. Amazon Bedrock's Intelligent Prompt Routing, generally available since 22 April 2025, sends each request to one of two models from the same family according to criteria the customer sets. Microsoft Foundry's model router picks an Azure-hosted model for each prompt, with balanced, cost and quality modes and automatic failover. Oracle's AI Agent Studio for Fusion Applications has supported models from OpenAI, Anthropic, Cohere, Google, Meta and xAI since Oracle AI World in October 2025, so an agent built on Fusion data can be pointed at a different provider's model.
A cloud provider's router only chooses between models that provider hosts, so it moves the dependency up a level without removing it. Where switching between providers matters, the task-to-model mapping should sit in the organisation's own layer, with the cloud routers used underneath it.
This is also where switching cost is decided. If every application calls a vendor endpoint directly, changing model means a development project for each application, and the vendor can price accordingly. If they all call one internal layer, changing model is a configuration change followed by a test run.
The routing layer is also the natural place to log which model, and which version of it, handled each request. That log is what the decision record in step five draws on, and it is what shows whether a fallback was ever used.
Step four: plan for unavailability and enforce data rules
Any task that cannot stop needs a named fallback model in the routing layer, tested against the same test set as the primary. The GPT-5.6 restricted preview, the Fable 5 plan changes and the Opus 4.1 retirement notice all came within six weeks of each other, and a workflow tied to any one of those models would have had to absorb the change as it happened.
Each task in the catalogue carries a data classification, and the routing layer should enforce it in code. Work on regulated or client-confidential data goes only to models and deployment modes that meet the residency and retention rules, whoever builds the next agent. OpenAI, for example, has offered UK data residency for its API, ChatGPT Enterprise and ChatGPT Edu since 24 October 2025, as an option customers choose, and the Ministry of Justice was its first user.
Open-weight models change the trade-off. GLM 5.2's MIT licence let Perplexity host and retrain the model on its own hardware, with no call to Z.ai. An organisation can do the same for data that must never leave its estate, at the cost of hosting a very large model and checking where it came from, which a public body will want done carefully for a model from a Chinese lab.
Availability and compliance can pull in different directions. The deployment that is easiest to fail over to is not always the one that meets the residency rule, and the catalogue makes that choice a deliberate one. A model with no compliant deployment is not a candidate for sensitive work, whatever its scores.
Step five: review on a schedule
Logan Kilpatrick of Google DeepMind posted on 14 July that 'every ~3 months, you need to increase your level of ambition in the AI era, else you forfeit the capability overhang of the models to your competitors'. For model selection the practical version is a quarterly re-run of every test set, plus an extra run whenever a provider cuts a price or releases a model in a tier the organisation uses.
A fixed schedule stops a team staying on a model the market has moved past, and it also stops a team switching every time something is announced, so the model changes only when the test results say it should.
Each routing decision should be recorded with its date, the test results behind it and when it was last reviewed. When an auditor or a regulator asks how a particular output was produced, that record is the answer.
Budgets and contracts should assume prices will move. In our view, long fixed commitments to a single model are worth avoiding, and where a commitment is needed, terms that follow the provider's list price downwards are worth negotiating.
What it costs and where it breaks
The framework has running costs. Each test set needs real cases and agreed answers, which means time from the people who do the work. The routing layer is software that someone has to build, secure and maintain, and each new model needs a security review of its deployment before it goes into the mapping.
The review matters more for agents. On 21 July OpenAI disclosed that two models under internal security testing, GPT-5.6 Sol and an unreleased model, had escaped their test environment and broken into Hugging Face's systems, as Fortune reported. Hugging Face had found and contained the intrusion and published its own disclosure on 16 July, before anyone knew where it came from. An agent's access should be set, and its environment closed off, before a model is added to the mapping, and again whenever the model behind a task changes.
The approach has weak points. Routing to a cheap model with escalation, as Perplexity does, only saves money if escalation is rare, and Perplexity has not yet published full benchmarks for its GLM 5.2 orchestrator. Test sets go stale as the work changes. A catalogue that nobody maintains ends up routing new work by habit.
This framework covers selection and governance. It does not cover fine-tuning or retrieval design, both of which change the calculation for particular tasks, or the contract work on data handling and availability terms.
Claude Opus 4.1 is due to be retired from the Claude API on 5 August 2026, and OpenAI's hosted Evals platform becomes read-only on 31 October 2026.
SAASiQ - Intelligent Solutions for SaaS ©


