Loading blog posts...
Loading blog posts...
Loading...

Claude Sonnet 5.5 and GPT-6.1 Sol now share the same $2 input and $10 output price. Procurement now turns on tooling fit, open-weight deployment and operational control as near-frontier intelligence gets cheaper.
Quick summary
Anthropic released Claude Sonnet 5.5 on September 28, followed by OpenAI's GPT-6.1 Sol at DevDay on September 29. Both cost the same at standard rates, but their API behavior differs enough to make benchmark-based switching careless.
| Model | Release | Input per 1M tokens | Output per 1M tokens | Operational distinction |
|---|---|---|---|---|
| Claude Sonnet 5.5 | September 28 | $2.00 | $10.00 | Terminal-Bench 4.0 score of 70.6% |
| GPT-6.1 Sol | September 29 | $2.00 | $10.00 | 1.05M context and $0.10 cached input |
| Claude Haiku 5.5 | October 7 | $0.10 | $0.50 | Standard pricing applies to prompts up to 100K |
| Mistral Large 4 | October 6 | $1.36 | $4.18 | API preview, with weights planned later |
Anthropic reports that Sonnet improved from 10.3% to 70.6% on Terminal-Bench 4.0, a benchmark for agents operating through terminal environments. Teams running with thinking disabled must migrate to the new between_tools setting because existing configurations won't behave identically.
GPT-6.1 Sol introduces different constraints:
272,000 input tokens cost twice the normal input rate.reasoning.effort values run from low to max.OpenAI says Sol approaches GPT-6 Astra at one-fifth of Astra's standard token price, according to the DevDay launch coverage. Most engineering teams should compare task completion rate, retry frequency and tool-call reliability alongside the published base token price.
Warning
Don't treat GPT-6.1 Sol's 1.05-million-token context as a flat-price feature. Crossing the 272,000-token threshold changes the price of the whole request, not just the excess tokens.

Anthropic released Claude Haiku 5.5 on October 7 and positioned it below Sonnet and Opus for complex agentic coding. Its target workloads are compaction, summarization and focused subagent tasks, according to the Haiku 5.5 announcement.
The price gap is too large to ignore:
$0.10 per million tokens$0.50 per million tokens100,000 tokensUse Haiku for bounded, high-volume tasks with validation. Escalate failed or uncertain outputs to Sonnet or Opus for ambiguous planning and difficult code changes. Sending every intermediate step to Sonnet 5.5 costs 20 times more for input and output without enough improvement in the final answer to justify it.
Mistral placed Large 4 into public preview on October 6 through the Mistral Studio API. The company describes it as a mixture-of-experts model with roughly one trillion total parameters and 52 billion active for each token, according to the official release.
Inference cost depends on active parameters, not total parameter count. Running only part of the model for each token can deliver large-model capacity without activating the full network on every request.
The API pricing is aggressive, but teams planning private deployment should wait for the actual license, serving requirements and memory footprint before committing architecture or budget:
$1.36 per million tokens$4.18 per million tokens
Reflection announced Beam on October 5, while Aleph Alpha released Kolibri on October 3. Both use sparse mixture-of-experts architectures. Kolibri is available to download now; Beam remains waitlisted.
| Model | Total parameters | Active parameters | License status | Current access |
|---|---|---|---|---|
| Reflection Beam | 501B | 23B | Apache 2.0 promised | Waitlist |
| Aleph Alpha Kolibri | 78B | About 3B | Apache 2.0 available | Downloadable weights |
| Mistral Large 4 | About 1T | 52B | Not announced | Hosted API preview |
Reflection reports an 80.9 score on SWE-Bench Verified for Beam. The model remains text-only and waitlisted until the promised Apache 2.0 release arrives later in October.
Kolibri is the less spectacular but more actionable release. It supports English and German, offers context up to one million tokens and is already published under Apache 2.0, though serving requires Aleph Alpha's inference plugin for vLLM.
Kolibri deployments beyond 262,144 tokens need extra serving flags. Teams should test memory use and latency at their real context lengths instead of treating the maximum specification as the normal operating point.
Note
Open weights don't remove deployment cost. They transfer responsibility for inference capacity, observability, updates, security controls and license review to your team.
Cloudflare released Clef and Clef-flash on October 1 as Apache 2.0 decision models. They return bounded, typed outputs with probabilities, as detailed in the Clef announcement.
The interface suits routing, policy checks, classification and workflow gates:
@cf/cloudflare/clef.This is the most important architectural release in the group. Small decision models are often the better production component when output must drive code. A support workflow can route high-confidence billing requests automatically while sending low-confidence cases to a person, without asking a general chat model to produce JSON and hoping it follows the schema.

272,000 input tokens before adopting Sol for long-context workflows.Equal headline prices expose the architectural differences underneath them, from tool APIs and context surcharges to deployment licenses and typed outputs.
The winning stack won't send everything to one flagship model. It will route narrow work to cheap subagents, reserve expensive reasoning for hard decisions, and use deterministic decision models where software needs a bounded answer. That is the same control problem behind the recent agent control fight, and these releases make it an immediate engineering priority.