Dev Radar: September's frontier models - what changed for teams that build

Published date:

Share directly to:

Dev Radar: September's frontier models - what changed for teams that build - Keter AI
Dev Radar: September's frontier models - what changed for teams that build - Keter AI
Dev Radar: September's frontier models - what changed for teams that build - Keter AI
Dev Radar: September's frontier models - what changed for teams that build - Keter AI
Dev Radar: September's frontier models - what changed for teams that build - Keter AI

Published date:

Share directly to:

Dev Radar: September's frontier models - what changed for teams that build - Keter AI
Dev Radar: September's frontier models - what changed for teams that build - Keter AI
Dev Radar: September's frontier models - what changed for teams that build - Keter AI
Dev Radar: September's frontier models - what changed for teams that build - Keter AI
Dev Radar: September's frontier models - what changed for teams that build - Keter AI

In September Anthropic and OpenAI shipped new top-tier models and Google announced one. For teams that run production workloads, the story is not the benchmark tables. Prices moved, API contracts tightened and several older models now have retirement dates.

What shipped

Anthropic. Claude Fable 5.1 arrived on 1 September, alongside Claude Mythos 5.1 for Project Glasswing participants. Both offer a 1M token context window, 128k max output tokens and always-on adaptive thinking at $10 input / $50 output per million tokens. Claude Opus 5.5 followed on 22 September at $4 / $20 (Opus 5 was $5 / $25), and Claude Sonnet 5.5 on 28 September at $2 / $10, both with the same context and output limits.

OpenAI. GPT-6 Astra reached the API on 3 September with a 1,050,000 token context window, 128,000 max output tokens and pricing of $10 input / $50 output per million tokens. GPT-6 Sol ($2 / $10) and GPT-6 Luna ($0.10 / $0.50) followed on 22 September, and GPT-6.1 Sol on 29 September, also at $2 / $10; the Sol and Luna prices apply to prompts up to 272K input tokens. The input price spread between Astra and Luna is 100x.

Google. Gemini 4 Argon was announced on 30 September, but access starts with trusted cyber defenders in Google's Fairwind Program, with no date given for developers and enterprises. Google lists a 1 million token output limit and introductory pricing of $2 / $10 per million tokens, then $4 / $20.

The benchmark scores published with Opus 5.5 and Gemini 4 Argon are vendor-reported and not independently reproduced. Anthropic reports 66.4% on Terminal-Bench 4.0 for Opus 5.5 at xhigh effort; Google reports 77.9% on DeepSWE v1.1 for Argon, which has no public API yet.

Breaking changes and API constraints

  • Opus 5.5: thinking cannot be disabled (thinking type "disabled" or "enabled" returns a 400 error), and forced tool use through tool_choice "any" or "tool" also returns 400.

  • Sonnet 5.5: Anthropic documents five breaking changes when moving from Sonnet 5.

  • Fable 5.1 and Mythos 5.1: forced tool use is not supported, thinking blocks are bound to the producing model or a newer one, generated text carries Anthropic's text watermark, and 30-day data retention is required. Zero data retention is unavailable unless Anthropic expressly authorises it.

  • GPT-6 Astra: no "none" reasoning effort, no custom temperature or top_p values, no logprobs, and tool calling requires the Responses API.

Deprecations and deadlines

  • Claude Sonnet 4.5: deprecation announced on 30 September, retirement on the Claude API on 30 November 2026.

  • OpenAI: gpt-5.3-codex, gpt-5.1 and gpt-5.4-nano were deprecated on 1 October and leave the API on 1 April 2027.

Why it matters

Two points for regulated buyers. First, Anthropic's top tier now carries compliance conditions (retention, watermarking, safety classifiers that can refuse requests) that the compliance lead should review, not only engineering. Second, availability is not guaranteed: access to Fable 5 was suspended in June after US export controls were applied, and restored on 1 July. Critical workloads need a tested fallback to a second model or vendor.

Pace is a governance issue too. On 25 September OpenAI fixed an image encoding bug in GPT-6 Sol and Luna and recommended rerunning evaluations for image workloads. Evaluation suites have to be cheap to rerun.

Migration checklist

  • Inventory every pinned model ID; find Sonnet 4.5, gpt-5.1, gpt-5.3-codex and gpt-5.4-nano.

  • Search the code for forced tool_choice, disabled thinking, temperature, top_p and logprobs.

  • Rerun your own evaluation set before changing any default model.

  • Re-cost workloads at the new prices.

  • Confirm retention and residency terms for each model tier.

  • Define and test a fallback model for each critical workflow.

Sources

Newsletter

Dev Radar, reviews and regulation notes for enterprise AI teams. No hype, only checked facts.

Newsletter

Dev Radar, reviews and regulation notes for enterprise AI teams. No hype, only checked facts.

Newsletter

Dev Radar, reviews and regulation notes for enterprise AI teams. No hype, only checked facts.

Book a readiness call.

Bring one process, product or function where AI should help. We will suggest the most practical next step.

Book a readiness call.

Bring one process, product or function where AI should help. We will suggest the most practical next step.