Dev Radar: Coding agents grow up - auto mode, self-hosting, honest benchmarks

Published date:

Share directly to:

Dev Radar: Coding agents grow up - auto mode, self-hosting, honest benchmarks - Keter AI
Dev Radar: Coding agents grow up - auto mode, self-hosting, honest benchmarks - Keter AI
Dev Radar: Coding agents grow up - auto mode, self-hosting, honest benchmarks - Keter AI
Dev Radar: Coding agents grow up - auto mode, self-hosting, honest benchmarks - Keter AI
Dev Radar: Coding agents grow up - auto mode, self-hosting, honest benchmarks - Keter AI

Published date:

Share directly to:

Dev Radar: Coding agents grow up - auto mode, self-hosting, honest benchmarks - Keter AI
Dev Radar: Coding agents grow up - auto mode, self-hosting, honest benchmarks - Keter AI
Dev Radar: Coding agents grow up - auto mode, self-hosting, honest benchmarks - Keter AI
Dev Radar: Coding agents grow up - auto mode, self-hosting, honest benchmarks - Keter AI
Dev Radar: Coding agents grow up - auto mode, self-hosting, honest benchmarks - Keter AI

In September the coding agents from three vendors changed their defaults, their execution model and their billing. A fourth item, a research preprint, questions how these agents are measured at all.

Claude Code: auto mode by default, policy through managed settings

Claude Code 2.1.284, released on 28 September, made Claude Sonnet 5.5 the default Sonnet model and changed interactive terminal and VS Code sessions to start in auto mode when no permission mode is configured, on every plan and provider. The permissions.defaultMode setting still overrides the new default.

Releases in the days around it added the controls administrators need. Version 2.1.283 (25 September) brought the availableModelsMatch: exact and deniedModels managed settings for pinning and blocking model versions. Version 2.1.285 (29 September) added allowedProviders, which limits the API providers a machine may use. Version 2.1.281 (23 September) gave the Claude apps gateway support for Amazon Bedrock assume_role and a guardrail applied to every request.

Cursor: self-hosted machines, Projects and Security Review

On 2 September Cursor added self-hosted machines, so agents can execute tools on personal machines or team pools inside customer infrastructure. Cursor Projects, in beta since 10 September, has a coordinator agent plan work, delegate it to subagents, keep shared context and subscribe to Slack channels, schedules or pull requests. On 23 September Rollouts and Security Review arrived for Teams and Enterprise plans. Security Review reads every pull request and reports exploitable bugs such as injection, authentication and authorisation bypasses, committed secrets, SSRF and vulnerable dependency changes. Rollouts monitors each pull request as it deploys and can open a revert PR for review; Cursor says it does not merge or roll back on its own.

GitHub Copilot: budgets and policy

GitHub announced policy and billing changes for Business and Enterprise customers on 28 August. Separately, the promotional AI Credits granted after the June move to usage-based billing ($30 a month on Business, $70 on Enterprise) covered June, July and August only, so included usage per seat is back at list levels: $19 and $39 of AI Credits. No earlier than 28 September, Copilot Chat on github.com, Copilot Chat in GitHub Mobile and the Copilot cloud agent are to be relaunched as one unified experience under a single policy, enabled by default. Chat data on github.com will then be retained for the life of the account instead of 28 days. By the end of September GitHub had posted no launch notice.

Benchmarks: a preprint urges caution

A paper submitted to arXiv on 8 September and revised on 16 September reports two sources of unreliability in the SWE-Bench Pro benchmark: reward hacking enabled by leakage of gold solutions or hidden evaluation information, and task quality issues such as misleading problem statements and improperly scoped tests. The authors introduce SWE-Bench Pro Verified, which adds anti-hacking safeguards and minimally corrects flawed tasks. On the verified benchmark, some models perform substantially worse than previously evaluated. This is a preprint and not peer reviewed, so read it as a strong caution, not a settled result.

What this means for AI-assisted engineering programmes

  • Set permission mode and model policy explicitly through managed settings. A default that changes with a release is not a policy.

  • Treat persistent, event-driven agents as production services: a named owner, scoped credentials and guardrails against prompt injection through Slack and pull request triggers.

  • Self-hosted machines keep tool execution inside the customer network. Confirm separately where prompts and context are processed.

  • Re-forecast Copilot spend now that the promotional credits have ended, and review the single policy and longer chat retention with your compliance lead.

  • Choose tools and models on evaluations run against your own repositories and tasks. Leaderboards are a rough filter only.

  • Keep human review gates. Security Review and Rollouts report and propose; the decision to merge stays with people.

Sources

Newsletter

Dev Radar, reviews and regulation notes for enterprise AI teams. No hype, only checked facts.

Newsletter

Dev Radar, reviews and regulation notes for enterprise AI teams. No hype, only checked facts.

Newsletter

Dev Radar, reviews and regulation notes for enterprise AI teams. No hype, only checked facts.

Book a readiness call.

Bring one process, product or function where AI should help. We will suggest the most practical next step.

Book a readiness call.

Bring one process, product or function where AI should help. We will suggest the most practical next step.