Welcome back to Agentic Coding Weekly. What if instead of sandboxing the agents we sandboxed ourselves?

Here are the updates on agentic coding tools, models, and workflows for the week of Sep 27 - Oct 3, 2026.

1. Tool and Model Updates

OpenAI dev day was last Tuesday. Important things were: Dots (something similar to Grok bot and Muse), GPT 6.1 Sol, voice conversations in Codex CLI with /voice, $200 plan now only gives 10x usage, a new $500 plan that gives 25x usage, and a feature to sign in with ChatGPT on third party websites and use your chatgpt token allowance there.

6.1 Sol came out just 7 days after 6 Sol. Considering the reception for 6 Sol was not good compared to 5.6 Sol and Opus 5.5 is beating every other model across the board, this makes sense. 6.1 Sol performs close to 6 Astra and might have been planned to release as Astra minor but got rebranded to 6.1 Sol as per rumors.

Anthropic released Sonnet 5.5. Google published a blog post about Gemini 4 Argon, not a model that can be used.

Claude Code can be modded now to customize the UI or to change how it behaves. There is a new built-in plugin called You Should Know that is implemented with this mod feature. This plugin continuously watches claude's output and highlights important infos buried in the response.

Pi 1.0 shipped with MCP support and deferred tool loading. Alongside 1.0, there is a new experimental package called Pi Durable for long running agent processes that can survive crash.

DeepSeek Harness now has desktop apps for mac and windows. It has telemetry on by default though. Use this fix before the first launch to disable it.

Here's the current state of coding benchmarks:

Model

DeepSWE 1.1

Frontier Code 1.1 Main

Terminal-Bench 4.0

Pricing

GPT 6.1 Sol

75.2%

50.2%

-

$2 / $10

Claude Sonnet 5.5

-

52.1%

70.6%

$2 / $10

Claude Opus 5.5

-

54.6%

66.4%

$4 / $20

MiMo 2.6 Pro

71.9%

-

34.9%

$0.435 / $0.87

GPT 6 Sol

68.8%

49.3%

31.2%

$2 / $10

DeepSeek V4.1 Flash

74%

-

31.2%

$0.15 / $0.6

GPT 6 Astra

74%

53.3%

57.9%

$10 / $50

Fable 5.1

-

50.9%.

55.8%

$10 / $50

Kimi K3

69%

44.2%

-

$3 / $15

2. Open Source Corner

  1. livenerf - long-running benchmark for detecting whether a frontier model gets nerfed

  2. audionaut - audio editing app with MCP support

  3. rhun - code editor written in assembly

3. Reading List

  1. Security in the LLM age - talk from Greg Kroah-Hartman (linux stable maintainer) fact-checking the marketing claim made by Anthropic in Feb this year that Mythos found 79 vulnerabilities in linux. None of them were serious bugs. The reported bug count was wrong, 3 were completely made up, 15 were picked up from discussions of bugs in public that were already fixed in the past, and only 10 resulted in minor bugfixes

  2. GLM-5.3 and the spread of advanced cyber capabilities - Anthropic wrote a great advertisement for GLM 5.3 blinded by its war against open-weights models

  3. How our vibe coded website looks like a designer made it - the author put so much effort that I wouldn't call it vibe-coded anymore

  4. Ask HN: What are you reading? - (bonus) are we still reading?

That’s it for this week. I’ll be back next Monday with the latest agentic coding updates.

— Prashant

Reply

Avatar

or to participate