2026 In LLMs (So Far)
AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

In a September 27 keynote, developer and writer Simon Willison described how LLM coding agents became reliable enough for daily use after late-2025 model releases, prompting developers to take on more ambitious projects. His account also points to ongoing concerns about security and the effects of rapid change on software engineers; the year is still underway.

Developer and writer Simon Willison published notes on September 27 from a keynote reviewing 2026’s developments in large language models, saying coding agents became reliable enough for day-to-day use after late-2025 model releases. His account describes a shift in how developers approach projects, while noting that concerns about agent security and the pace of change remain unresolved.

Willison delivered the closing keynote at the WeAreDevelopers World Congress North America in San Jose on the Friday before his September 27 post. He said he organized the talk as a chronological tour of the year and shared annotated slides and notes alongside a video of the presentation. He cautioned that 2026 was not over, framing the talk as a progress report rather than a final accounting.

In his timeline, the shift began with Claude Opus 4.5 and GPT-5.1, released in November 2025. Willison described the models as incremental improvements, but said that, paired with Anthropic’s Claude Code and OpenAI’s Codex, they crossed a practical threshold: coding agents moved from often making mistakes to being reliable enough for daily use. That assessment is his account of working with the tools, not a benchmark result presented in the material.

Willison said that after the December holidays, developers began experimenting with the new model-and-agent combinations, and entered January keen to put them to work. He described changing his own New Year’s resolution from staying focused on existing projects to taking on more projects, in part to test the technology’s limits. He also recounted building a JavaScript interpreter and a WebAssembly runtime in Python with coding agents, projects he said helped temper his initial enthusiasm about engineering agent harnesses.

At a glance
recapWhen: Published September 27, 2026; covers de…
The developmentSimon Willison published notes from a September 27 keynote reviewing how LLMs and coding agents developed during 2026 so far.

Coding Agents Change Project Ambitions

Willison’s account captures a practical change in software development: as he experienced it, coding agents became useful in routine work, encouraging developers to attempt projects they might previously have avoided. That could affect how software teams divide tasks and plan work, though the keynote notes do not quantify adoption or productivity gains across the industry.

The same shift has a human dimension. Willison said the pace of change prompted him to reconsider what it means to be a software engineer, and recalled a term coined on a podcast with Adam Leventhal: “Deep Blue,” describing a feeling of listlessness when AI seems able to do anything. Several conference speakers touched on the theme, he said. Those remarks suggest professional uncertainty is part of the story alongside new capabilities, but they do not establish how widespread the feeling is.

Amazon

Top picks for "llms"

As an affiliate, we earn on qualifying purchases.

From Late 2025 to January

Willison’s timeline begins before the calendar year. He said Claude Opus 4.5 and GPT-5.1 arrived in November 2025, and that coding-agent tools such as Claude Code and Codex had already existed. In his telling, the change came from combining improved models with those agent systems, rather than from the launch of coding agents themselves.

He also used a recurring informal challenge—asking models to generate an SVG of a pelican riding a bicycle—to illustrate limits in image generation. His November examples still had difficulty rendering bicycles. The exercise is explicitly a “stupid benchmark” in Willison’s words, and offers only a narrow example, not a general measure of model capability.

Earlier in the year, Willison had predicted that it would become undeniable that LLMs could write good code. In September, he said he thought that point had been reached. He also had forecast progress on sandboxing and warned of a possible severe coding-agent security incident. He said security had drawn substantial conference attention, but that the specific scenario he predicted—agents being hijacked and causing real-world economic damage—had not played out as of the talk.

““reliable enough to use on a day-to-day basis””

— Simon Willison, in his September 27 keynote notes

Security Risks Remain Unsettled

The keynote notes do not provide independent measurements of coding-agent reliability, productivity, or how widely developers use the tools. Willison’s assessment reflects his own experience, and the material does not set out comparable tests across models or workplaces.

Agent security remains an open concern. Willison said many conference sessions addressed sandboxing and agent security, but did not describe a specific incident count, a completed solution, or evidence that the risks have been resolved. He said the particular high-impact attack scenario he had predicted had not occurred by the time of his talk. The notes also do not establish how common “Deep Blue” or AI-related overwork is among software engineers.

The Year’s Remaining Tests

Willison presented his keynote as an interim review, so further developments may change the picture before the end of 2026. The next indications to watch include whether coding agents remain dependable in everyday software work, how developers and organizations handle security and sandboxing, and how the profession adapts to changing capabilities.

Willison said listeners could ask him at year’s end whether his decision to take on more projects had been a good one. That question, alongside the security issues he raised, leaves his personal experiment and the broader effects of the tools open for later assessment.

Key Questions

What is the main development in Willison’s 2026 recap?

Willison said coding agents became reliable enough for daily use when paired with improved models released in November 2025, leading developers to take on more ambitious projects.

Which models does he identify as a turning point?

He points to Claude Opus 4.5 and GPT-5.1, released in November 2025. He describes their improvement as incremental but says it made a practical difference when paired with coding-agent tools.

Did Willison report that coding-agent security had been solved?

No. He said sandboxing and agent security received attention at the conference, while the high-impact hijacking scenario he had predicted had not occurred by the time of his talk. The notes do not say the risks are resolved.

Is the 2026 account a final review of the year?

No. Willison published the recap on September 27, 2026, and said in the keynote that the year was not over.

Source: rss

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Why AI-Ready Desktop Workstations Matter for Smaller Teams

Offering powerful, customizable hardware, AI-ready desktop workstations enable smaller teams to handle demanding workloads—discover why they are essential for your team’s success.

What Mobile Workstation Laptops Still Do Better Than Desktops

Beneath their portability, mobile workstation laptops offer unique advantages over desktops that can transform your remote work experience—discover how inside.

The AI That Can Predict the Future With Eerie Accuracy

Future predictions powered by AI promise astounding accuracy, but what ethical dilemmas and limitations lurk beneath this technological marvel?

Rethink AI Standards: Sovereignty Isn’t About Country Of Origin

Europe shifts its AI sovereignty stance, redefining sovereignty beyond country origin, raising questions about measurement and legal standards.