AI agents can turn photos into 3D scenes, but still cannot reliably judge their own work
LEGO-Anything shows coding agents can build editable 3D scenes from photos, but LEGO-Bench exposes major gaps in accuracy and self-assessment.
Latest News and Analysis in Coding AI
LEGO-Anything shows coding agents can build editable 3D scenes from photos, but LEGO-Bench exposes major gaps in accuracy and self-assessment.
The Economist asks whether developers will outpace every other profession in AI use, highlighting coding’s unusual fit with measurable automation.
Block’s open-source Goose offers local AI coding agents without subscriptions, challenging Claude Code’s premium pricing while exposing trade-offs in quality and hardware.
DeepSeek is reportedly forming a team to develop AI agents that could compete with Anthropic’s Claude Code, intensifying pressure in coding tools.
An OpenAI-backed report says coding agents can speed research software upgrades dramatically, but experts still must verify scientific correctness.
OpenAI says scientists are using AI coding agents to update legacy research software and speed genomics work, signaling a new enterprise AI battleground.
Anthropic has introduced Claude Opus 5, pitching stronger coding performance, lower costs and new API features for agents and enterprise AI use.
Poolside released Laguna S 2.1, an open-weight coding model that it says rivals far larger systems by improving persistence and verification.
monday.com says it runs production AI agents on Amazon Bedrock, offering a rare look at the architecture and controls behind enterprise coding workflows.
OpenAI’s Codex now encrypts internal agent handoffs, limiting developer visibility into delegation and raising reliability questions for AI coding workflows.
OpenAI says about 30% of SWE-Bench Pro tasks may be broken, raising new doubts about how the AI industry measures coding models.
Fortune reports Japan is emerging as a key market for Devin as aging software systems and a shrinking developer base sharpen demand for AI coding agents.
Alibaba's open-source Qwen3.6-27B outperforms its 15x larger predecessor on most coding benchmarks with only 27 billion parameters.
Anthropic launches Claude Sonnet 4.6, delivering frontier AI performance in coding, computer use, and agents with 1M token context window, just 12 days after Opus 4.6.
Industry reports suggest Claude Sonnet 5 will match Opus 4.5 capabilities while costing 50% less, with enhanced coding and agent capabilities.
Anthropic introduces a fast mode for Claude Opus 4.6 that delivers responses up to 2.5 times faster, revolutionizing AI-powered software development and coding workflows.
Anthropic launches Claude Opus 4.6, featuring improved coding skills, longer agentic task sustainability, and innovative agent teams functionality for enterprise applications.