LEGO-Anything shows coding agents can build editable 3D scenes from photos, but LEGO-Bench exposes major gaps in accuracy and self-assessment.

Coding agents are getting better at turning a single photograph into an editable 3D scene, but new research suggests they still cannot reliably tell when their reconstruction is wrong.
The University of Maryland and AWS project, called LEGO-Anything, asks an agent to write Blender code from an image, render the result, inspect it, and revise the program. Its accompanying benchmark, LEGO-Bench, found that the strongest tested configuration reached 53.4% accuracy on indoor scenes and 39.6% on outdoor scenes. More importantly, the agents’ ability to judge whether one reconstruction was better than another landed near or below chance.
That combination matters for product teams building AI-assisted design, simulation, robotics, and spatial computing tools. An agent that can produce a usable scene is valuable, but an agent that cannot recognize geometric errors may make iterative workflows difficult to trust without independent checks.
LEGO-Anything treats image-to-3D reconstruction as a programming task rather than a single generative-model output. The agent receives one image and writes code for Blender, the open-ended 3D creation platform. It can then run that code, inspect the rendered scene, and make further edits.
The result is intended to be more useful than a static image or opaque 3D asset. A program can expose objects, geometry, layout, and camera position. Users can inspect the scene, change individual elements, or query it as part of a larger workflow.
The approach also reflects a practical direction for AI agents: generate an artifact, execute it, evaluate the result, and refine the underlying code. In principle, that loop should allow an agent to correct mistakes without requiring a new model for every scene type.
But a single photograph does not reveal exact depth or hidden geometry. The researchers therefore needed a controlled way to measure reconstruction quality without making the inputs look obviously synthetic.
LEGO-Bench contains 208 images representing 104 indoor and outdoor scenes, built from 443 registered assets. According to the research description reported by The Decoder, the images are rendered from professionally created simulator scenes. That gives the benchmark access to hidden ground-truth geometry, depth, and object assignments while preserving a more natural visual appearance.
The benchmark evaluates three dimensions. Validity checks whether the agent delivered a usable scene artifact. Reconstruction measures the accuracy of visible geometry. Appearance compares a re-rendered submission with the reference image at the pixel level.
All six tested GPT configurations generally produced working scenes. Their quality, however, varied widely. The reported leader, GPT-6 Astra, reached 53.4% on indoor scenes and 39.6% outdoors, while weaker configurations scored around 15%. The results are benchmark findings from this research, not evidence that the model will perform at the same level on arbitrary real-world photographs.
Complexity made performance worse, and outdoor scenes were harder than interiors. The researchers also reported that giving models a larger reasoning budget improved results substantially. On an office subset, GPT-6 Astra’s score rose from 32.3% to 61.8% under a higher reasoning budget.
The more revealing failure appeared in self-evaluation. When agents had to choose which of two versions better matched the original image, their geometric judgments were close to or below a coin flip. In practical terms, an agent could make a damaging edit and fail to recognize that the scene had become less accurate.
That finding places a limit on a common assumption about agentic workflows: that repeated visual inspection will automatically lead to convergence. The research indicates that iteration alone is not enough when the agent’s internal assessment is unreliable.
The researchers responded with LEGO-Plugin, an extension that does not require additional model training. It anchors the initial scene to the reference image, replaces the agent’s unreliable self-judgment with concrete measurements, and protects correct progress from later regressions.
The plugin improved all six tested models, according to the reported results. The largest gains reached 62.7% for weaker agents, while the strongest model gained roughly two percentage points. That spread is important: evaluation and workflow controls may offer more value for weaker systems than simply increasing model capability.
The reconstructed scenes were also tested as inputs for standard computer-vision tasks. Object detection performed best, reaching roughly half the performance of the specialized model DINO. Segmentation and depth estimation were further behind specialized systems including SAM 3 and Depth Anything 3.
Those comparisons suggest that executable scene programs are already usable as intermediate representations, but they are not yet reliable replacements for task-specific vision models. A scene can be valid enough to edit or inspect while still being too inaccurate for downstream measurement.
For developers, the immediate lesson is to separate artifact generation from artifact verification. A coding agent can write a Blender scene, but production systems may need geometric constraints, image-based scoring, object-level checks, and safeguards against regressions. Human review may remain necessary when the scene feeds manufacturing, architecture, robotics, or safety-sensitive simulation.
The work also points to a deployment trade-off. More reasoning can improve quality, but it may increase latency and inference cost. A measurement-driven plugin could provide a more targeted way to improve results than simply allocating a larger reasoning budget to every task.
Enterprise buyers should treat “editable 3D from a photo” as a capability with several levels of usefulness. A scene that looks plausible may support rough visualization or content blocking. It should not automatically be treated as an accurate digital twin, a dependable depth map, or a substitute for specialized reconstruction software.
The broader market is already exploring adjacent approaches. Unity has released plugins for Claude Code and Codex, while World Labs’ Atlas takes a model-based route to scene reconstruction. Google DeepMind’s GenCeption uses a video model for depth estimation and segmentation. These systems are not directly comparable to LEGO-Anything, but they show that 3D software integration, world models, and specialized perception pipelines are developing in parallel.
The key signal will be whether measurement-based refinement continues to work on real photographs rather than benchmark scenes with hidden simulator ground truth. Future evaluations should also show how performance changes with occlusion, unusual camera angles, reflective surfaces, clutter, and objects missing from an agent’s asset library.
Builders should watch for benchmarks that report cost, latency, edit stability, and downstream task performance alongside reconstruction scores. It will also matter whether agents can explain which parts of a scene are uncertain, rather than returning a single confidence judgment that masks errors.
Another important test is integration into production tools. If systems such as Blender, Unity, Claude Code, or Codex can combine generation with independent geometric checks, they may become useful for supervised workflows before they are trusted with autonomous scene construction.
LEGO-Anything is less a demonstration that agents have solved 3D reconstruction than evidence that evaluation is becoming the central engineering problem. Producing a scene is only the first step; determining whether the scene preserves the right geometry is what separates a creative prototype from a dependable tool.
The research also offers a practical design principle for AI products: do not ask an agent to be the sole judge of its own work when objective measurements are available. For teams building spatial or code-generating systems, independent verification may matter as much as the underlying model’s visual reasoning.