AI Agents Expand Model-Development Work, but Humans Still Control the Critical Calls

A Fudan-linked study finds AI agents propose much of model-development work, while humans retain final control—revealing autonomy’s limits in practice.

AI News

AI agents are taking on a growing share of the work involved in building AI models, but a new analysis suggests they remain far less influential when researchers decide what to pursue and how to proceed.

A research team involving China’s Fudan University reviewed 769 task logs from a model-development project involving 56 participants and the agents they used. The study found that AI supplied up to 55.4% of proposals for methods and parameters, while humans made 85.5% of the final decisions in those areas. Humans also selected project goals and scope in 93.4% of cases.

The findings offer a more qualified picture of agent autonomy than raw activity measures might suggest. Agents performed work, generated options, and handled revisions, but human judgment remained the main control point—particularly when the project faced ambiguity, failures, or strategic choices.

Atria Dawn Preview as a test case

The project centered on Atria Dawn Preview, an agentic language model built on a mixture-of-experts architecture with 744 billion parameters. The team designed it for research and engineering tasks, using a workflow in which each task was connected to an execution environment. Agents could call tools, produce intermediate results, and receive feedback from tests, metrics, or source evidence.

The researchers reported that AI was used in 96.5% of the reviewed tasks. Over four weeks, the median ratio of agent actions to human inputs increased from 11 to 28.5. That rise could appear to show agents becoming more independent, but the authors argue that it mainly reflects a different pattern: one human decision often triggered a longer chain of agent actions.

The model reportedly led on five of 16 benchmarks, including web search and cybersecurity. The source does not indicate that Atria Dawn Preview achieved an overall lead against competing systems, so the benchmark results should be treated as a limited performance claim rather than evidence of broad superiority.

Agents proposed options; people chose directions

The most common decision pattern was “AI proposes, human selects,” accounting for 55.4% of decisions about methods and parameters. Across the same category, AI made 9.2% of final decisions, compared with 85.5% for humans. The share of AI proposals varied by decision type, ranging from 17% to 55%, but its role in final selection remained in the single digits.

Human control was even stronger for project-level choices. Researchers made the final call on goals and scope in 93.4% of cases. Among 151 tasks that participants said would not have been feasible without AI, humans still selected the goal 95.4% of the time.

That distinction matters for model development. An agent can search a large space of possible implementations or carry out a complex experiment without deciding whether the experiment is worth running. The study’s evidence suggests that current systems are more effective as research operators and proposal generators than as owners of a research agenda.

AI made more work possible, not just faster

The study also points to a less obvious effect of AI agents. Of 455 completed AI-assisted tasks, participants judged 151—roughly one-third—to be infeasible without AI at the same scope and quality. Those tasks involved 27 of the 56 participants, rather than being concentrated among only a few advanced users.

This result suggests that agents may expand the set of work teams can attempt, rather than simply accelerating existing workflows. For AI builders, that could mean more experiments, evaluations, and engineering changes entering a project pipeline. It could also mean a larger oversight burden, because the number of actions requiring review grows along with the amount of work being attempted.

When tasks became difficult, human intervention was the usual way forward. In a set of 588 tasks with recorded difficulty, 76% progressed through human assistance, while agents solved 23% without intervention. The most common forms of assistance were supplying context or clarifying requirements, at 35.2%, and diagnosing problems or changing methods, at 34.7%.

Humans rarely took over execution directly. Partial edits represented 3.2% of interventions, and full takeovers accounted for 0.7%. After feedback was provided, agents handled revisions themselves in 75.4% of cases. In practical terms, the human bottleneck was usually judgment and information—not manually performing the task.

The oversight problem grows with agent activity

The findings do not eliminate the possibility of more autonomous model development. Instead, they identify a control problem that becomes sharper as agent chains grow longer. If no person can inspect every intermediate action, human review may degrade into approval of summaries or final outputs rather than meaningful supervision.

The researchers also observed that participants often used autonomous modes to avoid interrupting long-running processes. That boundary was reportedly chosen for convenience, not because teams had formally decided how much authority agents should receive. For enterprise AI teams, this is a warning that operational settings can quietly become governance decisions.

The study places current systems between two stages of development: AI has moved beyond being only an object of research or a narrow task tool, but it has not clearly become an independent research director. The authors describe a possible next stage as recursive self-improvement, in which stronger systems help produce their successors, while emphasizing that the path from better performance on training tasks to better model design remains unresolved.

What to watch next

The most important follow-up will be whether similar decision patterns appear in other model-development organizations and projects. This study examined one team’s work, so its observations should not be treated as a universal measure of AI autonomy.

Researchers and buyers should watch for three signals. First, future evaluations should separate agent actions from decisions about goals, methods, and resource allocation. Second, development platforms may need audit trails that expose how a final recommendation emerged from a long agent chain. Third, companies claiming automated AI research should disclose how often humans set objectives, reject proposals, change methods, or intervene after failures.

The debate over recursive self-improvement will also depend on evidence about research judgment, not only coding speed or experiment throughput. Claims from companies such as Anthropic about increasingly automated research direction remain a separate category of evidence from this independently described project and should not be treated as directly comparable without common measurements.

Creati.ai perspective

This study’s central lesson is operational: more agent activity does not necessarily mean more agent authority. AI agents can multiply the number of experiments and engineering tasks a team can run, while humans continue to determine which questions matter and whether the results justify a change in direction.

For builders and enterprise teams, the immediate priority is not simply maximizing autonomous runs. It is designing systems that preserve decision provenance, surface uncertainty, and make it possible to review the choices hidden inside long chains of agent work. Until agents can reliably evaluate research goals before results are available, human judgment remains the scarce resource in model development.

Ads