Research

AI Agents Proposed Over Half the Ideas, Humans Made 85 Percent of Calls

A team logged 769 tasks while building its own AI model: agents proposed up to 55% of methods, humans made 85%+ of final calls, and a third of tasks needed AI to exist at all.

AI agents do more of the work in model development, but humans still make the decisions
AI agents do more of the work in model development, but humans still make the decisionsstriatic / Openverse
By James Calloway5 min read

Updated

Why it matters

  • AI agents supplied up to 55 percent of method proposals across 769 analyzed task logs from building an AI model.
  • Humans made more than 85 percent of final decisions in the same workflow.
  • One third of the logged tasks would not have been attempted without AI.
  • The authors warn that more agent activity does not mean more autonomy.

A research team that logged its own AI model development work reports that AI agents supplied up to 55 percent of method proposals — yet humans made more than 85 percent of the final decisions. The study, based on 769 task logs recorded during the construction of an AI model, offers a rare inside view of how machine assistance actually functions in a frontier-adjacent research workflow.

The numbers cut against two popular narratives at once. AI agents are not idle spectators in modern model development — they generate a majority of the proposals for how to proceed. But neither are they running the lab. The decisive calls, the ones that determine what gets built and how, remained overwhelmingly in human hands.

The team also found that a third of the tasks in the logs would never have been attempted without AI. That figure may be the most consequential of the three. It suggests that agent assistance does not merely accelerate existing research agendas. It expands the set of problems a team is willing to touch at all.

What the logs actually show

The dataset consists of 769 task logs produced while the researchers built their own AI model. Each log captures a unit of work — a proposal for a method, an experiment, a decision point. By analyzing these logs, the team could measure, rather than assume, where machine contributions entered the pipeline and where human judgment terminated it.

Up to 55 percent of method proposals came from AI agents. That is a striking share for a domain that has historically prized human intuition about architectures, training procedures and evaluation strategies. If more than half of the ideas on the table originate with a machine, the traditional picture of the human researcher as the sole source of scientific creativity no longer describes practice.

The decision statistics tell the complementary story. Humans made more than 85 percent of final decisions. Proposals, in other words, are cheap; commitments are expensive. The workflow the logs reveal is one in which agents generate options and humans filter, reject and select. The bottleneck has shifted from ideation toward judgment.

The third finding — that roughly one third of tasks would not have been attempted without AI — indicates that the agents changed the boundary of the feasible. Work that was previously judged too costly, too marginal or too tedious entered the pipeline because an agent could carry part of the load.

Why the authors issue a warning

The research team draws a sharp line between activity and autonomy. More agent activity, they warn, does not mean more autonomy. An agent that proposes ten methods and has none of them adopted has high activity and zero authority. The 55 percent proposal share and the 85 percent human decision share, read together, are precisely the signature of that distinction.

This matters because casual observers often equate volume of AI involvement with delegation of control. The logs argue for a more careful accounting. Who proposed the step and who approved it are different questions, and the study's data show the answers diverge dramatically. Adoption, not generation, is where power over the research process sits.

The warning also has a measurement implication. If organizations track only how often agents participate, they will systematically overstate how much autonomy has been transferred. The authors' framework separates proposal from decision, and the gap between the two figures — 55 percent versus a human share above 85 percent — quantifies how much of the apparent AI takeover dissolves under closer inspection.

Why the stakes extend beyond one lab

The study lands amid an industry-wide push to deploy AI agents across software development, research and operations. Vendors market agents as systems that can plan, execute and complete multi-step work with limited supervision. Enterprises are making purchasing and staffing decisions based on those claims. Hard telemetry from a real research project — 769 logged tasks, counted and analyzed — is scarce relative to the volume of assertion in the market.

For research organizations, the findings sketch a plausible equilibrium: agents as proposal engines, humans as gatekeepers, and a measurable expansion of the task frontier. For policy and governance discussions about AI self-improvement, the data point in a specific direction. If humans retain the vast majority of final decisions even when agents dominate proposal generation, then fears of runaway autonomous research loops are, at least in this documented case, not borne out by the workflow evidence.

The one-third figure deserves equal attention in that debate. A system that enlarges what gets attempted changes research direction even without changing who decides. Influence over the agenda is a form of power that decision statistics alone do not capture. The authors' warning about conflating activity with autonomy cuts in this direction too: a team may retain formal decision authority while its effective research scope is shaped by what the agents make tractable.

The measurement question the study opens

Self-logged, self-analyzed data has limits, and the study concerns one team building one model. But its method — decomposing the workflow into proposals, tasks and decisions, then counting each — is portable. Other labs could apply the same instrumentation and produce comparable figures. A body of such measurements would replace the current argument-by-anecdote about how much work agents actually do and how much control they actually hold.

The three headline numbers also frame the natural follow-up questions. Does the human decision share fall as models improve? Does the proportion of previously unattempted tasks grow? And at what point, if any, does the proposal share translate into de facto autonomy, as humans rubber-stamp machine-generated options faster than they evaluate them? The present study establishes the baseline against which those shifts can be judged.

For now, the documented reality is unambiguous on one point: in this lab, at this stage, AI agents did a great deal of the work of imagining — and humans did almost all of the work of deciding.

Original: huggingface.co

Share this article:

More from James Calloway

James Calloway

Show full bio

News editor covering industry trends and analytics at AI In Context.

118 articles

Related articles

  1. OpenAI Predicts AI-Made Discoveries by 2026 as Intelligence Costs Plunge
  2. 37 Researchers Propose Replacing the Scientific Paper With AI-First Format
  3. OpenAI and PNNL Launch Benchmark for AI in Federal Permitting

« Previous article