← Latest papers
🤖 AI

Position: AI Agents in Scientific Teams Should Be Studied as Human-Agent Systems

This paper argues that AI agents in scientific discovery should be studied as integrated human-agent systems rather than autonomous entities, emphasizing that this collaborative perspective is essential to mitigate risks like reduced inquiry diversity and to foster effective human-AI synergy.

Original authors: Patrick Emami, Sameera Horawalavithana, Truc Nguyen, Gihan Panapitiya, Bruno Jacob, Siddhisanket Raskar, Saumya Sinha, Jared D. Willard, Andrew Glaws, Nithin Somasekharan, Ling Yue, Brian Lu, Shaowu P
Published 2026-08-18
📖 1 min read☕ Coffee break read

Original authors: Patrick Emami, Sameera Horawalavithana, Truc Nguyen, Gihan Panapitiya, Bruno Jacob, Siddhisanket Raskar, Saumya Sinha, Jared D. Willard, Andrew Glaws, Nithin Somasekharan, Ling Yue, Brian Lu, Shaowu Pan, Jason Eisner

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Technical Summary: AI Agents in Scientific Teams Should Be Studied as Human-Agent Systems

Problem Statement
The paper identifies a critical gap in the current literature regarding Large Language Model (LLM)-based "AI Scientists." While existing research increasingly focuses on the autonomous capabilities of agents to perform scientific discovery, this perspective largely overlooks the social and collaborative dynamics inherent in scientific teamwork. Current systems often treat human participation as limited to high-level goal setting and final verification, failing to account for the iterative, joint collaboration typical of human-human teams. The authors argue that studying AI Scientists in isolation ignores the "social aspects of scientific teamwork" and fails to address near-term risks, such as the homogenization of scientific inquiry and the reduction of topic diversity, which arise when agents are deployed without accounting for human-agent dynamics.

Methodology
The authors employ a multi-faceted approach combining literature review, empirical analysis, and case study examination through the lens of Human-Agent Systems (HAS).

  1. Literature Review & Taxonomy Application: The authors review recent AI Scientist systems (e.g., AI Scientist v2, Kosmos, Co-Scientist, Denario) and categorize them based on the primary channel of human feedback. They adopt a taxonomy from Zou et al. (2025a) to analyze these systems across three dimensions:

    • Type: Evaluative (quality assessment), Guidance (instructions/critiques), or Corrective (edits/fixes).
    • Granularity: Coarse (single assessment for an output) or Fine (targeted step-wise feedback).
    • Phase: Pre-task (planning), Mid-task (execution), or Post-task (evaluation).
      The review reveals that most systems rely on artifact-level (post-task) or discovery-phase-level (mid-task) feedback, with very few supporting flexible, asynchronous, or fine-grained intervention.
  2. Preliminary Empirical Experiment: To test agent behavior regarding feedback, the authors conducted an experiment using a ReAct-based agent harness with GPT-5-mini on the AstaBench End-to-End Discovery (E2E) benchmark. They compared three conditions:

    • A baseline planning-centric agent.
    • An agent with an "ask user" tool and feedback taxonomy in the system prompt.
    • An agent with explicit per-step instructions to always call the "ask user" tool.
      The study utilized an LLM-as-a-judge to classify the type, granularity, and phase of feedback requests.
  3. Case Study Analysis: The authors analyzed published accounts of scientists collaborating with agents (e.g., "Vibe-coding" a complexity theory result with Gemini 3 Pro; reduced physics modeling with GPT-5) to qualitatively illustrate patterns of mutual augmentation and synergy.

  4. Theoretical Modeling: The authors propose a toy utility model for HAS, defined as $U(H, A) = CA(H, A) - CD(H, A)$, where $CA$ represents the collaboration advantage (complementarity) and $CD$ represents the collaboration disadvantage (coordination costs).

Key Results

  • Feedback Limitations: The literature review indicates that the majority of AI Scientist systems restrict human interaction to supervisory roles (scoping, auditing, approval) rather than sustained, iterative collaboration. Artifact-level feedback is the most common, while asynchronous steering and fine-grained interruption are rare.
  • Agent Feedback Behavior: In the preliminary experiment, the agent without explicit per-step instructions failed to know when to ask for help, often making assumptions in ambiguous situations. When forced to ask at every step, the dominant feedback type was evaluative (occurring mid-task), followed by guidance (pre-task). Corrective feedback was rarely requested. This suggests current agents struggle to self-regulate when to seek fine-grained, corrective human input.
  • Risk Evidence: The paper cites evidence that AI-augmented research can produce three times as many papers with a 5% reduction in the diversity of explored topics. It also highlights risks of hallucination, bias proliferation, and the potential for "de-skilling" among junior researchers who offload deliberative processes.
  • Synergy Patterns: Case studies demonstrated that effective synergy occurs when humans provide conceptual framing, domain judgment, and epistemic verification, while agents accelerate implementation and drafting. However, this synergy is fragile; without expert oversight, agents exhibit failure modes like premature confidence and smoothing over complex issues.

Key Contributions

  1. Reframing the Unit of Analysis: The paper argues for shifting the research focus from autonomous "AI Scientists" to Human-Agent Systems (HAS), where the unit of analysis is the human-agent pair.
  2. Taxonomy of Interaction: It provides a structured categorization of current AI Scientist systems based on feedback channels (artifact, phase, planning, asynchronous) and dimensions (type, granularity, phase), highlighting the scarcity of systems supporting deep, iterative collaboration.
  3. Risk Mitigation Framework: It posits that an HAS lens is essential for mitigating near-term risks, including hallucinations, bias, and the erosion of scientific expertise, by ensuring human oversight remains integrated into discovery loops rather than relegated to post-hoc review.
  4. Mathematical Framework Proposal: The authors introduce a utility model ($U = CA - CD$) to guide future research in maximizing collaboration advantage while minimizing coordination costs.
  5. Evaluation Pathways: The paper proposes a staged evaluation strategy for HAS, moving from low-cost proxies (e.g., review time, intervention density, simulated oracle synergy) to more expensive Centaur evaluations (comparisons of HAS against humans and agents alone), to make the assessment of human-agent synergy feasible.

Significance and Claims
The paper claims that studying AI Scientists as HAS is both underexplored and undervalued. Its significance lies in:

  • Mitigating Near-Term Risks: By centering human participation, the HAS lens offers concrete pathways to detect hallucinations, prevent bias, and preserve scientific skills that might otherwise degrade through over-reliance on automation.
  • Unlocking Synergy: It suggests that true scientific breakthroughs may require a "human-agent synergy" where the combined team outperforms either member alone, a state currently inaccessible to fully autonomous agents or humans working in isolation.
  • Pragmatic Research Direction: The authors argue that adopting the HAS lens is a pragmatic route to achieving safe and effective human-agent co-discovery. They call for new research to develop mathematical frameworks and evaluation metrics that specifically target the dynamics of human-AI synergy in scientific discovery, rather than treating agents as isolated optimization solvers.

The paper concludes that while the evaluation of HAS is complex and costly, a staged approach using lightweight metrics can validate the utility of these systems, ultimately fostering a scientific environment where AI augments rather than replaces human deliberation.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →