Back to blog

Runtime AI Scientist: Can AI Scientists Coordinate at Runtime?

We are open-sourcing Runtime AI Scientist, the code behind our paper on Runtime Agent Coordination (RAC). Instead of a fixed workflow, the agent that just finished picks who acts next from the live state of the research. Runtime selection had the highest mean score on all three AI-scientist hosts we tested.

Multi-agent AI scientists can already take a research question from the literature to a written report. Most of them coordinate at design time: a workflow written in advance decides which agent acts next. Human research groups work differently. When an experiment fails or a result surprises them, they change who does what.

Our paper, Can AI Scientists Coordinate at Runtime?, asks whether AI scientists can do the same. The code is now open source as Runtime AI Scientist, with a project page and a 46-second video.

Key takeaways

  • Runtime coordination is a different decision from communication. Letting agents talk to each other is not the same as letting the current agent decide who works next. The study separates the two.
  • Same agents, different coordination. Runtime Agent Coordination (RAC) wraps existing AI-scientist hosts without replacing their agents, models, tools, or permissions, and holds the budget fixed.
  • Runtime selection scored highest on every host. On ResearchClawBench, choosing the next agent at runtime (R2) had the best mean score for ARK, Agent Laboratory, and EvoScientist.
  • More mechanism is not automatically better. Adding scoped contracts and verification (R3) lowered those means under the same budgets.
  • It is an exploratory, single-seed result. More seeds, tasks, and models are the most useful next experiments, and the repository is set up for other groups to run them.

At a glance

Item Detail
Paper arXiv 2610.00980, October 1, 2026
Code github.com/systemind-team/Runtime-AI-Scientist (MIT license)
Hosts ARK, Agent Laboratory, EvoScientist
Benchmark ResearchClawBench: 10 tasks (5 for ARK), seed 0
Conditions N0 native workflow, R1 + communication, R2 + runtime selection, R3 + contracts and verification

Same agents, different coordination: native fixed workflows compared with Runtime Agent Coordination

What "coordinating at runtime" means

An AI-scientist host is an existing system such as Agent Laboratory, made of specialised agents for literature review, planning, experiments, analysis, and writing. In its native workflow, each agent hands off to a fixed successor.

Under RAC, the agent that has just finished looks at a fresh checkpoint: the artifacts produced so far, the open problems, the execution history, and the remaining budget. From that checkpoint it selects the next authorised agent. There is no separate orchestrator. If a run fails, the work can go back to the coding agent; if results are ready, it can go forward to analysis.

RAC can also attach a work contract to a handoff, which states what the next agent should change and when it is done. The returned artifacts are then checked, and the verdict (supported, refuted, or inconclusive) is passed on to whoever acts next. The verdict is advisory: it never blocks a step, forces a retry, or throws work away.

Four cumulative conditions

Each host runs under four conditions, and each condition adds one mechanism to the one before:

Condition What it adds
N0: Native lifecycle The host's own scheduler, end to end
R1: Runtime communication Agents exchange requests, results, and artifacts; handoffs still follow the native order
R2: Runtime selection The current agent selects the next agent from a fresh checkpoint
R3: Contracts and verification Scoped work contracts plus artifact-grounded, advisory verification

Within a comparison, the model, tools, permissions, inputs, and budget are identical. Coordination overhead is charged to the same budget.

Results

Mean ResearchClawBench score per host and condition (Table 3 of the paper):

Condition ARK Agent Laboratory EvoScientist
N0: Native lifecycle 16.66 5.47 15.99
R1: Runtime communication 17.40 9.63 12.07
R2: Runtime selection 18.42 12.08 18.53
R3: Contracts and verification 17.98 9.88 15.96

Runtime selection had the highest observed mean on every host. Its estimated cost per task was lower than the native workflow's for ARK and Agent Laboratory, and about three times higher for EvoScientist. Contracts and verification reduced the means, and whether R3 beat the native workflow depended on the host.

These numbers come from a single seed on a small task set. They motivate runtime coordination; they do not settle it.

Run more experiments with us

The repository contains everything needed to rerun the ResearchClawBench study. That includes one shared N0–R3 policy, a Docker image per host, revision-pinned upstream projects, and an external scorer that alone sees the benchmark answers. The experiments we would most like to see are:

  • more seeds on the paper's tasks, to test whether R2's lead holds up;
  • more tasks and domains, to find where runtime selection helps and where it hurts;
  • other models and budgets, including comparisons at equal spend;
  • new hosts, by exposing another AI scientist through the HostBridge protocol.

To take one on, open an experiment proposal. Results that come with the evidence in the repository's reporting checklist are listed in the README, with credit.

Sources