Runtime AI Scientist: Can AI Scientists Coordinate at Runtime?
We are open-sourcing Runtime AI Scientist, the code behind our paper on Runtime Agent Coordination (RAC). Instead of a fixed workflow, the agent that just finished picks who acts next from the live state of the research. Runtime selection had the highest mean score on all three AI-scientist hosts we tested.
Multi-agent AI scientists can already take a research question from the literature to a written report. Most of them coordinate at design time: a workflow written in advance decides which agent acts next. Human research groups work differently. When an experiment fails or a result surprises them, they change who does what.
Our paper, Can AI Scientists Coordinate at Runtime?, asks whether AI scientists can do the same. The code is now open source as Runtime AI Scientist, with a project page and a 46-second video.
Key takeaways
- Runtime coordination is a different decision from communication. Letting agents talk to each other is not the same as letting the current agent decide who works next. The study separates the two.
- Same agents, different coordination. Runtime Agent Coordination (RAC) wraps existing AI-scientist hosts without replacing their agents, models, tools, or permissions, and holds the budget fixed.
- Runtime selection scored highest on every host. On ResearchClawBench, choosing the next agent at runtime (R2) had the best mean score for ARK, Agent Laboratory, and EvoScientist.
- More mechanism is not automatically better. Adding scoped contracts and verification (R3) lowered those means under the same budgets.
- It is an exploratory, single-seed result. More seeds, tasks, and models are the most useful next experiments, and the repository is set up for other groups to run them.
At a glance
| Item | Detail |
|---|---|
| Paper | arXiv 2610.00980, October 1, 2026 |
| Code | github.com/systemind-team/Runtime-AI-Scientist (MIT license) |
| Hosts | ARK, Agent Laboratory, EvoScientist |
| Benchmark | ResearchClawBench: 10 tasks (5 for ARK), seed 0 |
| Conditions | N0 native workflow, R1 + communication, R2 + runtime selection, R3 + contracts and verification |

What "coordinating at runtime" means
An AI-scientist host is an existing system such as Agent Laboratory, made of specialised agents for literature review, planning, experiments, analysis, and writing. In its native workflow, each agent hands off to a fixed successor.
Under RAC, the agent that has just finished looks at a fresh checkpoint: the artifacts produced so far, the open problems, the execution history, and the remaining budget. From that checkpoint it selects the next authorised agent. There is no separate orchestrator. If a run fails, the work can go back to the coding agent; if results are ready, it can go forward to analysis.
RAC can also attach a work contract to a handoff, which states what the next agent should change and when it is done. The returned artifacts are then checked, and the verdict (supported, refuted, or inconclusive) is passed on to whoever acts next. The verdict is advisory: it never blocks a step, forces a retry, or throws work away.
Four cumulative conditions
Each host runs under four conditions, and each condition adds one mechanism to the one before:
| Condition | What it adds |
|---|---|
| N0: Native lifecycle | The host's own scheduler, end to end |
| R1: Runtime communication | Agents exchange requests, results, and artifacts; handoffs still follow the native order |
| R2: Runtime selection | The current agent selects the next agent from a fresh checkpoint |
| R3: Contracts and verification | Scoped work contracts plus artifact-grounded, advisory verification |
Within a comparison, the model, tools, permissions, inputs, and budget are identical. Coordination overhead is charged to the same budget.
Results
Mean ResearchClawBench score per host and condition (Table 3 of the paper):
| Condition | ARK | Agent Laboratory | EvoScientist |
|---|---|---|---|
| N0: Native lifecycle | 16.66 | 5.47 | 15.99 |
| R1: Runtime communication | 17.40 | 9.63 | 12.07 |
| R2: Runtime selection | 18.42 | 12.08 | 18.53 |
| R3: Contracts and verification | 17.98 | 9.88 | 15.96 |
Runtime selection had the highest observed mean on every host. Its estimated cost per task was lower than the native workflow's for ARK and Agent Laboratory, and about three times higher for EvoScientist. Contracts and verification reduced the means, and whether R3 beat the native workflow depended on the host.
These numbers come from a single seed on a small task set. They motivate runtime coordination; they do not settle it.
Run more experiments with us
The repository contains everything needed to rerun the ResearchClawBench study. That includes one shared N0–R3 policy, a Docker image per host, revision-pinned upstream projects, and an external scorer that alone sees the benchmark answers. The experiments we would most like to see are:
- more seeds on the paper's tasks, to test whether R2's lead holds up;
- more tasks and domains, to find where runtime selection helps and where it hurts;
- other models and budgets, including comparisons at equal spend;
- new hosts, by exposing another AI scientist through the
HostBridgeprotocol.
To take one on, open an experiment proposal. Results that come with the evidence in the repository's reporting checklist are listed in the README, with credit.
Sources
- Paper: Can AI Scientists Coordinate at Runtime? (arXiv 2610.00980)
- Code: systemind-team/Runtime-AI-Scientist
- Project page and video: systemind-team.github.io/Runtime-AI-Scientist