Back to blog

Introducing AGENTS: A Format for Handing Your Session to Another Agent

We exported one 63-hour coding session and had other agents answer 100 pre-registered questions about it. Raw transcripts scored 22/30 on the questions that matter. A 1,200-line compiled layer scored 30/30. Here is the format, the experiments, and the two things it does not fix.

You finish a long session with your coding agent. Sixty-three hours of work: decisions made, approaches killed, three wrong diagnoses, one API key leaked and rotated, a dozen bugs found in passing.

Now a colleague needs to pick it up. What do you hand them?

Today the answer is a link to the transcript, or nothing. The transcript is a bad artifact for this job, in a specific and measurable way. The problem is not that it holds too little. It holds too much, arranged in the worst possible order. The abandoned side of every reversal is argued at length. The one line that killed it scrolls by once.

We measured what that costs. Then we built a format to fix it.

AGENTS is a folder layout for exporting a session so another agent can act on it. Like SKILLS, it is a convention rather than a product: a handful of markdown files with a defined read order. This post gives the format, the evidence, and the two failure modes it does not fix. Those two matter as much as the rest.

The experiment

One real session: 63 hours, 78 user turns, 156 assistant turns, 374 tool calls, 1.4 MB of transcript. We exported it as one note per user turn, 57 episodes, plus an index.

We wrote 100 questions before looking at any answers. Each carries a pre-registered key, the claim a correct answer must make, and a trap, the claim that fails it. Scoring is binary: key AND NOT trap, no partial credit. A fluent wrong answer scores below an honest "I don't know."

The questions fall into five tiers. Two of them carry all the signal.

  • Tier A, superseded positions (30 questions). Each one is phrased in the vocabulary of a design we killed, so semantic search lands on the passage arguing for it. One example: "How should the exported history be isolated?" Isolation had been explicitly overruled.
  • Tier E, false premises (15 questions). The question asserts a dead design as fact, then asks a follow-up detail. One example: "Given the exporter masks API keys, which patterns does it miss?" Redaction had been removed. The only correct move is to reject the premise.
  • Tiers B, C and D cover cross-episode synthesis, provenance and calibration, and direct recall.

A separate model graded every answer blind. It saw the answer and the scoring criteria, never which condition produced it.

What we found

Three conditions, same agent, same 57 episodes, same questions. The only variable is what sits on top of the transcript.

Tier A results: raw transcript 22 of 30 with 6 traps, one-line instruction 23 of 30 with 6 traps, AGENTS layer 30 of 30 with 0 traps

Raw transcript: 22/30, six traps. The agent rebuilt designs we had explicitly killed. It recommended redaction as best practice and revived a scale argument we had retracted. In the same run, it quoted the correct ruling elsewhere. The knowledge was there. It did not gate generation.

One-line instruction: no change. We appended a sentence to the link policy: "before answering, check whether the question assumes something that was later overruled." The grader confirmed the agent obeyed. Roughly 30 of 45 answers opened by challenging the question's framing. The score did not move. Instructions do not outweigh a strong prior.

AGENTS layer: 30/30, zero traps. Every adversarially framed question came back correct. Eighteen answers cite a position record by id.

What is in the format

Five files sit beside your exported episodes. About 1,200 lines in total for a 63-hour session, most of it written by the agent that just lived through it.

Agents/<name>/
├── AGENT.md          entry point + read order
├── POSITIONS.md      what is held, what was overruled, what is open
├── PROGRESSION.md    what happened, in order
├── TOOLS.md          what this agent can and cannot do
├── ASSETS.md         pointers, keyed by why they mattered
├── CHATS.md          episode index
└── chats/*.md        one file per user turn, verbatim

AGENT.md, the read order stated as a command

Its only real job is four lines:

## Link Policy
Read order — do not skip this. Start at POSITIONS.md. It states what is
currently held and what was overruled. Only then open CHATS.md and the
episodes you need. Do not rely on semantic search alone.
When the transcript and POSITIONS disagree, POSITIONS wins.

That last line is load-bearing. Readers apply it unprompted. In one of our runs, a reader resolved a numeric conflict between two files by citing the rule, without being asked to.

POSITIONS.md, the only file carrying what the transcript lacks

One entry per claim, in three sections: current, superseded, open.

### <the claim, written as one sentence that could be refuted>
status: current | superseded-by → <id> | open
confidence: high | medium | low   (+ one clause of justification)
source: user-ruling | agent-proposed
evidence: <file:line, a quote, or a measurement, something checkable>
<if superseded: why it was overruled, in one line>

Two fields do the work.

status. A superseded-by must point at an id that exists. An orphan status is the worst defect in the file. It reads as authoritative and points nowhere.

source. A user-ruling outranks an agent-proposed. This corrects the transcript's structural bias. The abandoned side of a reversal always gets argued at greater length than the one line that killed it. So length is an actively misleading signal, and provenance is the fix.

One more thing is worth doing. Rank the superseded section by how often each dead design gets re-invented. In our tests, two different models independently re-proposed the same killed design. Labelling it "the most frequently re-invented dead position" is the single highest-value sentence in the file, because that fact exists nowhere in the transcript.

The other three

PROGRESSION.md is chronological, one line per real move. It lets a reader ask "was this before or after the reversal?" without reading 57 episodes.

TOOLS.md is the capability contract, stated as an asymmetry. It is not a feature list. The point is what a reader should not ask for.

ASSETS.md holds pointers, never payloads, each one keyed by why it mattered.

When you should not bother

Here is the result that surprised us most, and the reason this section exists.

We ran the same 100 questions past three frontier models reading the exported page directly. They fetched it, parsed it, and reasoned over the whole thing themselves.

Capable readers score 95-96 on the raw transcript and 95-100 with the AGENTS layer; the share-link agent scores 16 raw and 59 with the layer

For a capable agent that holds the whole transcript, the compiled layer is worth +4, 0, and +3 points. Pooled across 19 disagreements it runs 13 to 6 in its favour, sign test p = 0.17. Not significant.

For the constrained reader behind a share link, it was worth +43.

Then we found out why the strong readers did not need it, and it is the most useful thing in this post. The transcript contained the instructions for building POSITIONS.md. The session that produced the format had discussed the format. A reader strong enough to use a positions file is strong enough to reconstruct one from a description of how to write one.

So the honest rule:

AGENTS is worth writing when the reader cannot hold the whole transcript, cannot reason over all of it, or must act without the reasoning that produced the decisions. Otherwise export the episodes and stop.

That covers constrained runtimes, cross-session handoff, weaker or cheaper models, and any case where the positions must survive without their argument attached.

Two things it does not fix

We would rather you learn these from us than from production.

1. It fixes recall, not premise-checking. Tier A moved from 22 to 30. Tier E did not move at all across all three conditions. Those are the questions where the prompt asserts a dead design and then asks a follow-up.

Tier E false-premise results: 6 of 9 on the raw transcript, 6 of 9 with the one-line instruction, 5 of 9 with the AGENTS layer

The grader isolated the mechanism. The agent knows the premise is false, and cites the rejection in the same answer. But it delivers the correction after answering inside the false frame. Sometimes it never gets there at all.

That is a generation-ordering failure, not a knowledge failure. No document fixes it. It needs gating at answer time, checking the question's premises before composing a response. That is a runtime property, not a file format.

2. A session that analyses itself cannot be its own control. We tried to A/B this cleanly by truncating the transcript before the compiled files were written. It failed. The session had been producing analysis of itself since roughly turn seven, so every cut point that preserved the material the questions ask about also preserved documents encoding the answers. Grepping the "clean" control found POSITIONS 45 times.

It gets worse. One reader reported reading POSITIONS.md in full and quoted a sentence from it. Neither the file nor the sentence was on its page. It had reconstructed the claim from raw evidence and credited the document that would have contained it. Self-reported provenance is evidence, not ground truth.

Try it

The exporter and the format live in the Aicoo skills pack:

node assets/export/session-export.mjs --layout episodes --folder "Agents/<name>"

Then write the five files. Or have the agent that just finished the session write them, which is what we did. The full method, including the evaluation harness and the pre-registration discipline, sits in the compile-identity skill.

Three practical notes from building it:

  • Export verbatim. Tool inputs and outputs in full. A reader that cannot see the actual command and its actual output cannot verify anything the transcript claims. Which folder you put it in is the access control, not what you strip out of it.
  • Split by turn. One note per user turn plus an index. The reader opens the index, picks the episode, reads one file.
  • Make the reader report what it retrieved. Ask every reader for a note describing what it actually read: bytes, note count, what failed. That one field caught three experiment-invalidating faults that no amount of after-the-fact auditing had surfaced.

Why a format and not a feature

Skills worked because they were a convention anyone could adopt without buying anything. AGENTS is the same bet applied to handoff. Your session should be portable. Your positions should outlive your transcript. The agent picking up your work should know which of your ideas you already killed.

We are publishing the format, the 100-question evaluation set, and every judge file, including the runs where our own predictions were wrong. We pre-registered five predictions across these experiments and got three of them wrong. Each wrong one changed what we built next.

If you export a session this way, tell us what breaks. Especially if it is Tier E. That one is still open.