We designed the harness around two goals:

  1. Throughput: the agent keeps up as work arrives.
  2. Reactivity: the agent responds to new information promptly instead of waiting for all ongoing work to finish.

Those goals conflict when new information arrives during ongoing work. Interrupting every running model call would hurt throughput. Waiting for all current work to finish would hurt reactivity. The harness accepts each arrival into an ordered queue without changing the call already in progress. At the next breakpoint, prepareTurn() combines the saved conversation with the oldest queued events that fit. Together, they form the next input. The loop can also adopt a shorter version of old history at the same breakpoint.

New work waits in order

Each arrival is an event: a message, a system change or a result from background work. It is added to an ordered queue before it can affect the conversation. In the diagram, A arrived before B, B before C and C before D.

Four colored event cards labeled A through D entering an ordered queue in the same order.

At the breakpoint, prepareTurn() reads the conversation so far and everything waiting in the queue. If A, B and C fit in the next model call, it prepares them together with the existing history. D stays in the queue. An older event is never skipped to reach a newer one.

The current history and event cards A through D passing through prepareTurn. The history and A through C form the next model call while D waits.

At this point A, B and C are still queued. A successful call records the input and response and advances the queue to D as one change.

The prepared history and event cards A through C becoming turn 43 after a successful call, while D becomes the head of the pending queue.

If this step or the model call fails, the queue remains A, B, C, D.

Taking several waiting events in one call helps the agent keep up. Returning to the queue at each breakpoint lets it react to what arrived during the previous call, without waiting for the rest of the task to finish.

Older history is compacted while the conversation continues

Turns 1–4 eventually stop fitting comfortably. A background job starts shortening them. The conversation keeps moving: turns 5 and 6 can be added while that work runs, and new events can continue to enter the queue.

Turns 1 through 4 from epoch zero flowing through background compaction into a shorter summary while turns 5 and 6 are added later.

The ordinary, unshortened history is already epoch 0. When the summary is ready, it waits for the current model call to finish. At the next breakpoint, the conversation adopts epoch 1. prepareTurn() can then build the next call from S1 followed by the unchanged turns 5 and 6.

Epoch zero containing turns 1 through 6 above epoch one containing summary S1 followed by turns 5 and 6. S1 points back to turns 1 through 4, while turns 5 and 6 carry forward.

Later, turns 7 and 8 arrive. If the history needs shortening again, S2 can replace epoch 1 through turn 6. Epoch 2 starts with S2 and carries turns 7 and 8 forward. S2 points to epoch 1, where S1 still points to epoch 0.

Three conversation epochs. Epoch zero contains the original turns. Epoch one contains S1, turns 5 and 6, then turns 7 and 8. Epoch two contains S2 followed by turns 7 and 8. S1 points to epoch zero and S2 points to epoch one.

The original turns remain stored. Following the links through later epochs leads back to the original conversation. Over several rounds, those links form a tree similar in shape to a Merkle tree, using ranges of turns rather than content hashes.

Throughput and reactivity

The result is an agent that can keep up with a busy conversation and react before a long task has finished. At each breakpoint, it takes the oldest waiting events that fit. Old history can be shortened in the background without stopping new turns. Each model call sees a stable input, while the conversation around it keeps moving.