Keyboard shortcuts

Press or to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Agent runtime

Wisp separates the stateful agent harness from the provider-neutral agent loop. The split keeps conversation policy and durable session concerns out of the model/tool cycle while giving every interface the same runtime behavior.

flowchart TD
  Interface["CLI, TUI, JSONL-RPC, or SDK"] --> Host["RPC command host"]
  Host --> Session["CodingSession"]
  Session --> Harness["AgentHarness"]
  Harness --> Loop["run_agent_loop"]
  Loop --> Provider["Provider adapter"]
  Loop --> Tools["Tool executor"]

Ownership at a glance

LayerOwnsDoes not own
run_agent_loopOne invocation’s turns, provider streaming, context estimates, tool batches, and continuation stateThe conversation between invocations, persistence, compaction policy, or frontend behavior
AgentHarnessThe in-memory transcript, steering and follow-up queues, cancellation, and coordination around one live invocationDurable storage, trust and safety policy, or provider-specific protocol behavior
CodingSessionPersistence, compaction orchestration, project context, trust, safety policy, and cost accountingProvider/tool control flow or frontend rendering
RPC command host and interfacesCommand scheduling, transport adaptation, and rendering typed eventsIndependent copies of agent policy

The practical rule is to change behavior in the narrowest layer that owns it. Provider-native request and replay behavior stays in the provider adapter even when the loop consumes it.

The agent loop

run_agent_loop is an async event stream. It receives a portable base history and an AgentLoopConfig, then repeats a provider/tool cycle until cancellation, failure, a limit, or a request-boundary decision stops it.

flowchart TD
  Start["Start a turn"] --> Estimate["Estimate context"]
  Estimate --> Stream["Stream one model response"]
  Stream --> Outcome{"Response outcome"}
  Outcome -->|Failure| Terminal["Emit terminal turn events"]
  Outcome -->|Tool calls| Execute["Execute the requested tool batch"]
  Execute --> Boundary["Consult the request boundary"]
  Outcome -->|No tool calls| Boundary
  Boundary --> Decision{"Boundary decision"}
  Decision -->|Stop| End["End the invocation"]
  Decision -->|Continue| Start
  Decision -->|"Replace or rebase"| Start

One turn is one provider response plus its requested tool batch, if any. TurnCompleted closes that turn; it does not necessarily end the loop invocation. The loop keeps only transient continuation state:

  • the provider’s native response cursor, when supported;
  • tool results and extra user messages waiting for the next request;
  • assistant and tool rows produced during the current invocation;
  • turn and tool-iteration counters.

The input messages sequence is not mutated. Persistence and the transcript used by a later invocation remain the caller’s responsibility.

Request boundaries

After a successful turn, the loop asks an optional request-boundary hook what the next request should use:

DecisionEffect
stopEnd the invocation. This takes precedence over other fields.
extra_messagesAdd plain user messages to the next continued request.
messagesUse a fresh portable history and discard native continuation state.
context_rebaseReplace the portable base while retaining the accepted native continuation tail.

Context-overflow recovery uses the same decision model through a separate hook. Decisions are validated centrally so unsupported structured-history transitions are rejected instead of being silently flattened for a provider. A hook that declines recovery can return ContextOverflowFailure(message) to supply the run’s error text; either way, the loop publishes the error and completes the rejected turn itself.

The harness

AgentHarness persists an in-memory conversation across calls to prompt(), prompt_message(), and continue_(). It constructs a loop configuration for each live invocation and projects completed assistant messages and tool results into its transcript as events arrive.

sequenceDiagram
  participant Session as CodingSession
  participant Harness as AgentHarness
  participant Loop as run_agent_loop

  Harness->>Loop: Start with normalized provider history
  Loop-->>Harness: Stream message and tool events
  Harness->>Harness: Append completed transcript rows
  Loop-->>Harness: TurnCompleted
  Harness->>Harness: Drain an eligible queue and arm the boundary
  Loop->>Harness: Request the next-boundary decision
  Harness->>Session: Ask for session policy when configured
  Session-->>Harness: Stop, continue, replace, or rebase
  Harness-->>Loop: Return the decision
  Loop-->>Harness: Start the next turn
  Harness->>Harness: Apply any accepted transcript transition

The boundary handshake deliberately separates two kinds of state:

  • The loop applies a decision to the request it is preparing.
  • The harness applies the corresponding transcript replacement when the next turn starts.

This keeps the provider-visible request and harness-visible transcript synchronized without letting the loop mutate conversation state owned by its caller.

Queued messages

The harness has two FIFO queues with a shared capacity:

  • Steering is considered after every completed turn and has priority over follow-up messages.
  • Follow-up is considered only when a turn has no tool calls and the run would otherwise stop.

Each queue can inject one message at a time or its current snapshot as a batch. Messages added after a batch is selected wait for a later boundary. Cancellation and stream closure do not inject queued messages that were never exposed through queue events.

The harness permits one live invocation. While it is running, callers use steering or follow-up rather than starting an overlapping prompt.

Failure recovery and message ownership

Invocation offsets are validated before transcript mutation. Once a valid prompt is accepted, runtime failures do not silently remove it or completed tool results. Startup and stream-cleanup errors still release the running flag and cancellation handles, allowing a later invocation.

The harness owns detached copies of incoming messages and exposes detached transcript and queue snapshots. Frozen message models can still contain mutable nested JSON, so copying at these boundaries prevents callers, completion-event observers, and session callbacks from modifying retained messages. Internal queue draining keeps entry identity; accounting does not build public snapshots. Transcript transitions remain deferred until the next accepted turn.

wisp.agent.loop

ModuleResponsibility
runner.pyTop-level turn lifecycle, context estimation, and provider/tool orchestration
model_response.pyProvider event translation, response outcomes, deltas, and retries
provider_request.pyOptional capabilities, request invocation, overflow normalization, and stream cleanup
provider_lifecycle.pyProvider response start, retry, tool-call, and terminal validation
response_projection.pyUsage, cost, context observation, and completed-message projection
tool_execution.pyToolBatch facade, sequential and truncated execution, and cancellation settlement
tool_lifecycle.pyShared tool event contracts, executor lifecycle validation, and the per-batch settlement record
prepared_tools.pyTwo-phase preparation and bounded scheduling
continuation.pyNative cursor, pending request data, and request-boundary transitions
stream_cleanup.pyOwned iterator close and cleanup exception precedence
config.pyProvider-neutral dependencies, limits, hooks, and cancellation contracts

Start with run_agent_loop in runner.py, then follow only the phase you need.

wisp.agent.harness

ModuleResponsibility
runner.pyTranscript updates, queues, cancellation, and loop-event orchestration
boundaries.pySession-policy preparation and synchronized transcript transitions
config.pyHarness dependencies, runtime limits, and queue policy

Start with AgentHarness._run in runner.py for the complete orchestration path.

The source directories also contain contributor-focused loop and harness READMEs with local change and test maps.

Contracts to preserve

Runtime event order is observable through the SDK, RPC, persistence, and frontends. In particular:

  • every started turn has exactly one terminal TurnCompleted, across the loop, harness, and session (all three track turns with wisp.agent.turn_lifecycle.TurnLifecycle);
  • completed tool executions emit ToolExecutionEnded immediately before the matching ToolResultReady;
  • approval events are ordered request, resolution, then terminal result;
  • steering drains before follow-up, with FIFO order within each queue;
  • interrupted tool exchanges are repaired before the next provider request;
  • optional provider capabilities are detected rather than assumed.

Focused assertions for these contracts live in tests/agent_runtime.py and tests/agent/test_agent_runtime_invariants.py.