Spec: embedding agents-exe in a web server
On Sun, 04 Oct 2026, by @lucasdicioccio, 6262 words, 12 code snippets, 5 links, 0images.
Generated from todos/web-server-embedding.md, the repository is the canonical source and may be ahead of this page.
Spec: embedding agents-exe in a web server
Status: proposed (2026-09-18)
Goal
Run agents from a long-lived HTTP server, with no TUI and no CLI, where:
- every session lives in a database, not in
conv.<uuid>.jsonfiles; - a client can start a session, send follow-up messages, watch progress live, list and complete deferred tool calls, resume, and cancel;
- the server can restart at any time without losing or corrupting sessions.
The deliverable is:
- a small library API in
agents-libfor hosting agents, with no HTTP dependencies; - a reference HTTP server,
agents-server, built on that API.
Non-goals (for this spec)
- Multi-tenancy: per-user isolation, per-user API keys, and sandboxing of bash
or MCP tools. The data model reserves an
ownercolumn so this can come later without a migration (see Milestone 2). - A Postgres backend. The interfaces below are designed so it can be added as a separate backend.
- Streaming LLM tokens. Clients get updates at step granularity.
- Agent definitions stored in the database. Agents still come from JSON files loaded at startup.
- Splitting the library to drop the
brick/vtydependencies.
Current state
What we can reuse as-is
| Piece | Where | Notes |
|---|---|---|
| Loading the agent tree once, sharing it across concurrent runs | AgentTree.withAgentTree; used by MCP/Server.hs:127 and :222 | The MCP server already runs one async per request over a shared tree. |
| Stepping a session | Session.Step.runStepM, Session.Loop.runUntilBlocked (Loop.hs:105), Session.Wake.resumeSession | runUntilBlocked returns the session when it waits only on deferred calls. |
| Injecting external results | Session.Wake.wakeSession | Pure load, modify, return. |
| Session storage interface | SessionStore.SessionBackend (SessionStore.hs:97) | store / load / list / delete; file, SQLite, and composite versions. |
| Continuation index | Session.Async.ContinuationStore (Async.hs:248), SQLite table tool_continuations | Written by the scheduler (Step.hs:490, :728). |
| Declarative policy per agent | toolCallPolicyConfig, executionMode, etc.; applied by Session.AgentConfig.applyAgentDurableConfig | |
| Progress hook | Combinators.StoreSessionProgress.agentWithSessionProgress | Fires SessionUpdated before every step. |
| Restart handling for background calls | Step.pollRunningCall (Step.hs:505) | A Running call whose process is gone becomes Failed ("orphaned"). |
| Background execution | Session.Base.withAsyncEngine; on-demand World at Step.hs:917 | Needs an OS World plus an AsyncEngine. |
Gaps
G1. A backend set on an agent after construction is ignored for progress
storage. OneShot.nodeToAgent wraps the agent with agentStoreSession
while ctxSessionBackend = Nothing. agentStoreSession
(StoreSessionProgress.hs:130) checks ctxSessionBackend on the agent it
receives when it wraps, not at store time. So
withSessionBackend backend =<< nodeToAgent … still writes every step to the
file store.
G2. Agent construction is copied three times with small differences:
OneShot.nodeToAgentWithThinking (OneShot.hs:293),
AgentTree.OneShotTool.nodeToAgent (OneShotTool.hs:511, used for
sub-agents), and the MCP server (MCP/Server.hs:~300). The one-shot copy also
prints thinking to stdout/stderr. Progressive-disclosure tool filtering
(agentEvaluateActiveTools) is applied only on the one-shot path.
G3. Sub-agent sessions always go to the file store. OneShotTool.nodeToAgent
takes the file SessionStore, generates a fresh ConversationId, and records
no link to the parent session.
G4. Session-reading tools and search read only the file store (tools
fixed in Phase 3; the search index is CLI-only and stays on files). The
SystemToolbox session tools (Tools/SystemToolbox/Session.hs, via
SessionIntrospectionConfig.introspectionStore) and the search index
(Session/Search/Index.hs, via indexSessionStore) call
SessionStore.listSessions / readSession directly. Props.sessionStore
passes this down from AgentTree.
G5. Sessions have two keys. The file store is keyed by ConversationId,
durable commands by SessionId, and the two are generated separately
(SessionDurable.hs handleStart). runOneShotWithConfig stores a paused
session under both.
G6. No concurrency control. Every mutation is load, modify, store. Two
concurrent requests on one session (two complete calls, or complete during
a run) silently lose an update.
G7. Continuation table and sessions drift apart. wakeSession never calls
csComplete, so tool_continuations rows stay pending forever. Finding a
session from a token scans every session (findSessionForToken in
SessionDurable.hs) even though tool_continuations.session_id already maps
it.
G8. Background calls only live as long as one loop call. The loops
create an engine on demand and shut it down on exit (withEngineShutdown,
Loop.hs:88). A server that pauses and resumes over several HTTP requests
needs a World and engine per live session that outlive a single request.
G9. No metadata beyond the JSON blob. The sessions table has only
session_id, created_at, updated_at, json. Listing sessions by agent or
status, or finding ones interrupted by a restart, means decoding every blob.
G10. The prompt-to-first-turn code lives in the CLI. handleStart builds
the initial UserTurn inside CLI/SessionDurable.hs, so the library has no
reusable version.
Design
Overview
+-------------------------------------------+
HTTP / SSE --> | agents-server (wai + warp) | examples/agents-server/
+---------------------+---------------------+
|
+---------------------v---------------------+
| System.Agents.Host.Runner | per-session lock, live runs,
| SessionRunner | events, cancellation, recovery
+---------+-------------------+-------------+
| |
+------------------v-----+ +-------v----------------------+
| System.Agents.Host | | SessionBackend (extended) |
| Host, on top of | | ContinuationStore |
| AgentFactory | | SQLite, one database file |
+------------------------+ +------------------------------+
The library adds three modules:
System.Agents.AgentFactory: the single agent factory (done).System.Agents.Host: loaded agents, database, and stores.System.Agents.Host.Runner: session lifecycle on top of aHost.
None of them imports anything HTTP-related. The HTTP layer is a new executable.
1. Keys: one id per session
Rule: a session’s SessionId is its only key. Wherever a
ConversationId is needed for the same session, it is
sessionIdToConversationId sid. The Host API never generates a separate
ConversationId for a session.
Sub-agent sessions get their own SessionId and record their parent (§3).
2. The agent factory and the Host
2.1 System.Agents.AgentFactory (implemented, Phases 1–2)
data AgentDeps = AgentDeps
{ adApiKeys :: LoadedApiKeys
, adSessionSink :: SessionSink
, adContinuationStore :: Maybe ContinuationStore
, adToolCache :: Maybe ToolCache
, adCompletion :: Maybe (OSAgentNode -> Completion)
-- ^ Replaces the LLM call (tests pass a mock).
}
defaultAgentDeps :: LoadedApiKeys -> AgentDeps -- SinkNone
fileAgentDeps :: SessionStore -> LoadedApiKeys -> AgentDeps -- SinkFiles
data AgentRole
= RootAgent
| SubAgent { subParentConversation :: ConversationId, subCallStack :: [CallStackEntry] }
-- | The only way to turn an OSAgentNode into a runnable Agent.
buildAgent :: Tracer IO Trace -> AgentDeps -> AgentRole -> ConversationId -> OSAgentNode
-> IO (Agent (LlmTurnContent, Session))buildAgent takes a ConversationId, not a SessionId: the CLI and TUI key
their file store by conversation ID, and changing that is not part of this
work. Callers that follow the §1 rule pass sessionIdToConversationId sid.
What buildAgent does, in this order:
-
Builds the base
Agentrecord: prompt, tools,toolCall,toolPortal, andcomplete(fromadCompletion, or else the agent’s OpenAI config). -
Applies
applyAgentDurableConfigfor the node’s JSON config, and installsadContinuationStore,adToolCache, and (forSinkBackend) the backend. -
Applies
agentEvaluateActiveTools(progressive disclosure) to every agent. -
Wraps the agent with
agentPersistSession adSessionSink convIdlast (fixes G1).SessionSinklives inCombinators.StoreSessionProgress:data SessionSink = SinkBackend SessionBackend | SinkFiles SessionStore | SinkNone agentPersistSession :: SessionSink -> ConversationId -> Agent r -> Agent ragentStoreSessionkeeps its old behaviour for existing callers. -
Never writes to stdout or stderr.
OneShot.nodeToAgentWithThinkingadds the thinking printer, the media injection, and the extra--session-filecopy as decorators on the result.
All front-ends use it: one-shot run and the TUI (through
OneShot.nodeToAgent), sub-agent tools, the MCP server, and the durable
session commands.
Sub-agents: turnAgentRuntimeIntoIOTool takes an AgentDeps instead of a
SessionStore and API keys. It builds the sub-agent with
SubAgent parentConvId stack and the same conversation ID it uses for the
call stack and the OS World entity. That ID now also names the stored
sub-session; it used to be an unrelated random ID. The parent link is kept at
runtime (ctxParentConversation), and backends store it as
parent_session_id (Phase 3).
2.2 System.Agents.Host (implemented, Phase 5)
data Host = Host
{ hostAgents :: Map Text OSAgentNode -- root agents, by slug
, hostDeps :: AgentDeps -- root agents: SinkNone, the runner stores
, hostSubAgentDeps :: AgentDeps -- sub-agents: SinkBackend hostBackend
, hostBackend :: SessionBackend
, hostContinuations :: ContinuationStore
, hostTracer :: Tracer IO HostTrace
, hostLiveSessionTtl :: NominalDiffTime
}
data HostConfig = HostConfig
{ hcAgentFiles :: [FilePath], hcApiKeysFile :: FilePath, hcDatabasePath :: FilePath
, hcCompletion :: Maybe (OSAgentNode -> Completion), hcLiveSessionTtl :: NominalDiffTime }
defaultHostConfig :: [FilePath] -> FilePath -> FilePath -> HostConfig -- TTL 15 minutes
withHost :: HostConfig -> Tracer IO HostTrace -> (Host -> IO a) -> IO awithHost opens the database with journal_mode = WAL and
busy_timeout = 5000, runs the session and continuation migrations, loads
each agent file (sub-agent tools get hostSubAgentDeps, and session tools a
backendCatalog), and refuses duplicate root slugs (HostError). Both
dependency sets share the continuation store and the completion override.
The bundled SQLite is built with THREADSAFE=1 (serialized), so one
connection is shared by all threads; no statement sequence relies on
changes().
3. Storage
3.1 Extended SessionBackend (implemented, Phase 3)
SessionStatus, sessionStatusOf, isBlockedOnDeferredCalls, and
hasBackgroundCalls live in Session.Types (pure, re-exported by
Session.Base); the metadata types live in SessionStore.
data SessionStatus = StatusIdle | StatusReady | StatusRunning | StatusWaitingExternal | StatusFailed
data SessionLabels = SessionLabels -- Nothing keeps the stored value
{ slAgent :: Maybe Text, slParent :: Maybe SessionId, slOwner :: Maybe Text }
data SessionMeta = SessionMeta
{ smSessionId :: SessionId
, smAgent :: Maybe Text
, smParent :: Maybe SessionId
, smOwner :: Maybe Text -- reserved, unused for now
, smStatus :: SessionStatus
, smStatusDetail :: Maybe Text -- why it failed
, smVersion :: Int -- incremented on every write; 0 = never stored
, smCreatedAt, smUpdatedAt :: UTCTime
}
data SessionQuery = SessionQuery
{ sqAgent :: Maybe Text, sqStatuses :: Maybe [SessionStatus], sqParent :: Maybe SessionId
, sqUpdatedBefore :: Maybe UTCTime, sqLimit :: Maybe Int }
data VersionConflict = VersionConflict { vcSessionId :: SessionId, vcExpected, vcActual :: Int }
data SessionBackend = SessionBackend
{ sbStore :: SessionId -> Session -> IO () -- unconditional
, sbLoad :: SessionId -> IO (Maybe Session)
, sbList :: IO [(SessionId, UTCTime)]
, sbDelete :: SessionId -> IO ()
, sbStoreLabelled :: SessionLabels -> SessionId -> Session -> IO ()
, sbLoadMeta :: SessionId -> IO (Maybe (Session, SessionMeta))
, sbCompareAndStore :: SessionMeta -> Session -> IO (Either VersionConflict SessionMeta)
, sbQuery :: SessionQuery -> IO [SessionMeta]
}Unconditional stores (sbStore, sbStoreLabelled) increment the version,
keep stored labels where the new ones are Nothing, and set the status from
sessionStatusOf, unless the stored status is StatusRunning.
sbCompareAndStore writes the given metadata exactly, only if the stored
version equals smVersion (a missing session counts as 0).
Backend implementations:
- SQLite: compare-and-store is one conditional statement
(
UPDATE … WHERE version = ? RETURNING …, or an upsert withDO UPDATE … WHERE sessions.version = 0for version 0), so it is atomic without a lock. Rows from before the migration have version 0. - File: metadata in a sidecar
meta.<uuid>.json(notconv.*, so session listings ignore it). A session file without a sidecar reads as version 0. Compare-and-store is read-compare-write, not atomic across processes. - Composite: writes and queries go to the primary backend; loads fall back in order.
AgentFactory stores through sbStoreLabelled with the agent’s slug and,
for sub-agents, the parent session. A sub-agent’s session ID is its
conversation ID, so its own sub-agents name it correctly as their parent.
3.2 SQLite schema and migrations (implemented, Phase 3)
SessionStore.runMigrations conn component migrations applies each
not-yet-applied Migration in a transaction and records it in
schema_migrations(component, version, applied_at), so the continuation
store and tool cache can have their own migration lists later.
initializeSessionSchema (called by mkSqliteSessionStore) runs the
sessions migrations:
-- migration 1: the original table and index (CREATE IF NOT EXISTS)
-- migration 2
ALTER TABLE sessions ADD COLUMN agent_slug TEXT;
ALTER TABLE sessions ADD COLUMN parent_session_id TEXT;
ALTER TABLE sessions ADD COLUMN owner TEXT;
ALTER TABLE sessions ADD COLUMN status TEXT NOT NULL DEFAULT 'ready';
ALTER TABLE sessions ADD COLUMN status_detail TEXT;
ALTER TABLE sessions ADD COLUMN version INTEGER NOT NULL DEFAULT 0;
CREATE INDEX IF NOT EXISTS idx_sessions_status ON sessions(status, updated_at);
CREATE INDEX IF NOT EXISTS idx_sessions_parent ON sessions(parent_session_id);
-- then each existing row's status is derived from its JSONThe tool_continuations(session_id, completed_at) index belongs to the
continuation store’s migrations (Phase 4). Connection settings
(journal_mode=WAL, busy_timeout) are set by withHost
when it opens the database (Phase 5), not by the library backends, which
work on connections their callers own.
3.3 Reading sessions from tools (G4, implemented, Phase 3)
A read-only interface both stores implement:
data CatalogEntry = CatalogEntry
{ ceConversationId :: ConversationId, ceUpdatedAt :: Maybe UTCTime
, ceSession :: Maybe Session -- Nothing when unreadable (locked file)
, ceBusy :: Bool } -- locked file, or status running
data SessionCatalog = SessionCatalog
{ catList :: IO [CatalogEntry], catRead :: ConversationId -> IO (Maybe Session) }
fileCatalog :: SessionStore -> SessionCatalog
backendCatalog :: SessionBackend -> SessionCatalogEntries are keyed by conversation ID, the vocabulary of the session tools
(and the same UUID as the session ID in a backend).
SessionIntrospectionConfig.introspectionCatalog and
AgentTree.Props.sessionCatalog take a catalog; the CLI and TUI pass
fileCatalog store, and the Host will pass backendCatalog.
The search index (Session/Search) stays on the file SessionStore: only
the session-index and session-search CLI commands use it, and it is built
around file paths and modification times.
3.4 Continuations stay consistent (G7, implemented, Phase 4)
The session JSON is the source of truth. tool_continuations is an index
from token to session.
ContinuationStore.csFindSession :: ContinuationToken -> IO (Maybe SessionId)answers for pending and completed tokens (csLoadonly returns pending ones, so it cannot tell a completed token from an unknown one).Session.Wake.wakeSessionWith :: Maybe ContinuationStore -> Maybe ToolCache -> Session -> [(ContinuationToken, UserToolResponse)] -> IO WakeOutcomereports, per token, whether it was applied, already completed, or unknown, and callscsCompletefor the applied ones. A token is already completed when a call of the session still carries it but is no longer deferred, or when the store knows it for this session (once a turn is complete, the session no longer holds its tokens).wakeSessionandwakeSessionWithCacheare wrappers returning the session.Session.Wake.findSessionForToken :: Maybe ContinuationStore -> SessionBackend -> ContinuationToken -> IO (Maybe SessionId)askscsFindSessionfirst and scans the backend’s sessions only for tokens the store does not know.csCompleteusesUPDATE … RETURNINGinstead ofSELECT changes(), which another user of the same connection could change in between.- The continuation table’s schema goes through
runMigrations(componentcontinuations); migration 2 adds the(session_id, completed_at)index.
The CLI complete command keeps its scan over the file store: CLI agents
have no continuation store. The runner (§4.3) uses findSessionForToken and
wakeSessionWith.
Agents store their session before each step, so the session a run stops in
is not stored by the agent: callers store it (the session commands and
one-shot run do; the runner stores after every step, §4.2).
4. System.Agents.Host.Runner: session lifecycle (implemented, Phase 5)
data SessionRunner -- opaque
newSessionRunner :: Host -> IO SessionRunner
shutdownSessionRunner :: SessionRunner -> IO () -- cancels active runs, stores their sessions
withSessionRunner :: Host -> (SessionRunner -> IO a) -> IO a
runnerStats :: SessionRunner -> IO RunnerStats -- live sessions, active runs
data RunMode = StepOnce | UntilBlocked
createSession :: SessionRunner -> Text -> NewMessage -> Maybe RunMode -> IO (Either RunnerError SessionMeta)
postMessage :: SessionRunner -> SessionId -> NewMessage -> Maybe RunMode -> IO (Either RunnerError SessionMeta)
resume :: SessionRunner -> SessionId -> RunMode -> IO (Either RunnerError SessionMeta)
completeCall :: SessionRunner -> ContinuationToken -> UserToolResponse -> Bool {- auto-resume -} -> IO (Either RunnerError SessionMeta)
cancelRun :: SessionRunner -> SessionId -> IO (Either RunnerError SessionMeta)
getSession :: SessionRunner -> SessionId -> IO (Maybe (Session, SessionMeta))
awaitRun :: SessionRunner -> SessionId -> NominalDiffTime -> IO (Either RunnerError (SessionMeta, Bool {- run still active -}))
deleteSession :: SessionRunner -> SessionId -> DeleteMode -> IO (Either RunnerError DeletionPlan)
subscribe :: SessionRunner -> SessionId -> IO (IO SessionEvent) -- blocking "next event"
recoverOnStartup :: SessionRunner -> IO [SessionId]
data NewMessage = NewMessage { nmText :: Text, nmMedia :: [MediaAttachment] }
data DeleteMode = DryRun | DeleteForReal
data DeletionPlan = DeletionPlan
{ dpSessions :: [SessionId] -- ^ the session and all descendants, deepest first
, dpContinuations :: Int -- ^ continuation rows removed (or that would be)
, dpDryRun :: Bool
}
data RunnerError
= UnknownAgent Text | UnknownSession SessionId | UnknownToken ContinuationToken
| TokenAlreadyCompleted ContinuationToken
| RunInProgress SessionId -- ^ a run owns the session (on delete: one of the tree, or an ancestor)
| NoActiveRun SessionId -- ^ cancel without a run (HTTP 409 no_active_run)
| NotAcceptingMessages SessionId SessionStatus
| Conflict VersionConflictAgents built by the runner always run asynchronously
(withExecutionMode Asynchronous), like the CLI session commands, so that
runs can stop and resume.
4.1 Per-session state
data LiveSession = LiveSession
{ lsLock :: MVar () -- serialises every change
, lsRun :: TVar (Maybe (Async ())) -- the active run, if any
, lsAgent :: TVar (Maybe (Agent …)) -- kept between runs (G8)
, lsLatest :: TVar (Maybe (Session, SessionMeta)) -- last version stored or loaded
, lsInbox :: TVar [(ContinuationToken, UserToolResponse)] -- see 4.3
, lsLastTouched :: TVar UTCTime
, lsEvicted :: TVar Bool
}
-- held in: TVar (Map SessionId LiveSession)- The runner creates a
LiveSessionon first access. A reaper thread (every half TTL, between 50 ms and 60 s) evicts sessions idle for longer thanhostLiveSessionTtlthat have no active run and noRunningcall, and shuts their engine down. An operation that took aLiveSessionjust before it was evicted noticeslsEvictedonce it holds the lock, and retries with a fresh one. - The first run of a session builds its agent with
buildAgentand keeps it inlsAgent.runStepMinstalls a World and an engine on demand and returns the agent holding them; the runner keeps that agent, so the engine survives between runs. The runner has its own loop overrunStepM, with the stop conditions ofrunUntilBlocked, and shuts the engine down only on cancel, eviction, deletion, or runner shutdown. Runningcalls whose engine was lost (restart or eviction) are resolved as orphaned bypollRunningCallon the next step. No new code is needed.
4.2 Mutation protocol
Every operation that changes a session:
- Takes
lsLock. - If a run is active and the operation is not
cancelRun, returnsRunInProgress.completeCallis the exception: see 4.3. sbLoadMeta, then applies the change (wakeSessionWith, or pushes aUserTurn), thensbCompareAndStore.- Emits a
SessionEvent, then releases the lock.
A run takes the lock for each step’s store, not for the whole run. Inside the
run loop each step: runStepM, then sbCompareAndStore under lsLock, then
emit session.updated. A VersionConflict inside a run stops it with
StatusFailed with detail “concurrent modification”. That cannot happen with a single
server process; it guards against a second process or the CLI writing to the
same database.
The runner sets status = running when a run starts and sessionStatusOf s
when it ends. On an exception it sets StatusFailed with displayException e as detail.
Progress storage installed by buildAgent (§2.5) is disabled for
runner-built agents (SinkNone),
because the runner does the versioned stores itself.
4.3 Completing a deferred call during a run
runUntilBlocked only stops once nothing can progress in-process, so an
active run and a pending deferred call can coexist. For example, a deferred
call waits while a background call is still running.
Writing the woken session while the run is in a step would make the run’s
next versioned store conflict. So completeCall on a session with an active
run, under the lock:
- checks the token against the run’s latest version (and the queue), and
answers
UnknownTokenorTokenAlreadyCompletedright away; - otherwise appends the result to
lsInboxand returns.
Before each step and before deciding to stop, the run, under the lock,
applies the queued results with wakeSessionWith (which also marks the
continuations completed) and stores the result. A run that was about to
stop on deferred calls therefore carries on when their results are queued.
cancelRun applies the queue too.
Without an active run, completeCall applies the result and stores it at
once, and with auto-resume starts an UntilBlocked run when the session is
ready. Two concurrent completions of one turn are serialised by the lock:
the first stores, the second sees the first’s version (tested).
4.4 Follow-up messages
postMessage is valid only when the status is StatusIdle (head is a final
LlmTurn). It pushes
UserTurn (UserTurnContent sysPrompt sysTools (Just (UserQuery text media)) []) Nothingonto turns (newest first, as handleStart does), taking sysPrompt and
sysTools from the live agent. naiveTilNoToolCallStep then sends it to the
LLM on the next step (Step.hs:1081). Agents built by the runner keep
usrQuery = pure Nothing.
The code that builds the first turn from a prompt moves out of
CLI/SessionDurable.handleStart into
Session.Types.newSessionFromPrompt :: SystemPrompt -> [SystemTool] -> NewMessage -> IO Session
(fixes G10). createSession and handleStart both use it.
4.5 Events
data SessionEvent
= RunStarted SessionId RunMode
| SessionUpdated SessionId SessionMeta Turn -- ^ the new head turn
| CallsDeferred SessionId [DeferredCallView]
| RunStopped SessionId SessionStatus
| SessionFailed SessionId TextDeferredCallView is the same data agents-exe session pending prints: tool
name, call id, token, disposition, and the call with its arguments. It lives
in Session.Types with pendingDeferredCalls :: Session -> [DeferredCallView];
the CLI’s extractDeferredCalls wraps it.
Every stored version emits SessionUpdated (with the head turn), including
the stores at the start and end of a run. Events go through one runner-wide
broadcast channel: subscribe returns an action yielding the next event of
one session, so a subscriber keeps its stream when the session is evicted
and loaded again. subscribeSTM gives the same stream as a transaction that
consumes one event of any session and yields Nothing for another session’s,
to combine with timers or flags (the HTTP events stream uses it). Filtering
must consume other sessions’ events one transaction at a time: a retry
after readTChan rolls the read back, which in the first version of
subscribe blocked a subscriber for good on another session’s event. Each
event is also traced as HostRunnerTrace kind sid.
4.6 Startup recovery
recoverOnStartup queries status = running. For each session found that
this runner is not running itself:
- Load it.
- Set its status to
sessionStatusOf s. - Emit nothing, since no subscribers exist yet.
It does not resume automatically. Rows left running mean the process
died mid-step. The partially started step is lost, and any Running calls
become orphaned on the next resume. The server logs how many sessions it
recovered. Automatic resume can be added later as --resume-interrupted.
4.7 Cancellation
cancelRun cancels the run’s Async without holding the lock (the run takes
it to store its steps), then, under the lock and only if that run still
owns the session:
- shuts the engine down, which marks background calls cancelled in the World, and keeps the agent (and World) with no engine, so the next run gets a fresh engine on the same World;
- applies queued external results;
- refreshes the head partial turn from the World, so its cancelled calls
become
Failed "async tool call was cancelled"; - stores the session with the status its turns imply, and emits
RunStopped.
Background calls below the head (the LLM already got a placeholder for them)
stay Running in the stored session. The next run’s late-result collection
finds them cancelled in the kept World and tells the LLM, in a user message,
as for any late result (tested). If the session is evicted first, they are
reported as orphaned instead. Sub-agent runs execute inside the parent’s
tool call, so cancelling the parent cancels them.
4.8 Waiting for a run
awaitRun blocks until the session’s active run stops or the timeout
expires, whichever comes first. It returns the current meta and whether a run
is still active. With no active run it returns at once. It watches
lsRun (STM waitCatch on the Async, raced against a timer), so waiting
never holds lsLock. A waiter going away (for example an HTTP client
disconnecting) never cancels the run.
4.9 Deleting sessions
Deletion always cascades: a session goes together with every descendant
(sub-agent sessions, found recursively through sbQuery on parent) and
their continuation rows. deleteSession:
- Collects the tree and builds a
DeletionPlan, deepest sessions first. - Returns
RunInProgressif any session in the tree has an active run. This applies in both modes, so a dry run shows whether the real delete would succeed. - With
DryRun, returns the plan and changes nothing. - With
DeleteForReal, walks the plan in order. For each session it takes that session’slsLock, deletes its continuation rows, callssbDelete, drops itsLiveSession(shutting down its engine), and releases the lock.
Deleting leaves first means a crash part-way through never leaves a
sub-session whose parent is gone; running the delete again finishes the job.
This needs two new ContinuationStore fields:
csCountSession :: SessionId -> IO Int for dry runs, and
csDeleteSession :: SessionId -> IO Int, which deletes every row of a session
(pending or completed) and returns the count. Deletion is also refused while
an ancestor of the session has an active run: its sub-agent calls write
into the tree, and would recreate what was deleted. The runner does the cascade,
not SQL foreign keys, so the file backend behaves the same.
5. HTTP API (agents-server) (implemented, Phase 6)
The server lives in examples/agents-server/, on wai and warp. The
application code is a private sub-library, agents-server-internal
(AgentsServer.Api, AgentsServer.Log, AgentsServer.Server), shared by the
agents-server executable and its agents-server-tests suite; agents-lib
does not depend on wai or warp. SSE is written by hand with responseStream
(no wai-extra). All bodies are JSON. Session objects embed the raw
Session JSON that is already stored, so clients can reuse session-print
logic. The user guide is documentation/agents-server.md.
agents-server --agent-file a.json [--agent-file b.json …] --api-keys keys.json \
[--db ./agents-server.db] [--port 8080] [--bind 127.0.0.1] \
[--live-session-ttl 900] [--shutdown-grace 10]
The server binds to 127.0.0.1 by default. There is no authentication in
milestone 1 (see non-goals); the server.started log line says so.
| Method & path | Body | Success | Errors |
|---|---|---|---|
GET /v1/agents | 200 [{slug, description, tools:[name]}] | ||
POST /v1/sessions?wait=&timeout= | {agent, prompt, media?:[{mime, base64, filename?}], run?: "none"\|"step"\|"until_blocked"} (default until_blocked) | 201 SessionView, Location header | 404 unknown_agent, 400 bad_request |
GET /v1/sessions?agent=&status=&parent=&limit=&before= | 200 {sessions:[SessionMetaView], next_before} | 400 | |
GET /v1/sessions/:id | 200 SessionView | 404 unknown_session | |
POST /v1/sessions/:id/messages?wait=&timeout= | {prompt, media?, run?} | 202/200 SessionView | 404, 409 run_in_progress, 409 not_accepting_messages |
POST /v1/sessions/:id/resume?wait=&timeout= | {mode?: "step"\|"until_blocked"} or no body | 202/200 SessionView | 404, 409 run_in_progress |
POST /v1/sessions/:id/cancel | 200 SessionMetaView | 404, 409 no_active_run | |
GET /v1/sessions/:id/pending | 200 {calls:[DeferredCallView]} | 404 | |
POST /v1/continuations/:token?wait=&timeout= | {result: UserToolResponse \| string, resume?: bool} (default true) | 202/200 SessionView | 404 unknown_token, 409 token_already_completed, 409 conflict |
GET /v1/sessions/:id/events | 200 text/event-stream | 404 | |
DELETE /v1/sessions/:id?dry_run= | 200 {sessions:[id], continuations, dry_run} | 404, 409 run_in_progress (also for a dry run, when the real delete would fail) | |
GET /healthz | 200 {ok:true, live_sessions, active_runs} |
Other errors: 404 not_found (unknown path), 405 method_not_allowed, 413
payload_too_large (bodies over 32 MiB), 500 internal_error. A malformed
session id answers 404 unknown_session, a malformed token 404
unknown_token.
SessionMetaView is the SessionMeta JSON: {session_id, agent, parent_session_id, owner, status, status_detail, version, created_at, updated_at} (owner is always null in milestone 1).
SessionView = SessionMetaView + {session: <Session JSON>, pending: [DeferredCallView]}.
Errors are {error: "<code>", message: "<text>"}.
Listing. Newest first by updated_at. status takes a comma-separated
list; limit is 1–500, default 50. When a page is full, next_before is the
updated_at of its last session, used as before= (strictly older) for the
next page. A session with exactly the boundary’s updated_at would be
skipped; timestamps have nanosecond precision, so that takes two writes in
the same instant.
Waiting. Every endpoint that can start a run (create, messages, resume, continuations) takes the same two query parameters:
wait=true|false, defaultfalse(?waitalone means true). Withfalse, the server answers as soon as the run has started. Withtrue, it answers when the run stops (idle, blocked on deferred calls, or failed), usingawaitRun.timeout=<seconds>, only used withwait=true. Default 120, maximum 600; larger values are clamped. If it expires, the server answers with the current state (status: "running") and the run carries on. Clients then follow it through the events stream or by polling.
The response body is always the SessionView at the time of answering, so a
waiting client gets the final turn and any pending deferred calls in one
round trip. A client disconnecting while waiting does not cancel the run.
Creation always answers 201. The other endpoints answer 202 when the
stored status is running at the time of answering (a run is still going,
including a continuation queued into an active run) and 200 otherwise.
On shutdown, waiting requests answer at once with the current state.
Deleting. dry_run=true returns the same body as a real delete,
listing every session and the number of continuation rows that would be
removed, and changes nothing.
SSE stream. Each SessionEvent becomes event: <kind> plus one
data: <json> line, where kind is run.started ({session_id, mode}),
session.updated (SessionMetaView + head_turn), calls.deferred
({session_id, calls}), run.stopped ({session_id, status}), or
session.failed ({session_id, message}). On connect the server first sends
event: snapshot with the SessionMetaView, so clients never need a
separate GET to sync; it subscribes before loading the snapshot, so no event
falls in between. A run’s session.updated for the running version comes
just before its run.started. The server sends a : keepalive comment
after 15 s without events (warp pauses its idle timeout while the handler
runs, so quiet streams and long waits are not cut). Streams end on shutdown.
Events are not replayed on reconnect. The stream reads events with
subscribeSTM, combined with the keepalive timer and the shutdown flag in
one transaction, so a timer firing never drops an event.
Shutdown. On SIGTERM or SIGINT: stop accepting connections, end event
streams and release waiting requests, give open requests
--shutdown-grace seconds, then shutdownSessionRunner (cancels active runs,
storing their sessions) and close the database.
6. Tracing and logs (implemented, Phase 6)
HostTrace wraps the existing traces (OneShot.Trace, tool registration,
OpenAI) plus runner events. The server prints them as JSON lines on stderr,
one object per line with ts, kind, and session_id when known.
AgentsServer.Log summarises each trace field by field instead of showing
it: LLM traces give byte and token counts, the HTTP client trace gives
method, host, path, and status only (its request carries the API key in a
header), and tool and agent-tree traces give their constructor name. Prompts,
payloads, headers, and API keys are never printed; a smoke test with a fake
key checked the log. Each HTTP request logs method, path, status, and time
to first byte.
Phases
Each phase leaves cabal build all and cabal test agents-tests passing with
-Wall -Werror, and gets its own entry in a todos/web-server-embedding.progress.md
tracker.
Phase 1: storage wiring fix (G1, G5) ✅
SessionSinkandagentPersistSession; keepagentStoreSessionas a wrapper.session startderives its conversation ID from the session ID.- Tests: an agent built with a SQLite backend and a mock LLM runs to
completion and writes no files under a temp session directory; the row
is keyed by the session’s
SessionId.
Phase 2: single agent factory (G2, G3, G10) ✅
System.Agents.AgentFactorywithAgentDeps,buildAgent, andAgentRole; the three copies moved onto it; sub-agents get their parent and call stack fromSubAgent.newSessionFromPromptin the library.- Tests: the existing suite unchanged; new
AgentFactoryTests.
Phase 3: metadata, versions, migrations, catalog (G4, G9) ✅
- Extended
SessionBackend,SessionLabels,SessionMeta,sessionStatusOf, SQLite migrations, and the file sidecar. SessionCatalog; the session tools andPropsswitched to it.- Sub-agents store their agent slug and
parent_session_id. - Tests (
SessionMetadataTests): status derivation; version increments; labels kept; CAS success and conflicts (stale, and version 0 twice); running status kept by unconditional stores;sbQueryfilters and limit; composite fallback; migrating a database from before metadata; migrations run once; file sidecar and legacy files; file and backend catalogs; list-sessions over a backend catalog; a parent agent calling a sub-agent tool (mock LLM) leaves a sub-session row naming the parent.
Phase 4: continuation consistency (G7) ✅
csFindSession,wakeSessionWith/WakeOutcome,findSessionForToken, continuation migrations,RETURNINGincsComplete.- Tests (
ContinuationConsistencyTests, on a real paused run: async agent, defer policy, mock LLM, one SQLite database): a wake marks the continuation completed; a second wake reports the token as already completed and leaves the session unchanged; an unknown token is reported; the woken session resumes to a final answer; a token lookup through the index loads no session; without an index the sessions are searched; migrations run once.
Phase 5: SessionRunner (G6, G8) ✅
System.Agents.Host(withHost) andSystem.Agents.Host.Runner(§4).DeferredCallView/pendingDeferredCalls;csCountSessionandcsDeleteSession.- Tests (
RunnerTests, mock LLM through the host’s completion override): deferred call completed with auto-resume, with the event sequence; two concurrent completions of one turn both land; messages refused during a run and accepted after, giving a second LLM turn; a background call started in one run picked up by the next; cancel during a background call, then the cancellation reported on the next run; recovery of a session left running, its call orphaned on resume;awaitRunwith timeout and completion; delete cascade (refused during the parent’s run, also for the sub-session; dry run changes nothing; real delete removes sessions and continuation rows); idle eviction and reload;withHostover agent and database files.
Phase 6: agents-server executable ✅
- wai/warp app, routes, SSE, CLI flags, graceful shutdown on SIGTERM (stop
accepting,
shutdownSessionRunner, close the database). documentation/agents-server.mduser guide, plus links fromdocumentation/durable-workflows-howto.md.- Tests (
agents-server-tests, threaded, the application on a random port with a mock LLM): the demo flow over SSE (snapshot, then per runrun.started…calls.deferred,run.stopped, then the continuation’s run endingidle); the same flow withwait=trueand no events stream (the create call returns the blocked session with its pending call, the continuation call the final answer); keepalives; listing pages and filters; delete with dry run; agents and health; error codes; shutdown releasing waiting requests and ending streams.
Milestone 2: the later work
Recorded 2026-09-19, when the user asked to continue with the later work. Each item becomes a phase, in this order: each phase is useful on its own, and the riskier ones come after the ones they build on. Choices marked default were made without asking; they are listed again under Decisions.
Phase 7: authentication and owners ✅
agents-server --auth-tokens tokens.json:{"tokens": [{"owner": "alice", "sha256": "<hex>"}, {"owner": "bob", "token": "<plain>"}]}. With it, every endpoint but/healthzneedsAuthorization: Bearer <token>(401unauthorizedotherwise, withWWW-Authenticate: Bearer). Without it, nothing changes. Tokens are compared by SHA-256 digest (default: a static file, like the API keys file; no token issuing or expiry).- Sessions created by a caller record its owner (
createSessionAs). Sub-sessions record none: a session belongs to the owner of its root session (sessionOwnerwalks upparent_session_id). This needs no change to how sub-agents store themselves. - Another owner’s session answers 404, as if it did not exist, for every
endpoint including continuations (404
unknown_token). Listing filters by owner (SessionQuery.sqOwner, index(owner, updated_at), migration 3);?parent=lists the sub-sessions of an owned session. - Sessions stored before authentication was turned on have no owner and are invisible to every caller (default).
- Not in this phase: per-owner API keys and a default
runIsolatedpolicy for bash and MCP tools. Both need agents built per owner, including sub-agent tools, which the host builds once at load time today. They stay in Remaining later work.
Phase 8: MCP over HTTP ✅
POST /mcponagents-server(AgentsServer.Mcp) speaks MCP’s Streamable HTTP transport. It handles JSON-RPC with aeson directly and reuses theMCP/Base.hstypes for tools, tool lists, and the initialize result; tool results are written by hand, becauseTextContentImplalways writes"annotations": null. Each request, or batch, gets a plain JSON response. Messages that need no answer (notifications, client responses) get 202.GET /mcpanswers 405: there are no server-initiated messages. Protocol versions 2025-06-18, 2025-03-26, and 2024-11-05: the client’s is echoed when supported, else the latest.- Methods:
initialize,ping,tools/list,tools/call, and emptyresources/listandprompts/list. Anything else is -32601; a bad or unknown tool is -32602; malformed JSON answers 400 with -32700. - Tools: one
ask_<slug>per root agent, input{prompt}. A call creates a session through the runner, owned by the caller, and waits up to 120 s. The result carries_meta.session_id. An idle session gives the final answer;waiting_externalgives a sentence plus the pending calls as JSON (not an error); a run still going gives the session id;failedgivesisError: true. - The same bearer tokens and owners apply.
Mcp-Session-Idis not used. - The Streamable HTTP transport requires validating
Originagainst DNS rebinding. Without authentication, every endpoint (not only/mcp) refuses anOriginother than localhost, 127.0.0.1, or [::1] with 403forbidden_origin. With authentication, origins are not checked.
Phase 9: TUI out of the core library ✅
- New public sub-library
agents-tui(source directorytui/) with the 25 modules that need brick or vty, or import them:System.Agents.TUI.*(exceptTUI.ToolCallActivity, which is pure and used by the tests),System.Agents.CLI.TUI,System.Agents.CLI.Config, andSystem.Agents.CLI. Module names do not change. Onlyagents-exeimports them. agents-libdropsbrick,vty,text-zipper, and the unuseddata-clist; nothing in its dependency closure pulls brick or vty any more.agents-exedepends on both libraries.- The moved files live in their own source directory: with a shared
src/, GHC would compile core modules again insideagents-tuiinstead of usingagents-lib. - Default:
agents-libitself is the core instead of a newagents-corename, so nothing that depends onagents-libbreaks.
Phase 10: Postgres backend ✅
- Public sub-library
agents-postgres(postgres/System/Agents/Postgres.hs, onpostgresql-simpleand theresource-poolalready used byagents-lib):withPostgresStores,openPostgresPool(10 connections),mkPostgresSessionStore,mkPostgresContinuationStore,isPostgresUrl,runPostgresMigrations. - Same schema and semantics as SQLite. Postgres starts at migration 1 with the
full current schema:
TIMESTAMPTZ, and session JSON asTEXT, notJSONB, which rejects\u0000. CAS isINSERT … ON CONFLICT DO UPDATE … WHERE version = 0 RETURNINGorUPDATE … WHERE version = ? RETURNING. Timestamps are truncated to microseconds, as Postgres stores them, so the metadata handed out matches what a query returns. - Migrations run in one transaction per component, under a
pg_advisory_xact_lock, withclient_min_messages = warningso that Postgres notices do not reach stderr. System.Agents.Host.withHostStorestakesHostStoresfrom the caller;withHostis the SQLite case.agents-server --dbtakes apostgres:///postgresql://URL; the startup log drops its user and password.- Tests (
agents-postgres-tests, threaded): a throwaway cluster (initdb+pg_ctlfrompg_config --bindir, a free port on 127.0.0.1), one fresh database per test;AGENTS_TEST_POSTGRES_URLuses an existing server instead, and without either the suite skips. They cover migrations running once; CAS with conflicts, labels, and the kept running status; queries (agent, owner, parent, status, limit, before); 16 concurrent CAS with exactly one winner; and a runner flow (deferred call, completion, token already completed, cascade delete). - Coordinating runs between servers came later: see “Several server processes on one Postgres database” under Remaining later work.
Phase 11: token streaming ✅
System.Agents.LLMs.OpenAIStream: reads a streamed chat completion (server-sent events,data: [DONE]) and folds the chunks back into the JSON of a non-streamed completion. It handles content, reasoning, tool calls assembled byindexfrom name and argument pieces,finish_reason, the finalusagechunk, and anerrorobject sent mid-stream. Parsing, tool calls, and storage do not change. Each non-empty text delta goes to a callback.HttpClient.Runtime.postStream: POST with a body reader for successful answers (others are read whole, so the overloaded-retry check still works).OpenAI.callLLMPayloadStreamingshares the retry logic (withOverloadedRetry) withcallLLMPayload.OpenAICompletionConfig.cfgOnTextDeltaswitches a completion to streaming and adds"stream": true(plusstream_options.include_usagefor theOpenAIv1flavor only, as other providers may refuse it).AgentDeps.adOnTextDeltapasses it throughbuildAgent.HostConfig.hcStreamTokens/hostStreamTokens; the runner gives each session’s agent a callback emittingTextDelta sid text, which is not traced (one per token, and it is content).agents-server --stream-tokens;event: text.deltawith{session_id, text}.- Tests: five unit tests of the fold (text, tool call pieces, null content,
a body read in 7-byte pieces with CRLF, comments and
[DONE], a stream error), and an end-to-end server test against a fake streaming endpoint (warp) that checks the deltas, the stored answer, and the request’sstreamfields.
Phase 12: agents from the database ✅
System.Agents.AgentStore:StoredAgent(config,updated_at,updated_by),AgentStore(asList,asPut,asDelete),mkSqliteAgentStore(tableagents, componentagents),fileBasedFields.agents-postgreshasmkPostgresAgentStore.HostStores.hsAgents :: Maybe AgentStore;withHostandwithPostgresStoresprovide one.AgentTree.loadAgentTreeFromConfig: a tree of one agent from an in-memoryAgent, reusing the loader’s later phases (create, wire tools, build) on a one-node graph, with no file discovery.Host:hostStoredAgents(loaded stored agents in aTVar, the store, a loader, and a lock serialising edits);hostAllAgents,lookupAgent,putStoredAgent,deleteStoredAgent,AgentEditError. The runner looks agents up throughlookupAgent, so new sessions see changes at once.agents-server:GET/PUT/DELETE /v1/agents/:slug,sourcein listings, MCPtools/listincludes stored agents, and--admin-owners.- Differences from the plan:
extraAgentsis refused, like the other file-based fields: stored agents have no sub-agents yet.- Editing needs
--admin-owners, which needs--auth-tokens. The plan let anyone edit agents without authentication. An agent definition can start MCP servers, which are commands on the server’s machine, so edits are off by default. - A file agent hides a stored agent with its slug: the stored one is skipped at startup, and the skip is logged. Refusing to start instead would leave no way to fix the database through the API.
- Known limitation: replacing or deleting a stored agent does not stop MCP servers it started (the tree loader has no cleanup); they end with the server.
- Tests:
- runner level: store, replace, and refuse (file-based fields, a file agent’s slug); run a session on a stored agent; reload after a restart; a file agent hiding a stored one, with the skip traced; delete persists;
- server: admin and non-admin, 201 and then 200, listing with sources, refusals, sessions and MCP on a stored agent, delete, and edits disabled by default;
- Postgres: put, replace, list, and delete.
Phase 13: self-description and a chat page ✅
The server answered 404 at / and published nothing a client could read:
the only reference was documentation/agents-server.md, which a caller holding just a
URL does not have.
- Routes as a servant type (
AgentsServer.Routes). Every endpoint whose shape OpenAPI can express is now described byDocumentedAPI, andGET /openapi.jsonis generated from it, so the published document cannot drift from the routes. The event stream andPOST /mcpstay hand-written (AgentsServer.OpenApi): neither server-sent events nor JSON-RPC is expressible, and a wrong schema would be worse than a described one. - Bodies as types (
AgentsServer.Types), carrying both the JSON encoding the API already used and an OpenAPI schema, with per-field descriptions and enums added by hand because openapi3 does not read Haddock. The turn tree stays an opaque object: it belongs to the library, not to the protocol. - A chat page at
/(AgentsServer.UI): one self-contained HTML document, no build step, no assets. It is served on a loopback bind, or when--auth-tokensis on;--no-uiturns it off. access_tokenon the event stream only, becauseEventSourcecannot set headers. The request log records no query strings, so the token does not reach the logs.
Deviations from the plan:
- The rewrite is staged. The servant types describe and generate the
document; the wai router still dispatches. Moving dispatch onto servant
needs
UVerbfor the 200/201/202 answers andRawfor the stream, and would rewrite 646 tested lines for no change a client can see. A test asserts the document’s paths are exactly the router’s, so the two cannot drift in the meantime. openapi3needed the freeze’sindex-statebumped: 3.2.4 caps QuickCheck below the pinned 2.16, and 3.2.5 (which allows it) was published after the pinned index date. Every previously pinned version is unchanged.
Database agents with tools and helpers ✅
Lifts the two refusals of Phase 12 that decision 10 recorded.
- Tool files are stored with the agent.
StoredAgent.saFiles(contents by relative path; columnfiles, SQLite and Postgres migration 2);asPuttakes them.PUT /v1/agents/:slugreads them from afilesobject in the body, andconfig.filesshows them. - Loading writes them to disk. Each load of a stored agent gets its own
directory under a temporary directory of the host
(
materializeToolFiles); the loaded configuration’stoolDirectoryandbashToolboxespaths point there. The directory is removed when the node is released, and with the host. extraAgentsname stored agents by slug (pathis now optional in the JSON, and refused for a stored agent).resolveHelperscomputes what an agent reaches;AgentTree.loadAgentTreeFromConfigsloads the root and those helpers as one tree, so each root has its own copy of its helpers, like roots from files.putStoredAgentWithFiles;AgentEditErrorgainsAgentInvalidPaths,AgentHelperError(UnknownHelper,HelperCycle), andAgentInUse.- Choices made here (default, none was planned in detail):
- Cycles are refused, self-reference included, although agent files allow them: an edit is checked against the agents stored at that time, and a helper must exist first, so a cycle can only be an error.
- Replacing a helper reloads its dependents; a dependent that fails to load again keeps its previous version, and the failure is traced.
- A helper in use cannot be deleted (
409 agent_in_use). - Only stored agents can be named, not agents from files.
openApiToolboxes,postgrestToolboxes,skillSources, andautoEnableSkillsare still refused.
- Tests: runner level (a stored agent with a tool directory and a stored
helper with a single tool, run end to end, across a restart, and after
replacing the helper; unknown helper, cycles, paths); server (
filesin and out,unknown_helper,helper_cycle,agent_invalid_paths,agent_in_use, a session that runs the tool); Postgres (files round-trip).
Remaining later work
- Per-owner API keys and isolation: build agents per owner, including
sub-agent tools; default policy
runIsolatedfor bash and MCP tools, backed bydockerRunner. Partly done:- Keys:
HostConfig.hcOwnerApiKeysFiles(--owner-api-keys OWNER=FILE), loaded intoHost.hostOwnerApiKeys. The runner already builds one agent per session, so it picks the keys there (Runner.newAgent), from the owner of the session’s root. Sub-agents needed one fix to be covered:createSessionForNodeWithnow always builds the agent from the node it was given. Before, a helper that was not also a root agent failedlookupAgentand silently fell back to the in-tool path, which is built at load time and knows no owner. That fallback remains for real failures, and gets no key at all when per-owner keys are in use. - Isolation:
System.Agents.Tools.Isolated,AgentDeps.adToolIsolation,HostConfig.hcToolIsolation(--isolate-tools docker:IMAGE|process:PATH). Enforced inbuildAgent, around the tool execution rather than throughToolCallPolicy: a policy is the agent’s own (it would let an agent file opt out, and replacing it would dropdeferrules), and aRunIsolateddisposition without a runner runs in-process. - Still open (owner decisions): isolation is opt-in, not the default, as there is no worker image to default to; it covers bash and MCP tool calls only; MCP servers still start on the host; no worker ships with the repository; unlisted owners use the shared keys.
- Keys:
- ~~Several server processes on one Postgres database: live-run ownership
through a lease column (
run_owner,run_lease_until).~~ Done:System.Agents.Host.Coordination: what a host needs from a shared database (acquire, renew, release a run lease; who holds it; which running sessions lost their owner; signals from the other processes).noCoordinationis the single-process case (SQLite, the TUI) and changes nothing there.HostStores.hsCoordination,Host.hostCoordination.agents-postgres: migration 4 addsrun_ownerandrun_lease_untiltosessions; lease times are the database’snow(). Each process gets a random instance name. Session writes and accepted mail are announced withpg_notifyonagents_sessions; one connection outside the pool listens and reconnects.- The runner takes the lease in
startRunand releases it wherever a run ends. A heartbeat (every third ofrcLeaseTtl, default 30 s) renews, stops runs whose lease was taken (abandonRun, nothing stored), syncs the mail of running sessions, and takes over expired sessions (takeOverExpired, same recovery asrecoverOnStartup, nothing resumed).recoverOnStartuponly recovers sessions without a live lease. A run held elsewhere is busy:postMessageandcompleteCallgo to mail,resumeanddeleteSessionare refused,cancelRunpostsStopRun,awaitRunwaits on forwarded events. - Writes are fenced by the versioned compare-and-store, not by the lease:
a takeover stores a version, so the previous owner’s next write
conflicts, and
failRunstores nothing once the lease is gone. - Mail:
msAppendanswers with the envelope as stored. Postgres settlessequnder an advisory lock per session, since two processes each number from what they last loaded;mbSyncmerges what others appended. - Events: only
session.updated,run.stoppedandsession.deletedare forwarded to the other servers. Text deltas and tool events stay on the owner. A watch skips the forwarded copies (srRemoteEvents), so a registration known to two servers yields one mail per event. - Not done: resuming a taken-over run by itself; sharing stored-agent
edits and
watch-sessionregistrations between servers; anagents-serverflag for the lease duration.
- ~~Database agents with files or sub-agents~~: done, see Database agents with tools and helpers.
- Stopping a stored agent’s MCP servers when it is replaced or deleted.
- Dispatching through servant, replacing the wai router (see Phase 13).
Decisions
Recorded 2026-09-18.
- Waiting for a run is chosen per request with the
waitandtimeoutquery parameters (§5), on every endpoint that can start a run. The default is not to wait. The body’srunfield still controls how far the run goes. - Deleting cascades to sub-sessions and continuation rows, and supports
a dry run (
?dry_run=true) that reports what would be deleted (§4.9). - Idle live sessions keep their engine for 15 minutes
(
hcLiveSessionTtl, configurable). A background call still running when its session is evicted ends up orphaned. - The MCP server over HTTP is wanted, but later (see Phase 8).
Recorded 2026-09-19 (defaults taken when the user asked to continue with the later work; see Milestone 2):
- Authentication is a static bearer-tokens file mapping tokens (or their SHA-256) to owners. A session belongs to the owner of its root session; other owners get 404. Sessions without an owner are invisible once authentication is on.
- MCP over HTTP answers each request with plain JSON and does not use MCP sessions.
agents-libbecomes the core by moving the TUI into a newagents-tuilibrary, rather than creating anagents-corelibrary.- Postgres lives in its own
agents-postgreslibrary; the host takes its stores from the caller. - Token streaming is opt-in (
--stream-tokens). - Database agents cannot use file-based tools in their first version,
nor
extraAgents. Storing agents needs--admin-owners(and so authentication). A file agent hides a stored agent with the same slug.
Recorded 2026-09-20:
- The OpenAPI document is generated from servant types, rather than hand-written or derived from a full servant rewrite of the dispatcher.
- The chat page is loopback-only by default and off with
--no-ui. It is plain HTML and JavaScript in one document, with no build step.
Related docs
todos/durable-workflows.md,todos/durable-workflows.progress.mdtodos/async-tool-calls.mddocumentation/durable-workflows-howto.md,documentation/async-tool-calls.md,documentation/sessions.md