Plan: Durable Workflows for agents-exe
On Sun, 04 Oct 2026, by @lucasdicioccio, 1127 words, 12 code snippets, 1 links, 0images.
Generated from todos/durable-workflows.md, the repository is the canonical source and may be ahead of this page.
Plan: Durable Workflows for agents-exe
See todos/durable-workflows.progress.md for implementation progress.
Goal
Goal
Turn agents-exe into a system suitable for:
- Durable executions: sessions can be stored in and resumed from databases.
- Isolated deployments: tool calls can run inside Docker or as local processes, transparently.
- Transparent synchronous/asynchronous tool calls: the runtime can pause a turn after any subset of tool calls, yield to the outside world, and wake up later when results arrive.
Requested example flow:
- user-turn 1: some question
- agent-turn 1: reply with 3 tool calls
- user-turn: perform 1 tool call, mark 2 as deferred
- system: yields (possibly even reboots)
- system: wakes up, executes deferred tool calls and places results, then wakes up again
- user-turn: provides 3 tool results to agent
- agent: continues
This document lists the primitives and architectural changes needed to support that flow.
Current state
The codebase already has scaffolding:
ExecutionMode(Synchronous/Asynchronous)PartialUserTurnContinuationToken,AsyncToolResponse(ToolComplete/ToolYield)ContinuationStore/ToolContinuationToolCache- OS persistence layer with SQLite/Postgres backends
- File-based
SessionStore
What is missing:
- The async scheduler does not actually use
ctxAsyncToolCallorAsyncToolResponse; it just runs tools synchronously and pauses after each one. - There is no per-call decision primitive for sync vs. deferred vs. isolated execution.
- There is no “wake up and inject external results” primitive.
ContinuationStorehas stubbed load/list implementations.ToolExecutionContextcontains non-serializable fields (portal, world, event queue), making durable continuations fragile.- Session storage is file-only.
- There is no abstraction for running a tool call outside the current process (Docker / subprocess).
Phase 1 — Core primitives (types & interfaces)
1.1 Rich tool-call state inside a turn
Introduce an explicit state machine for every LLM-issued tool call within a user turn:
Ready -> Running | Deferred | Completed
Running -> Completed | Failed
Deferred -> Ready | Completed (via external wake)
Completed -> terminal
Add a TrackedToolCall type wrapping LlmToolCall with:
tcId :: ToolCallId— stable ID for matching external resultstcState :: ToolCallStatetcResult :: Maybe UserToolResponsetcContinuation :: Maybe ContinuationTokentcPolicy :: AppliedPolicy— record of why it ran this way
Refactor PartialUserTurnContent to hold [TrackedToolCall] instead of separate pCompletedResponses, pPendingCalls, and pPendingContinuations lists.
1.2 Tool-call policy / decorator
Add a pure decision primitive:
data ToolCallDisposition
= RunSync
| RunAsync -- yield a continuation token
| RunIsolated IsolationSpec -- run outside the current process
| Defer Reason -- intentionally pause (e.g., approval)
| Decorate [Decorator] ToolCallDisposition
type ToolCallPolicy =
ToolExecutionContext -> LlmToolCall -> ToolCallDispositionThis is the requested decoration primitive. It lets the runtime decide, per call, whether to execute immediately, defer, or run in isolation.
Provide defaultToolCallPolicy = const $ const RunSync for backward compatibility.
1.3 Pluggable executor interface
Decouple how a tool call runs from the session loop:
data ToolExecutor = ToolExecutor
{ execSync :: ToolExecutionContext -> LlmToolCall -> IO UserToolResponse
, execAsync :: ToolExecutionContext -> LlmToolCall -> IO AsyncToolResponse
}Concrete executors to provide:
inProcessExecutor— current behavioryieldingExecutor— always returnsToolYieldcachingExecutor :: ToolCache -> ToolExecutor -> ToolExecutorisolatedExecutor :: DeploymentRunner -> ToolExecutor
Replace the single Agent.ctxAsyncToolCall with an optional ToolExecutor selected by the policy.
Phase 2 — Async scheduler in the session loop
2.1 Replace the one-at-a-time async loop
Current runStepMAsync executes exactly one pending call per step and then pauses.
New scheduler behavior, given a set of TrackedToolCalls:
- Classify each call via
ToolCallPolicy. - Run all
RunSynccalls and collect results. - For
RunAsync/Defercalls, generate continuation tokens and move them toDeferred. - Yield a
PartialUserTurncontaining:- completed calls,
- deferred calls with tokens,
- any remaining ready calls.
This matches the requested flow:
perform 1 tool call, mark 2 as deferred → system yields → wakes up, executes deferred calls, places results → user-turn provides 3 tool results.
2.2 Continuation snapshot
When a call is deferred, generate a ContinuationToken and store a serializable ToolContinuationSnapshot:
tokensessionIdtoolCallIdllmToolCallcacheKeypolicy/isolationSpec- serializable context fields:
ctxSessionId,ctxConversationId,ctxTurnId,ctxCallStack,ctxAllowedTools,ctxParentConversation
Do not store the non-serializable fields (ctxToolPortal, ctxWorld, ctxEventQueue). Re-hydrate those on wake.
Phase 3 — Durable session storage
3.1 Generalize SessionStore
The current SessionStore is file-only. Introduce a backend interface:
data SessionBackend = SessionBackend
{ sbStore :: SessionId -> Session -> IO ()
, sbLoad :: SessionId -> IO (Maybe Session)
, sbList :: IO [(SessionId, UTCTime)]
, sbDelete :: SessionId -> IO ()
}Backends:
FileSessionStore FilePathSqliteSessionStore ConnectionCompositeSessionStore [SessionBackend]for read fallback with a single write target
3.2 Use OS persistence (optional)
The OS layer already has SqliteBackend, component tables, and migrations. Consider adding a SessionComponent persisted via persistComponent/loadComponent. This reuses ECS migrations and unifies backends.
If session storage stays separate, at least align IDs: SessionId/TurnId/ConversationId are UUID-based, and OS EntityId is also UUID-based, so conversion is trivial.
Phase 4 — Wake / resume API
4.1 Inject external results
wakeSession ::
Session
-> [(ContinuationToken, UserToolResponse)]
-> IO SessionBehavior:
- Find the
PartialUserTurnand matching deferredTrackedToolCalls. - Move them to
Completedwith the provided result. - Update the cache if configured.
- If all calls are complete, convert the
PartialUserTurnto a fullUserTurn. - Otherwise leave remaining deferred/ready calls in place.
4.2 Resume execution
resumeSession ::
ConversationId
-> Agent r
-> Session
-> IO (Either r Session)- If the latest turn is
PartialUserTurn, run the scheduler on remaining ready/deferred calls. - If the latest turn is
UserTurn, continue to the LLM step. - Returns
Left ron completion,Right Sessionwhen it needs to yield again.
4.3 Complete continuations from external workers
completeContinuation ::
ContinuationStore
-> ToolCache
-> ContinuationToken
-> UserToolResponse
-> IO BoolThis already exists as resumeAsyncToolCall but the ContinuationStore implementation is stubbed. Finish the load/list implementations and ensure proper JSON round-tripping.
Phase 5 — Isolated deployment primitives
5.1 DeploymentRunner abstraction
data DeploymentRunner = DeploymentRunner
{ drName :: Text
, drExecute :: IsolationSpec -> LlmToolCall -> IO (Either IsolationError UserToolResponse)
}Implementations:
localProcessRunner— fork a worker processdockerRunner— run inside a containerfunctionRunner— for serverless/FaaS (future)
5.2 Serialization contract for isolated calls
Define a stable JSON envelope consumed by any external runner:
{
"token": "...",
"toolCall": { ... },
"contextSnapshot": { ... },
"policy": { "tag": "RunIsolated", "spec": { "image": "...", "sandbox": "..." } }
}Result envelope:
{ "token": "...", "status": "success", "result": { ... } }This makes Docker vs. local execution transparent to the agent.
5.3 Tool-call decorator for isolation
Example policy:
isolateBash :: ToolCallPolicy
isolateBash ctx call
| toolName call == "bash_command" = RunIsolated (Docker "agents-exe/bash-runner:latest")
| otherwise = RunSyncPhase 6 — Integration & agent combinators
6.1 Extend Agent record
Add to Agent r:
ctxToolCallPolicy :: ToolCallPolicyctxSessionBackend :: Maybe SessionBackendctxContinuationStore :: Maybe ContinuationStorectxDeploymentRunner :: Maybe DeploymentRunner
Keep defaults so existing agents compile unchanged.
6.2 New combinators
withToolCallPolicy :: ToolCallPolicy -> Agent r -> Agent r
withSessionBackend :: SessionBackend -> Agent r -> Agent r
withContinuationStore :: ContinuationStore -> Agent r -> Agent r
withDeploymentRunner :: DeploymentRunner -> Agent r -> Agent r6.3 Update agentStoreSession
System.Agents.Combinators.StoreSessionProgress should store via the configured SessionBackend rather than only files, while still supporting file fallback.
Phase 7 — CLI / operator API
New commands for real operation:
agents session pause <session-id>— yield after current stepagents session resume <session-id>— continue executionagents session pending <session-id>— list deferred calls / continuation tokensagents session complete <token> <result-file>— inject an external resultagents session run-isolated <session-id>— poll and executeRunIsolatedcalls
Phase 8 — Testing strategy
Add tests for:
- Policy classification: a policy can mark some calls sync and some deferred.
- Partial turn serialization: a session with deferred calls round-trips through SQLite/file backends.
- Wake/resume:
wakeSessioninjects results and eventually produces a fullUserTurn. - Cache integration: deferred calls whose result is later cached are resolved on resume.
- Isolation contract: a
DeploymentRunnercan execute a tool call via a subprocess and return a result envelope. - Determinism: the same session state resumed twice produces the same outcome.
Open questions
-
Should
PartialUserTurnContentkeep its current shape or be rewritten aroundTrackedToolCall?
Rewriting is cleaner but touches JSON serialization and existing tests. -
Should session storage move into the OS ECS persistence layer or stay separate?
Moving into ECS unifies migrations/backends; keeping it separate is less invasive. -
How do we handle non-serializable
ToolExecutionContextfields on resume?
Recommended: store a serializable snapshot and re-hydrate portal/world/event queue in the runtime. -
Do we want true concurrency for sync calls or sequential execution?
Sequential is simpler and matches current cache semantics; concurrency can be added later behind a policy flag. -
Should
ToolCallIdbe a UUID or an index into the turn?
A UUID is safer for external workers and database keys; an index is smaller and deterministic. Hybrid:ToolCallId = TurnId + Int index.
Suggested first milestone
The smallest vertical slice that proves the design:
- Add
TrackedToolCalland refactorPartialUserTurnContent. - Implement
ToolCallPolicywithRunSync/Deferdecisions. - Update the async scheduler to batch sync calls and yield deferred ones.
- Implement
wakeSession. - Add a SQLite-backed
SessionBackend. - Write a test exercising the full requested flow:
- 3 tool calls issued,
- policy executes 1 and defers 2,
- session persists,
- wake injects 2 results,
- session resumes and produces a
UserTurnwith all 3 results.
This milestone delivers the requested primitives without yet building Docker runners or Postgres backends.