Architecture
The states an agent moves through, and what survives sleep, crashes, and upgrades.
A pi() agent runs Pi Durable, Earendil’s durable agent harness, inside a Rivet Actor. Pi saves each step of a run to the Actor’s SQLite database before the next step starts.
States
A conversation is idle until a prompt starts a run, and goes back to idle when the run ends or is cancelled. When the context fills up during a run, the agent compacts it and keeps going.
Sleep
An idle agent sleeps and frees its memory. The next action on its key wakes it: the agent opens Pi from SQLite, resumes interrupted work, and reconnects to the same sandbox the first time a tool needs it, so clients don’t do anything different. While the agent sleeps, its sandbox is suspended.
During a retry or poll wait longer than a minute, the Actor also sleeps, and a scheduled wake resumes the work when the wait ends.
What survives
| Conversations | Sandbox | |
|---|---|---|
| Sleep | Every step so far. | Reconnected on first use. |
| Crash | Every step saved before the crash. A run in progress resumes on wake from the step it reached. | Reconnected on first use. |
| Upgrade | Every step so far. A run in progress gets time to finish first. See Upgrades. | Reconnected on first use. |
| Destroy | Deleted. | Deleted. |
A model or instructions set with conversation.configure are part of the conversation, so they are kept too. Without a sandbox, files live in the Actor’s SQLite database and survive the same way as conversations. If the sandbox no longer exists when the agent wakes, a new one is created and the old files are lost.
Actions that don’t run a tool, such as conversation.context or conversation.configure, don’t connect to the sandbox. They keep working while the sandbox provider is down.
Resuming a run
After a crash or a cut-off upgrade, the agent reads the saved steps on wake and carries on from the step it reached:
| Stopped while | On wake |
|---|---|
| The model was answering | The model call is sent again from the start. The partial answer is lost. |
A tool marked replay: "safe" ran | The tool runs again. |
| Any other tool ran | The tool isn’t run again. Its result tells the model the call stopped partway. |
| A task phase ran | The phase runs again, and earlier phases are kept. |
The run was cancelled with abort | Nothing resumes. |
Nothing runs while a crashed Actor is down. The run resumes when the Actor starts again. A client that is still connected gets a fresh snapshot on its event stream.
Upgrades
When you ship new code, running work gets up to sleepGracePeriod (15 minutes by default) to finish before the Actor restarts on the new version. See Versions. Work still running when the grace period ends resumes on wake, after your onWake runs.
A client waiting on prompt gets an error if the Actor stops first. It can send the prompt again with the same requestId to get the first prompt’s result instead of a second run. Actions wait up to ten minutes, and a timed-out wait leaves the run going.
Destroying the Actor stops running work at once, without a grace period.
Roll forward only. An older version of Pi Durable can’t open data that a newer version wrote, so don’t roll back after an upgrade has run.