// Case file · cross-agent continuity
Resuming a Live-Money Session Across Agents
A 7,244-turn Codex session stopped during an IB Gateway outage. A fresh Claude session was asked to continue it. Instead of guessing, it queried AgentPM for the exact final state and found a prior stalled handoff too.
7,244
turns in the origin session
2
agents, models, and tools involved
2nd
resume attempt found in history
0
blind writes to the live host
Situation brief
A live-money system, left in a bad state
The origin work was a long Codex session on a live trading system: add a live account and test live trades. Deployment happened over SSH to a single host through the usual git push and process restart path.
The session did not end at a clean stopping point. It ended inside an incident.
Recovered from the source transcript: around 23:45 UTC, the gateway on the host lost its session. The API socket was down and the live brokerage account was disconnected. On a live-money system, that kind of state quietly stops trading.
[7264] tool BrokerAPIError: Timed out connecting to IB Gateway at 127.0.0.1:4001. Confirm IB Gateway is logged in and API socket clients are enabled. [7265] assistant IB Gateway shut down/disconnected around 23:45 UTC and systemd restarted it. It's now at the GUI level, but not authenticated enough to open 4001...
Trigger
Continue someone else's session, blind
A new session opened in a different tool, on a different machine, with none of the origin context in its window. The instruction was just a conversation ID.
On a live-money system, this is the dangerous moment: a capable agent, no memory of what happened, and one SSH hop from production. The safe move is to reconstruct the exact state before acting.
I want to use agentPM to continue work on goc_71ec51d033ff81398744ae6a from a codex session.
Retrieval
Querying its own history for ground truth
The agent launched the agentpm-conversations skill and ran read-only AgentPM commands. It located the origin session, inspected recent project work, checked mutating git activity, and read the final turns.
The important detail: AgentPM surfaced a sibling conversation opened minutes earlier, titled as an attempt to continue the same origin session. The new agent stopped to inspect that attempt before doing anything.
There's already a follow-on conversation that tried to continue this. Let me check it plus the exact final state.
Launching skill: agentpm-conversations
$ agentpm conv search goc_71ec51d033ff81398744ae6a
1 conversation(s)
1. goc_71ec51d0... 350-nice "...add a live alpaca account and test live trades"
2026-07-09 turns=7244
$ agentpm conv work
350-nice ... conversations=3
goc_07649831... "Continue work on goc_71ec51..." <- 21:27
goc_fbef2a72... "Review NewTradingThing security concerns"
goc_3de9e90... "Review VNC configuration work from Codex"
$ agentpm conv git goc_71ec51... --mutating-only
turn 5652 ssh will@<host> "cd NewTradingThing && git push origin main"
cmd=push branch=main remote=origin mutating=true ...
$ agentpm conv show --around 7244 --limit 40 goc_71ec51...
[7264] BrokerAPIError: Timed out connecting to IB Gateway at 127.0.0.1:4001...
[7265] IB Gateway shut down/disconnected around 23:45 UTC...What history exposed
Three findings a single transcript would hide
The reconstruction folded together three separate facts, each anchored to a different conversation.
First, the exact final state: the origin ended inside an outage, with socket 4001 down. Second, an unshipped safeguard: a prevention plan existed in the record, but had not been implemented. Third, a stalled prior handoff: another resume attempt had already reached the same point and stopped.
I've reconstructed where this Codex goal left off and found that an earlier attempt to continue it stopped right at this same decision point before it was interrupted.
- Origin ended inside an IB Gateway outage, with socket 4001 down.
- A layered prevention plan had been proposed but not implemented.
- A prior resume attempt had already stalled at the same decision point.
- Without AgentPM, the new session could have become the third restart.
Consequence
A safe, deduplicated handoff instead of a third restart
Knowing the system was live and mid-incident, and knowing a prior attempt had already stalled, the agent did not act. It converted the recovered context into an explicit decision for the operator.
That is the practical difference between resuming from a title and resuming from ground truth. A title says continue work. The evidence says what is safe to touch, what is known, what is inferred, and what needs human direction.
Which thread should I pick up from the Codex session? IBKR prevention plan, re-login IBKR now, or Alpaca live trades?
- No blind writes to the live host.
- Exact final state pulled from source turns.
- Prior stalled attempt detected and read instead of repeated.
- Operator handed a scoped, reversible decision.
Evidence ledger
What is known, inferred, and still open
The case deliberately separates recorded facts from inferred state. Recorded in a conversation is not the same thing as verified in production.
Directly observed: the AgentPM commands and outputs, the origin final state from source turns, the prior resume attempt, and the structured decision returned to the operator.
Agent-stated or inferred: that the prevention plan was never implemented and that the prior attempt stopped before it was interrupted at the same point.
Still open: whether the watchdog or hard gate was implemented later, and whether the outage recurred. Those require a later production check, not a claim from this case file.
Attribution
Who did what
AgentPM preserved the 7,244-turn origin session, connected sibling resume attempts, and retrieved the exact final state on request.
The coding agent chose to reconstruct before acting, read the prior attempt, and framed the decision.
The operator directed which thread to pick up and whether anything should touch the live host.
A live check would still be required to confirm current production and repository state. AgentPM did not restart the gateway or fix the trading system. It made the prior work retrievable and comparable.
Conclusion
The useful unit of memory was the actual session history
Static notes would not have known the system was mid-outage, that a safeguard was proposed but never shipped, or that another agent had already tried and stalled minutes earlier.
Cross-session retrieval did. It turned a risky blind continuation into a scoped, evidence-backed handoff.
The useful unit of memory here was not a note someone wrote down. It was the actual history of every session: searchable, comparable, and tied to exact turns.
Continuity means the next agent can see the state it is inheriting.
AgentPM lets agents and humans resume from evidence instead of titles, guesses, or fragile summaries, especially when the work touches production.
Real AgentPM case study, shared with permission. Organization, host, account, application, repository, and directory names have been obscured. Commands and quotes are from the AgentPM record; later production state is out of scope.
Start capturing agent work