Skip to content

fix(messaging): a Codex session start never removes a socket it cannot prove is its own - #570

Merged
kevintseng merged 10 commits into
mainfrom
fix/router-version-check
Oct 1, 2026
Merged

kevintseng merged 10 commits into
mainfrom
fix/router-version-check

Conversation

@kevintseng

@kevintseng kevintseng commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Refs #518.

A Codex session start no longer removes a socket it cannot prove is its own, and a router that is older than the installed MeMesh says so instead of failing silently.

  • When the router is outdated, the session is told why it could not reach it, and the router steps aside for the installed version.
  • The companion socket has a new name, of the same length. A record from before this change no longer blocks a start, and its socket file is left in place.
  • A companion records its socket right after binding it, and is marked registered once the router accepts it. A start waits for that mark. A second start refuses while the first is still starting.
  • A socket that no record owns is left alone, and the start explains why in one line. This includes the case where the bind loses with EEXIST, which macOS sometimes reports instead of EADDRINUSE.
  • When a recorded companion is running but does not answer, the message names the process, and it no longer suggests deleting the record, except for a record from before this change.
  • Each cleanup is written to codex-companion.log.

Known limit: a socket left with no record, and a record whose process id was reused by another program, still need the user to stop that process; MeMesh will not guess.

…pgrade

After an upgrade, the local message router that was already running kept
routing with the old code, and nothing noticed. The router now reports its
MeMesh version when a session connects, and each session sends its own. An
older router from this release on steps aside, and the installed one starts.
A router from 4.10.11 or earlier is refused, with the command to stop it, and
the Claude Code and Codex hosts print that reason.

A session send that falls back to its principal now also tells the sender
when the intended session is not connected, since only that session can take
the message in.

Fixes #518
A host that kept running across an in-place downgrade kept claiming the newer version it started with, so the router installed now kept stepping aside. The version is now read each time the host registers. A host that loses its router and finds an outdated one when reconnecting now says so once.
…pgrade

A package.json missing for a moment during an upgrade was taken for a missing router, so the host started another one. It is now retried like any other passing connection failure. The note about an outdated router is printed again after a later reconnect.
…ed router

Under Codex the router connection is made by a background process with no
output, so a session that met a router from before the upgrade only printed
"session registration failed". The reason and the command to stop the old
router now reach the session start output.

The background process hands the reason to the waiting session start through
a private file that is named after that one launch (a reused process id never
picks up an earlier launch's reason), written whole by rename (never read
half-written), and written only after the background process has closed its
socket and removed its state (stopping it cannot leave a socket behind). A
reason nobody collected is removed by a later launch once it is a minute old.

Refs #518
- A companion stopped at the session start's deadline, by a signal or by a
  failed start now runs one cleanup and exits with the right code; the
  session start waits for it to finish before it reports.
- The reason it could not register is written before cleanup, so a stop
  that is already under way cannot lose it.
- A control socket is removed only while it is still the same socket this
  companion recorded (same inode). When a newer companion already holds
  the path, the old one exits without unlinking it, so the newer one stays
  reachable.
- A dead companion's record is cleared only after the record itself is
  moved aside atomically and proven to be that companion's. A record from
  an older MeMesh that did not say which socket was its own is removed and
  the socket left in place, with one line saying how to check it and what
  to delete. A socket nothing owns is never probed or removed.
- A running companion whose socket does not answer gets its own reason,
  with the pid and the steps to stop it.
- A peer that disconnects from the control socket can no longer crash the
  companion.
- A temporary reason file left by a companion stopped mid-write is swept.
- A companion that cannot stop cleanly writes why to codex-companion.log
  in the MeMesh data folder, and exits 1.
- A session end whose companion had already exited finishes normally even
  when its record came from an older MeMesh; a session end that does fail
  says "session end failed", not "registration failed".

Refs #518
Also in this merge, so a Codex session start never removes a socket it cannot prove is its own:
- The companion socket has a new name, of the same length. A record from before this change no longer blocks a start; its socket file is left in place.
- A companion records its socket right after binding it, and is marked registered once the router accepts it. The start waits for that mark. A second start refuses while the first is still starting.
- A socket that no record owns is left alone, and the start says why in one line. This also holds when the bind loses with EEXIST, which macOS sometimes reports instead of EADDRINUSE.
- When a recorded companion is running but does not answer, the message no longer suggests deleting its record, unless the record is from before this change.
- Each cleanup is written to codex-companion.log.
Comment thread tests/host-adapters/codex-session-failure-channel.test.ts Fixed
Comment thread tests/host-adapters/codex-session-failure-channel.test.ts Fixed
Comment thread tests/host-adapters/codex-session-failure-channel.test.ts Fixed
…s socket

A companion that stops removes its control socket first and its lifecycle record once the close completes. The test waited only for the socket, so on a loaded runner it could read the record directory between the two. It now waits for both, within the same bound, and still asserts both are gone.
…t code

The fault-injection preload reads its settings from a JSON file beside it, and the socket holder takes its path from an environment variable, instead of having those values written into the code they run.
@kevintseng
kevintseng merged commit c805145 into main Oct 1, 2026
12 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants