Lab notes · 2026-09-20

Two drains, two days apart: 85 and 88 % of the probes came back, and the same ten never did

A maintenance window on an anycast node ends the same way everywhere. The operator re-enables what was disabled, the router reports the sessions up and the routes exported, the collectors show the prefix again, and the ticket is closed. We wanted to see the other side of that moment, so we drained one of our own sites, put it back, and kept measuring for six hours. AS218833 is our lab network: one /24 and one /48 announced from two sites, a node in Frankfurt and a node in New Jersey, serving one zone whose answers carry the site name in the NSID option. Every answer therefore says which of the two produced it.

We did it twice. On 17 September we drained Frankfurt for 30 minutes and then New Jersey for 30 minutes. On the night of 19 to 20 September, about 52 hours later, we drained Frankfurt again, with 20 minutes of undisturbed baseline in front of it and six hours of observation behind it. Both times the routes were withdrawn over the existing sessions, the way an operator drains a site rather than the way a link fails: all four BGP sessions stayed Established throughout, only the exported routes went away. Watching: a fixed set of 100 RIPE Atlas probes chosen by continent, 96 of which returned results in the second run, one DNS query per probe, family and minute, sent straight to our anycast address with no resolver in the path, then one query every five minutes for five and a half hours; RIPE RIS Live for the control plane; and nine hosts of our own every ten seconds.

Read this first. This is a two-site lab on one prefix per family, not a production anycast network, and the cohorts are small. Before the second drain, 20 probes on IPv4 and 59 on IPv6 were served by Frankfurt in every answered measurement of the 20-minute baseline; those are the two cohorts, and every percentage below has one of those two denominators. Probe shares are not traffic shares. What follows is a behaviour we measured on our own prefix, not a claim about how often it occurs elsewhere.

Leaving is fast and complete

Nothing surprising happened on the way out. In the 30 minutes of the second drain, 20 of 20 IPv4 probes and 59 of 59 IPv6 probes were answered by the other site, and all of them by the other site: not one landed on a stranger's nameserver, not one stayed on Frankfurt, not one fell silent for the whole window. Half the IPv4 cohort had moved within 20.3 seconds and half the IPv6 cohort within 28.3 seconds, 90 % within 68.3 and 74.3 seconds, all of them within 77.3 and 87.3 seconds. The measurement grid is 60 seconds, so each of those numbers is an upper bound on the switch, not the switch itself.

The control plane agreed. Of the RIS peers whose best path ran through the Frankfurt upstream, 91 of 93 on IPv4 and 180 of 183 on IPv6 were seen to leave, the first of them 10.3 seconds after the local withdraw, half of them within 26.3 and 32.3 seconds. The usual event was not a withdrawal but an announcement carrying the other site's path, which is what anycast is supposed to do.

Coming back is fast, and then it stops

The return starts just as quickly. Half of each cohort saw Frankfurt again within 25.4 seconds (IPv4) and 23.4 seconds (IPv6) of the restore. The last return arrived at 80.4 seconds on IPv4 and 82.4 seconds on IPv6. After that, in five hours and 58 minutes of further measurement, not one further probe came back. 17 of 20 IPv4 probes (85.0 %) and 52 of 59 IPv6 probes (88.1 %) had returned, and those two numbers are the same at 30 minutes, at one hour, at three hours and at six hours.

Return curve after the restore: the share of each cohort back at the restored site rises to 85 % on IPv4 and 88 % on IPv6 within the first two minutes and stays flat for the remaining six hours, never crossing the 90 % line.
Share of each cohort back at the restored site, six hours after the restore. The whole movement happens in the first two minutes; the dashed line is the 90 % mark that neither family reaches.

This is not an artefact of the coarse grid. The entire movement happens inside the first two minutes, which were measured every 60 seconds; the five-and-a-half-hour measurement at 300 seconds contributed not a single additional return. "Time to 90 % back" is therefore not a number we can report, because 90 % was never reached, in either family.

Meanwhile the node looked healthy from the inside. All four sessions Established, four routes exported, the static protocols up, the local log clean, no errors in the trail. Everything our own router could tell us said the maintenance was over.

The same ten entries, twice, in opposite directions

The interesting part is who stayed away, because we had asked the same question two days earlier. After the New Jersey drain on 17 September, ten entries did not return to New Jersey within the 31.5 minutes we observed. After the Frankfurt drain on 19 September, ten entries did not return to Frankfurt within six hours. They are the same ten: three on IPv4 and seven on IPv6, eight distinct probes, two of which appear in both families. The symmetric difference between the two lists is empty in both families, and all ten were members of the cohort in both runs, so each of them was actually asked both times.

In between, they did not drift back on their own. About 52 hours after the first experiment ended, in the 20-minute baseline before the second drain, every one of them was answered by Frankfurt again in every single answered measurement, which is where the first experiment had left them.

NetworkFamilyAfter the
New Jersey drain
52 h later,
before the next
After the
Frankfurt drain
ISP (ID)IPv4Frankfurt 31/31Frankfurt 18/18New Jersey 96/96
small network (NO)IPv4Frankfurt 31/31Frankfurt 19/19New Jersey 96/96
hosting network (US)IPv4Frankfurt 31/31Frankfurt 19/19New Jersey 96/96
cloud provider (CH)IPv6Frankfurt 31/31Frankfurt 19/19New Jersey 95/96
ISP (CA)IPv6Frankfurt 31/31Frankfurt 19/19New Jersey 96/96
ISP (BO)IPv6Frankfurt 31/31Frankfurt 19/19New Jersey 95/96
small network (NO)IPv6Frankfurt 31/31Frankfurt 19/19New Jersey 96/96
network (BR)IPv6Frankfurt 31/31Frankfurt 19/19New Jersey 96/96
ISP (CH)IPv6Frankfurt 31/31Frankfurt 19/19New Jersey 96/96
hosting network (US)IPv6Frankfurt 31/31Frankfurt 19/19New Jersey 96/96

The counts are answered measurements carrying a site name in each window: 31 in the 31.5 minutes after the first restore, 18 or 19 in the baseline, 96 in the six hours after the second. The two entries at 95 of 96 each had one timeout; none of the ten produced a single answer from the site it had left. The labels are deliberately coarse, and the country code is all the geography we give: the probe in Brazil carries no IPv6 AS number in the Atlas metadata, so we could not name its network even to ourselves, and the other nine we do not name here. The measurements are public and anyone can derive them, but a network hears from us before its name goes into a note, and two of them have been asked.

What we can say about the cause, which is not much

What is measured is the behaviour: these ten entries stay where the last drain put them, in both directions, for at least the hours we watched and, across the gap between the two experiments, for at least two days. What is not measured is why.

The obvious candidate is a tie-breaker in BGP best-path selection. Both large router vendors prefer, among external routes that are otherwise equal, the one learned first. Cisco's published algorithm puts it as "prefer the path that was received first (the oldest one)" and says the step exists to minimise route flap; Junos has the same step in the same position, "prefer the oldest path, in other words, the path that was learned first". The status of that rule is worth stating precisely: it is not in the standard. RFC 4271 section 9.1.2.2 lists AS path length, origin, MED, external over internal, interior cost, lowest BGP identifier and lowest peer address, and route age appears nowhere in it. The age preference is vendor practice on top, narrowed into a standard only by RFC 5004, which says not to leave the current best external path merely because a newcomer wins the router-identifier tie-break. Where such a rule applies, the route re-announced after a restore is the younger one, and the route the network already holds wins by age until something else changes.

That is a hypothesis and we cannot prove it: we have no RIS peer inside any of these ten networks, so we cannot see which path they hold, and the data plane only says where the answer came from. One competing explanation we can argue against. Route flap damping decays, so a suppressed route is released once the penalty falls under the reuse threshold, which would show up as stragglers returning over the following hour or two; we saw no stragglers at all in six hours. A caching resolver is not a candidate either, because there is none in the path: every probe asked our anycast address directly. A static preference at the operator, or a routing policy at the operator or one of its upstreams that has nothing to do with route age, are both still open.

Three things would settle it, none of which we have done. A third drain with the roles swapped again, following the same ten entries, would show whether the direction flips a second time. A cheap standing measurement, one probe per network every five minutes, would show whether the state holds across days or springs back at some point. And the operators could tell us which rule decides path selection in their network, which is worth asking before guessing.

What this means if you run anycast

After a maintenance window, your own equipment tells you the window is closed. Sessions up, routes exported, prefix visible in the collectors. In our second run all of that was true at 23:50:30 UTC, and at 05:50:30 the next morning 3 of 20 IPv4 probes and 7 of 59 IPv6 probes were still being served by the other site, six hours later, without a single error anywhere in our own logs. There is no view from inside the network that shows this. The withdrawal is visible from inside; the incomplete return is not, because nothing is broken, the traffic simply goes somewhere else that also works.

If you want that number for a real network rather than a lab, it is measurable from outside, before, during and after a planned maintenance, and we can run it.

The lab, its prefixes, etiquette and opt-out are at anycast.org. Both drains were on our own prefix and affected nobody's traffic but our probes'.

Limits

Two sites, one lab prefix per family, cohorts of 20 and 59 probes: the IPv4 list of the ones that stayed away is three probes in three networks, and one probe more or less would move it by five percentage points. The two runs are not cleanly comparable either. The first ran from 17:53 to 19:35 UTC, the second from 23:20 to 05:50 UTC, so the load and the maintenance windows of the transit networks in between were different ones. The first experiment also drained the other site ten minutes earlier, in a catchment it had already changed, while the second had 52 hours of quiet and 20 minutes of baseline in front of it.

The baseline itself was not quiet: in those 20 minutes, 7 of 95 IPv4 probes and 6 of 95 IPv6 probes moved between the two sites without any action from us. They are excluded from the cohorts, which is why "stable at Frankfurt" is a statement about 20 minutes and not about the normal state. The gap between the experiments is not measured at all; 51.4 hours passed without a single query, so what is documented is the state at both ends, not the path between them. Two probes behaved oddly and are reported rather than hidden: one in Egypt answered 145 times with SERVFAIL from nameservers that are not ours and never reached either site, and one in the United States appeared inside both IPv6 measurements with an IPv4 address family and nothing but timeouts, 61 rows, having already failed in the first experiment.

Method, so you can check it

Everything above comes out of one append-only trail, 2026-09-19-fra1, 76,623 entries and 438 checkpoints, hash-chained and signed with a key we published before the run, fingerprint ed25519:70ca7594ff7ebc9e, at /.well-known/logsiegel-pubkey.txt. logsiegel verify <trail> --pubkey logsiegel-pubkey.txt against that published key returns PASS: 76623 entries, 438 checkpoints. 32 RFC 3161 timestamps over checkpoint roots, taken during the run, all verify as OK against DigiCert's chain, which puts the trail in this shape in front of a clock that is not ours. Step-by-step instructions, including how to break the check on purpose so you know it measures something, are at /.well-known/logsiegel-verify.txt. The trail itself is published, log, checkpoints, timestamp tokens and a manifest of file hashes, at /witness/2026-09-19-fra1/; the run of the first experiment this note compares against, the New Jersey drain, is at /witness/2026-09-18-ewr1/, and /witness/ lists every published run with its key.

It holds the declaration of intent, written 1,746 seconds before the local withdraw, so the plan is on record before the event rather than after it; the state of the control plane beforehand, 759 RIS peers carrying the prefix, 343 over IPv4 and 348 over IPv6 across 23 collectors, so the denominator of every percentage is inside the chain; all 27,777 Atlas results live, with no backfill; and 1,767 RIS events, with a cross-check recording that 1,767 of 1,767 rows of our own database made it in and none were missing.

The foreign timestamps can be pulled independently. RIPE Atlas measurements 213363773 (IPv4, 60 s) and 213363774 (IPv6, 60 s) cover the drain and the first 30 minutes after the restore, 213371666 and 213371667 the following five and a half hours at 300 s; the first experiment is 212700284 and 212700285. Those results are public at RIPE, and where our numbers and theirs disagree, theirs are right. That matters most for the first experiment, whose numbers are the weaker half of the comparison: only about a fifth of its Atlas results were written into the run trails while they arrived, the rest were fetched afterwards into a separate trail that we do not publish, because no signing key for it had been published before it was written. The "31 of 31" column above therefore rests on results anyone can pull from measurements 212700284 and 212700285 at RIPE, not on a trail signed under a key published in advance; the published run 2026-09-18-ewr1 holds the part of them that was recorded live. What the chain does not prove is completeness: it protects what went in, not what could have been left out before it went in. Only the foreign times, RIPE's and the timestamping authority's, are outside our control.

What comes next

The measurement that would turn this from a lab curiosity into something an operator can act on is the same drain on a network with more than one transit provider and more than two sites, run by somebody who can then look up what their own routers did.