Skip to content

node: concurrent router dials, ready on the first, redial a dropped router; pace router rollouts - #613

Merged
aojea merged 5 commits into
google:mainfrom
aojea:router-rollout
Oct 9, 2026
Merged

aojea merged 5 commits into
google:mainfrom
aojea:router-rollout

Conversation

@aojea

@aojea aojea commented Oct 9, 2026

Copy link
Copy Markdown
Collaborator

Follow-up to #612, from what its instruments showed on bananas in the first hours.

Findings

  1. Cold join cost the sum of the router dials, plus a full timeout for any router that had gone. Start dialed the configured routers one after another and returned only when every dial had ended. With four live routers a member took 6.05 s to /readyz; while the control plane still listed the previous GCE router (it does for its lease's 15 minutes after every deploy), 7.6 s. The six members of the first harness run all showed the same 6.05 s.
  2. A router rollout took every session a member had. The first rollout under the soak fleet restarted the three routers in 45 s (next pod as soon as the previous passed its 5 s startup probe). The node only re-handshakes a router from the connection monitor, which acts once no router is left and only on its one-minute interval, so all 20 members lost all three routers with no handshake in between and were off the mesh for 7–63 s. A router restarting alone was never re-handshaken at all: a member kept one router fewer after each rollout. SamCanaryMemberDisconnected (5 m) correctly stayed silent.
  3. The deploy failed on a 120 s control-plane rollout wait. The new pod's containers were ready in 15 s, but GKE's NEG readiness gate took minutes (34 s on the previous deploy). The rollout finished on its own; every later step, canaries and the soak fleet included, was skipped until a rerun.

Changes

  • node: dial the routers at once, be ready on the first, and redial one that drops — router addresses are resolved and grouped by peer (a router advertised under several addresses, or by several entries resolving to the same records, is one peer with all of them: one dial over whichever address answers, one handshake; static relays are now 4 peers, not 20 entries). The routers are dialed concurrently and Start returns as soon as one admits the node. A dropped router is redialed in the background with a doubling delay from 2 s, six attempts, one redial per router at a time; RouterRedialDelay / RouterRedialAttempts expose it. Cold join on bananas: 6.05 s → 2.3 s.
  • deploy, chart: hold a router rollout 90s between pods — minReadySeconds: 90 on the router StatefulSet (testnet manifest and router.minReadySeconds in the chart), so a rollout waits for members to re-attach before taking the next router.
  • deploy: give the control plane and router rollouts the time they take — both waits 600 s.

Tests

  • TestStartIsReadyOnceOneRouterAdmitsTheNode: a live router and a listener that accepts and never speaks; Start returns in 20 ms rather than the dead router's 3 s timeout (fails at 3.01 s against the old loop).
  • TestRouterPeersGroupsAddressesByPeer: three entries for one router → one peer, two distinct addresses.
  • TestRouterDropIsRedialed: the router closes the connection, the node re-handshakes on its own within the redial delay, twice in a row (fails without the redial). TestRouterDisconnectClearsAuthenticatedSession opts out of the redial, since it tests the state in between.
  • go test ./... green, make lint clean, router-path tests under -race ×2, each commit builds on its own. Router template dry-run on kind; helm lint clean.

aojea added 3 commits October 9, 2026 13:41
… that drops

Start dialed the configured routers one after another and returned only
when every dial had ended, so a member's time to ready was the sum of
the dials, and a router the control plane still listed but that had gone
(the GCE router after each deploy, for its lease's 15 minutes) added its
whole dial timeout to every cold join. Measured on bananas: 6.05 s to
ready with four live routers, 7.6 s with one gone.

Router addresses are now resolved and grouped by peer first, so a router
advertised under several addresses or by several entries that resolve
to the same records is one peer, dialed once over whichever address
answers and handshaken once; the static relays are the same list, four
peers rather than twenty entries. The routers are then dialed at once
and Start returns as soon as one admits the node, the rest finishing in
the background. ConnectAndAuthWithRouter dials the routers behind one
entry the same way and waits for all of them, which also halves the
enrollment's own router step. Cold join on bananas: 2.3 s.

The first router rollout under the soak fleet showed the other half.
The connection monitor acts only once a member has no router left, and
only on its one-minute interval; the StatefulSet restarted the three
routers in 45 seconds, so every member lost all three with no handshake
in between and was off the mesh for up to a minute. A router restarting
alone was never re-handshaken at all, the member keeping one router
fewer after each rollout. A dropped router is now redialed in the
background with a doubling delay from two seconds, six attempts, one
redial per router at a time; a router that stays away is left to the
monitor as before. RouterRedialDelay and RouterRedialAttempts expose the
schedule; a negative delay turns the redial off.
A router that is ready has not been re-handshaken by anyone yet. Members
redial a dropped router with a backoff that reaches a minute, and older
members only on their monitor interval of one minute, so a rollout that
takes the next router sooner takes every session a member has before
any is back. The first rollout under the soak fleet did exactly that:
three routers in 45 seconds, every member off the mesh for up to a
minute. minReadySeconds: 90 on the StatefulSet, in the testnet manifest
and as router.minReadySeconds in the chart, makes a rollout wait for
members to re-attach before it goes on.
A control plane pod is Ready only once the load balancer's NEG
controller has seen it healthy, which has taken anywhere from half a
minute to several minutes; the 120s wait failed a deploy on bananas
while the rollout itself went on to finish, and everything after the
step, canaries and the soak fleet included, was skipped. The routers
are now held 90s apiece between pods. Both waits are 600s.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a concurrent router dialing mechanism and a background redial loop with exponential backoff to handle dropped router sessions, improving node resilience during router rollouts. It also updates Kubernetes and Helm templates to configure minReadySeconds for routers, and adds comprehensive tests to verify these behaviors. Feedback on the changes suggests adding an early check in ConnectAndAuthWithRouter to handle empty router lists gracefully and avoid returning a nil error, as well as logging failures when loading identity credentials during background redials.

Comment thread internal/node/node.go
Comment thread internal/node/node.go
aojea added 2 commits October 9, 2026 14:13
… without identity

Review follow-ups: ConnectAndAuthWithRouter returned "failed to
authenticate with any router addresses: <nil>" for an address that
resolved to no router peer; it now names the address. A redial that
cannot load the node's identity said nothing; it now logs why it gave
up.
govulncheck flags GO-2026-6603, -6611, -6612 and -6617 in x/net
v0.59.0, all fixed in v0.60.0. Nothing else moves.
@aojea

aojea commented Oct 9, 2026

Copy link
Copy Markdown
Collaborator Author

Addressed both review comments in 68e8aa7 (an address that resolves to no router peer is now named in the error; a redial that cannot load the identity logs why it gave up) and bumped golang.org/x/net to v0.60.0 in d612b64 for the govulncheck findings GO-2026-6603/6611/6612/6617. Full go test ./... green.

@aojea
aojea merged commit 9474b96 into google:main Oct 9, 2026
20 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant