impl connection recovery

This commit is contained in:
iceBear67
2026-08-15 17:31:35 +08:00
parent e63a34d53a
commit 7bd84af48d
33 changed files with 3858 additions and 200 deletions
+89 -18
View File
@@ -114,21 +114,34 @@ per-connection counter with no coordination.
### 3.1 Pool & allocation
The client keeps 1…`maxConn` worker connections and places each new stream on
the **least-loaded** one. It opens an additional connection only when the
least-loaded connection is *saturated* (more than 8 active streams) and the pool
is below `maxConn`:
the **least-loaded** one. The pool grows **breadth-first**: it dials out to
`maxConn` before stacking streams, so that no single TCP connection ever becomes
the shared point of failure for every player on the tunnel (PROTOCOL.md §7.1):
```
pick least-loaded conn
if leastLoaded.streams > 8 and pool.size < maxConn:
dial a new worker conn and use it
else:
use leastLoaded
pick least-loaded conn; use it
if leastLoaded.streams >= 1 and pool.size + dialsInFlight < maxConn:
dial another worker conn in the background # the stream just placed does not wait
```
The new connection becomes the least-loaded one and picks up subsequent streams.
Once the pool is at `maxConn`, streams stack on the least-loaded connection;
going past 8 active streams there is logged as pool saturation but is not an
error.
Two properties of the dialing path matter as much as the placement rule:
* A dial is **never performed while holding the pool lock** — session
establishment is network I/O, and one unresponsive hub must not park every
other player behind it.
* Only when the pool is *empty* does a caller dial synchronously, and then
exactly one caller dials while the others wait on its result, so a burst of
arrivals cannot open a burst of redundant connections.
The e2e test `TestConcurrentStreamsUseMultipleConns` drives 20 simultaneous
streams with `maxConn=4` and observes them deterministically spread over 3
connections (9 + 9 + 2), confirming the algorithm.
streams with `maxConn=4` and confirms they spread over more than one connection
without exceeding the cap; `TestAllocateDoesNotWedgePoolOnStalledHub` covers the
stalled-dial path.
## 4. Encryption
@@ -162,7 +175,7 @@ connections (9 + 9 + 2), confirming the algorithm.
## 6. Back-pressure & flow control
Two mechanisms operate at different granularities:
Three mechanisms operate at different granularities:
* **Per-stream credit windows** (PROTOCOL.md §7.3; the windows are exchanged
at session establishment): each stream direction has an independent byte
@@ -178,10 +191,44 @@ Two mechanisms operate at different granularities:
socket itself is congested (total bandwidth, not one stream), the hub parks
all sending players until it drains, and the client's `WriteFrame` blocks.
This is fair — when the pipe is genuinely full, everyone should slow down.
* **Client egress shaping** (optional, `maxBandwidth`; `client/shaper.go`): a
rate cap on everything the client sends to the hub, across all worker conns.
The first two mechanisms have no time dimension. A credit window bounds how many
bytes are *in flight*, and TCP back-pressure only reacts once the pipe is already
full — which on a residential uplink is too late. One player loading chunks fills
the line, the standing queue grows to seconds, and every other player's keepalive
times out. Nothing in §7.3 prevents that: each stream is individually
well-behaved, and collectively they still overrun the link.
The shaper closes that gap with a token bucket for the rate and start-time fair
queueing for the split. A global virtual clock advances with each grant; every
stream remembers where its last request finished, and a new request is stamped
`max(stream.vfinish, vclock)`. Lowest stamp wins. A stream that keeps sending
pushes its own stamp further out and yields; a stream returning from idle is
clamped back to the clock, so it cannot bank credit for time it did not use, but
is not penalised for the idleness either. A stream sending a few hundred bytes
gets a nearer stamp than one sending a full chunk, so keepalives and chat overtake
bulk terrain data for free. One stream alone still gets the entire rate.
Two details keep bursts cheap. The bucket banks 200 ms of transmission, so a
player joining spends it at once instead of paying for the cap in visible
chunk-loading latency. And the DATA chunk shrinks to ~20 ms of transmission when
the rate is low (floor 4 KiB), because a fixed 32 KiB chunk is a 256 ms slot at
1 Mbps — long enough dead air to drag the other players towards the very timeout
the cap exists to prevent. Above ~13 Mbps the chunk stays at the usual 32 KiB.
This is entirely client-local: nothing about it appears on the wire, and the hub
is unaware. Only DATA is shaped — delaying a `FIN`, `WND` or `PONG` would cause
the false-death detection §7.4 exists to avoid.
The window also bounds memory: a stream can hold at most one window of
undelivered data per direction (the client's pre-connect handshake buffer is
covered by the same bound).
covered by the same bound). With stream resumption enabled (§7) the *sender*
holds a second window — the bytes it has sent but the peer has not yet credited,
kept so they can be retransmitted after an outage. That is not a new bound so
much as the existing one made symmetric: the region is exactly what flow control
already declared outstanding, which is why resumption needs no cap of its own.
Per-stream flow control is mandatory: the hub rejects a session whose Rekey
lacks the STREAM_FC flag, and the client rejects a hub that does not echo it —
@@ -194,17 +241,41 @@ transport (QUIC) would be the escape hatch if it ever matters.
## 7. Failure & recovery
* **Control session drop:** the client reconnects with capped exponential
backoff and re-registers all patterns. Existing worker connections and their
live streams are unaffected.
* **Worker connection drop:** every stream on it is torn down (destinations
closed); the hub closes the corresponding player sockets; the client removes
the connection from the pool and will dial a fresh one on the next allocation.
* **Control session drop:** the client retries immediately, then backs off to a
10s cap, and re-registers all patterns. Existing worker connections and their
live streams are unaffected — they ride worker conns, which a control-session
close never touches. The hub meanwhile keeps that session's routes as
*orphaned* for `registrationGraceMs` (PROTOCOL.md §5.2) and **holds** players
arriving on them instead of refusing them, replaying the control request once
a client re-registers the pattern. Without that, the reconnect window is one
in which every new player is told there is no such server.
* **Worker connection drop:** the connection leaves the pool either way. What
happens to its streams depends on whether STREAM_RESUME was negotiated:
* *without it* — every stream is torn down (destinations closed) and the hub
closes the corresponding player sockets, as it always did;
* *with it* — the streams are **hung** instead (PROTOCOL.md §7.5). The
destination sockets stay open, the hub pauses the player sockets and holds
them for its grace period, and the client reattaches each stream over a
freshly dialed connection, replaying byte-exactly from the offset the peer
reports. Players see a stall rather than a disconnect.
* **Pending timeout:** if no worker takes over a matched player within
`pendingTimeoutMs`, the hub drops the pending entry and closes the player.
* **Bad PSK / bad timestamp / bad magic:** the hub closes the TCP connection;
the client's session establishment fails fast.
Resumption is worth the machinery because a worker connection is only the
*middle* leg of every stream it carries. When it dies both terminal sockets are
usually still healthy, so the old behaviour discarded working connections
because a replaceable transport failed — one conntrack expiry disconnected every
player sharing that connection. It also has to be byte-exact rather than
best-effort: bytes handed to a dying socket are lost with no notification and the
frame cipher cannot be resynchronized, so an approximate reattach would splice
the tunneled protocol mid-packet, which is worse than a clean close.
Notably, resumption does not depend on the control session. A blip usually kills
both, and the reattach path needs only a worker connection, so recovery does not
wait on the control reconnect backoff.
## 8. Known limitations
1. No AEAD — payload integrity/authenticity is not cryptographically guaranteed.