292 lines
17 KiB
Markdown
292 lines
17 KiB
Markdown
# redapricot architecture
|
|
|
|
This document explains *how* redapricot is built and *why*. For the exact bytes
|
|
on the wire, read [PROTOCOL.md](../PROTOCOL.md).
|
|
|
|
## 1. Roles and topology
|
|
|
|
```
|
|
┌───────────────────────── public internet ─────────────────────────┐
|
|
│ │
|
|
┌──────────┐ MC handshake (Intent 2/…) ┌───────────────┐ │
|
|
│ Player │ ───────────────────────────────▶│ │ │
|
|
└──────────┘ raw Minecraft bytes │ Hub │ │
|
|
│ (Java/Vert.x)│ │
|
|
┌──────────┐ Intent 17, magic 0x01 │ │ │
|
|
│ Client │ ◀──────── control session ──────│ • pattern reg │ │
|
|
│ (Go) │ ────────────────────────────────│ • CID table │ │
|
|
│ │ Intent 17, magic 0x02 │ • mux demux │ │
|
|
│ │ ═════════ worker conns ═════════│ │ │
|
|
└──────────┘ multiplexed player streams └───────────────┘ │
|
|
│ │
|
|
▼ MC bytes (+ optional HAProxy v2) │
|
|
┌───────────────┐ │
|
|
│ Real MC server│ (behind NAT, next to the client) │
|
|
└───────────────┘ │
|
|
```
|
|
|
|
Everything reaches the hub on **one TCP port**. The hub distinguishes three
|
|
kinds of inbound connection purely from the first Minecraft **Handshake**:
|
|
|
|
| Handshake `Intent` | Handled as |
|
|
|--------------------|------------|
|
|
| `17` + magic `0x01` | a **control session** from a client |
|
|
| `17` + magic `0x02` | a **worker connection** from a client |
|
|
| `18` | reserved (management/status) — never treated as a player |
|
|
| anything else | a **player** to be pattern-matched and tunneled |
|
|
|
|
Because players use ordinary intents (`1` status, `2` login, `3` transfer),
|
|
**vanilla clients need no changes**.
|
|
|
|
## 2. Connection lifecycle
|
|
|
|
### 2.1 Client establishes a control session
|
|
|
|
```
|
|
Client Hub
|
|
│ TCP connect │
|
|
│─ Handshake(Intent=17, addr=hex(SHA3-224(PSK))) ─▶ verify addr == expected
|
|
│ │
|
|
│ (both derive Phase-A keys = ChaCha20(SHA3-256(PSK ‖ dir)))
|
|
│─ Frame#1 [magic=0x01, rand, ts] ──────▶ check |now-ts| ≤ window
|
|
│ (both switch to Phase-B keys = ChaCha20(SHA3-256(rand‖ts ‖ dir)))
|
|
│◀──────────── Frame [SessionReady] ─────│
|
|
│─ Register("mc\.example\.com") ────────▶ patterns["mc\.example\.com"] = (regex, session)
|
|
│◀──────────── RegisterAck ──────────────│
|
|
│ ... periodic Ping/Pong ... │
|
|
```
|
|
|
|
Only frame #1 is encrypted with the PSK-derived key; a fresh random `rand‖ts`
|
|
becomes the per-connection key for everything after, so two connections never
|
|
share a keystream beyond that first frame.
|
|
|
|
### 2.2 A player arrives and is tunneled
|
|
|
|
```
|
|
Player Hub Client Destination
|
|
│─ Handshake(addr="mc.example.com", Intent=2)─▶ normalize + regex-match
|
|
│ │ pause player socket,
|
|
│ │ buffer bytes, mint CID
|
|
│ │─ ControlRequest(CID, pattern, ip:port) ─▶
|
|
│ │ allocate worker+stream
|
|
│ │◀──────── SYN(streamId, CID) ────────────│
|
|
│ │ takePending(CID) → bind dial destination,
|
|
│ │ forward buffered bytes write HAProxy v2 hdr
|
|
│ │─ DATA(streamId, handshake…) ───▶ ── handshake ──▶│
|
|
│ resume ─────────────────────────────│ bridge stream ⇄ dest
|
|
│══════════════ player bytes ══ DATA ══▶│════ DATA ═══▶ dest.write │
|
|
│◀═══════════ dest bytes ═══ DATA ══════│◀═══ DATA ════ dest.read │
|
|
│ player closes ──────────────────────│─ FIN(streamId) ────────▶ close dest │
|
|
```
|
|
|
|
Key points:
|
|
|
|
* **Patterns are regexes.** Each registered pattern is a case-insensitive
|
|
regular expression, matched against the *whole* normalized hostname (anchored,
|
|
first match wins). The hub echoes the **matched pattern string** — not the
|
|
player's hostname — in `ControlRequest`, so the client can look it straight up
|
|
in its own route table. Invalid patterns are rejected at registration with a
|
|
non-zero `RegisterAck` status.
|
|
* The hub **pauses** the player socket the instant it matches, so no player
|
|
bytes are lost while the takeover is arranged; the buffered handshake is
|
|
forwarded **verbatim**, so the real server sees exactly what the player sent
|
|
(including the original hostname — used for virtual-host routing there).
|
|
* **CID** is 16 random bytes minted by the hub and delivered only over the
|
|
encrypted control session, so only the intended client learns it. Any worker
|
|
connection presenting the correct CID is allowed to take over — that secrecy
|
|
is what binds a worker stream to the right pending player without any explicit
|
|
client identity.
|
|
* Disconnects are symmetric: player-close → hub sends `FIN` → client closes the
|
|
destination; destination-close → client sends `FIN` → hub closes the player.
|
|
|
|
## 3. Multiplexing (worker connections)
|
|
|
|
A worker connection is one encrypted TCP link carrying many **streams**. The
|
|
frame is intentionally tiny (PROTOCOL.md §7):
|
|
|
|
```
|
|
[plaintext VarInt length][ FrameType u8 | StreamID VarInt | Data… ] (payload encrypted)
|
|
```
|
|
|
|
Only the client opens streams (`SYN`), so stream-id allocation is a simple
|
|
per-connection counter with no coordination.
|
|
|
|
### 3.1 Pool & allocation
|
|
|
|
The client keeps 1…`maxConn` worker connections and places each new stream on
|
|
the **least-loaded** one. The pool grows **breadth-first**: it dials out to
|
|
`maxConn` before stacking streams, so that no single TCP connection ever becomes
|
|
the shared point of failure for every player on the tunnel (PROTOCOL.md §7.1):
|
|
|
|
```
|
|
pick least-loaded conn; use it
|
|
if leastLoaded.streams >= 1 and pool.size + dialsInFlight < maxConn:
|
|
dial another worker conn in the background # the stream just placed does not wait
|
|
```
|
|
|
|
The new connection becomes the least-loaded one and picks up subsequent streams.
|
|
Once the pool is at `maxConn`, streams stack on the least-loaded connection;
|
|
going past 8 active streams there is logged as pool saturation but is not an
|
|
error.
|
|
|
|
Two properties of the dialing path matter as much as the placement rule:
|
|
|
|
* A dial is **never performed while holding the pool lock** — session
|
|
establishment is network I/O, and one unresponsive hub must not park every
|
|
other player behind it.
|
|
* Only when the pool is *empty* does a caller dial synchronously, and then
|
|
exactly one caller dials while the others wait on its result, so a burst of
|
|
arrivals cannot open a burst of redundant connections.
|
|
|
|
The e2e test `TestConcurrentStreamsUseMultipleConns` drives 20 simultaneous
|
|
streams with `maxConn=4` and confirms they spread over more than one connection
|
|
without exceeding the cap; `TestAllocateDoesNotWedgePoolOnStalledHub` covers the
|
|
stalled-dial path.
|
|
|
|
## 4. Encryption
|
|
|
|
* **Cipher:** ChaCha20 (RFC 8439) as a raw stream cipher over frame *payloads*.
|
|
The length prefix is plaintext, which makes the cipher **phase switch** at
|
|
rekey trivial (a reader always knows exactly how many ciphertext bytes belong
|
|
to the current frame and never decrypts the next frame with the wrong key).
|
|
* **Keys:** `SHA3-256(phaseKey ‖ 0x01)` for client→server and
|
|
`SHA3-256(phaseKey ‖ 0x02)` for server→client. Distinct per-direction keys
|
|
with a fixed zero nonce avoid a two-time pad without nonce management.
|
|
* **Interop:** Java's JCE `ChaCha20` and Go's `x/crypto/chacha20` produce byte-
|
|
identical keystreams (including across partial-block, arbitrarily-split
|
|
writes), and both `crypto/sha3` implementations agree — verified directly and
|
|
pinned by unit tests on both sides against a shared SHA3-224 vector.
|
|
|
|
## 5. Threading model
|
|
|
|
* **Hub:** a single Vert.x verticle instance. All accepted connections are
|
|
handled on that verticle's one event loop, so the pattern registry, CID table,
|
|
and per-connection state are touched by a single thread — no locks on the hot
|
|
path (concurrent maps are used only defensively). Every socket operation is
|
|
non-blocking; crypto is CPU-cheap. This trades multi-core scaling for
|
|
simplicity and correctness.
|
|
* **Client:** goroutine-per-concern. One goroutine reads each connection
|
|
(control or worker); `WriteFrame` is mutex-serialized so many stream goroutines
|
|
can share a worker connection safely. Each stream has two goroutines: `run`
|
|
pumps destination → hub, and `writeLoop` is the only writer to the
|
|
destination, draining a per-stream queue fed by the worker readLoop. The
|
|
readLoop itself never writes to a destination, so a stalled destination can
|
|
never block frame dispatch for other streams.
|
|
|
|
## 6. Back-pressure & flow control
|
|
|
|
Three mechanisms operate at different granularities:
|
|
|
|
* **Per-stream credit windows** (PROTOCOL.md §7.3; the windows are exchanged
|
|
at session establishment): each stream direction has an independent byte
|
|
budget equal to the receiver's advertised window (default 256 KiB). A sender
|
|
that exhausts a
|
|
stream's window pauses *only that stream's source* — the hub pauses the one
|
|
player socket, the client parks the one destination-reader goroutine. Credit
|
|
is granted back (`WND` frames, batched at half-window) as bytes are actually
|
|
written to the terminal socket. The result: a slow player or slow destination
|
|
jams its own stream at a bounded buffer size and nothing else. This is what
|
|
eliminates head-of-line blocking between streams.
|
|
* **Aggregate TCP back-pressure** on each worker connection: when the shared
|
|
socket itself is congested (total bandwidth, not one stream), the hub parks
|
|
all sending players until it drains, and the client's `WriteFrame` blocks.
|
|
This is fair — when the pipe is genuinely full, everyone should slow down.
|
|
* **Client egress shaping** (optional, `maxBandwidth`; `client/shaper.go`): a
|
|
rate cap on everything the client sends to the hub, across all worker conns.
|
|
|
|
The first two mechanisms have no time dimension. A credit window bounds how many
|
|
bytes are *in flight*, and TCP back-pressure only reacts once the pipe is already
|
|
full — which on a residential uplink is too late. One player loading chunks fills
|
|
the line, the standing queue grows to seconds, and every other player's keepalive
|
|
times out. Nothing in §7.3 prevents that: each stream is individually
|
|
well-behaved, and collectively they still overrun the link.
|
|
|
|
The shaper closes that gap with a token bucket for the rate and start-time fair
|
|
queueing for the split. A global virtual clock advances with each grant; every
|
|
stream remembers where its last request finished, and a new request is stamped
|
|
`max(stream.vfinish, vclock)`. Lowest stamp wins. A stream that keeps sending
|
|
pushes its own stamp further out and yields; a stream returning from idle is
|
|
clamped back to the clock, so it cannot bank credit for time it did not use, but
|
|
is not penalised for the idleness either. A stream sending a few hundred bytes
|
|
gets a nearer stamp than one sending a full chunk, so keepalives and chat overtake
|
|
bulk terrain data for free. One stream alone still gets the entire rate.
|
|
|
|
Two details keep bursts cheap. The bucket banks 200 ms of transmission, so a
|
|
player joining spends it at once instead of paying for the cap in visible
|
|
chunk-loading latency. And the DATA chunk shrinks to ~20 ms of transmission when
|
|
the rate is low (floor 4 KiB), because a fixed 32 KiB chunk is a 256 ms slot at
|
|
1 Mbps — long enough dead air to drag the other players towards the very timeout
|
|
the cap exists to prevent. Above ~13 Mbps the chunk stays at the usual 32 KiB.
|
|
|
|
This is entirely client-local: nothing about it appears on the wire, and the hub
|
|
is unaware. Only DATA is shaped — delaying a `FIN`, `WND` or `PONG` would cause
|
|
the false-death detection §7.4 exists to avoid.
|
|
|
|
The window also bounds memory: a stream can hold at most one window of
|
|
undelivered data per direction (the client's pre-connect handshake buffer is
|
|
covered by the same bound). With stream resumption enabled (§7) the *sender*
|
|
holds a second window — the bytes it has sent but the peer has not yet credited,
|
|
kept so they can be retransmitted after an outage. That is not a new bound so
|
|
much as the existing one made symmetric: the region is exactly what flow control
|
|
already declared outstanding, which is why resumption needs no cap of its own.
|
|
|
|
Per-stream flow control is mandatory: the hub rejects a session whose Rekey
|
|
lacks the STREAM_FC flag, and the client rejects a hub that does not echo it —
|
|
peers that predate the mechanism cannot connect at all.
|
|
|
|
What remains (by design) is TCP-level head-of-line blocking: a lost packet on
|
|
a worker connection stalls all its streams for one retransmit. That is inherent
|
|
to mux-over-TCP; the connection pool is the mitigation, and a datagram
|
|
transport (QUIC) would be the escape hatch if it ever matters.
|
|
|
|
## 7. Failure & recovery
|
|
|
|
* **Control session drop:** the client retries immediately, then backs off to a
|
|
10s cap, and re-registers all patterns. Existing worker connections and their
|
|
live streams are unaffected — they ride worker conns, which a control-session
|
|
close never touches. The hub meanwhile keeps that session's routes as
|
|
*orphaned* for `registrationGraceMs` (PROTOCOL.md §5.2) and **holds** players
|
|
arriving on them instead of refusing them, replaying the control request once
|
|
a client re-registers the pattern. Without that, the reconnect window is one
|
|
in which every new player is told there is no such server.
|
|
* **Worker connection drop:** the connection leaves the pool either way. What
|
|
happens to its streams depends on whether STREAM_RESUME was negotiated:
|
|
* *without it* — every stream is torn down (destinations closed) and the hub
|
|
closes the corresponding player sockets, as it always did;
|
|
* *with it* — the streams are **hung** instead (PROTOCOL.md §7.5). The
|
|
destination sockets stay open, the hub pauses the player sockets and holds
|
|
them for its grace period, and the client reattaches each stream over a
|
|
freshly dialed connection, replaying byte-exactly from the offset the peer
|
|
reports. Players see a stall rather than a disconnect.
|
|
* **Pending timeout:** if no worker takes over a matched player within
|
|
`pendingTimeoutMs`, the hub drops the pending entry and closes the player.
|
|
* **Bad PSK / bad timestamp / bad magic:** the hub closes the TCP connection;
|
|
the client's session establishment fails fast.
|
|
|
|
Resumption is worth the machinery because a worker connection is only the
|
|
*middle* leg of every stream it carries. When it dies both terminal sockets are
|
|
usually still healthy, so the old behaviour discarded working connections
|
|
because a replaceable transport failed — one conntrack expiry disconnected every
|
|
player sharing that connection. It also has to be byte-exact rather than
|
|
best-effort: bytes handed to a dying socket are lost with no notification and the
|
|
frame cipher cannot be resynchronized, so an approximate reattach would splice
|
|
the tunneled protocol mid-packet, which is worse than a clean close.
|
|
|
|
Notably, resumption does not depend on the control session. A blip usually kills
|
|
both, and the reattach path needs only a worker connection, so recovery does not
|
|
wait on the control reconnect backoff.
|
|
|
|
## 8. Known limitations
|
|
|
|
1. No AEAD — payload integrity/authenticity is not cryptographically guaranteed.
|
|
2. TCP-level head-of-line blocking within a worker connection (lost packets;
|
|
see §6) — per-stream flow control removes the application-level variant only.
|
|
3. Single-event-loop hub (see §5) bounds throughput to one core.
|
|
4. `Intent 18` is reserved but only stubbed (the hub logs and closes).
|
|
5. Pattern ownership is last-writer-wins; two clients registering the identical
|
|
pattern string will silently reassign it. Overlapping-but-distinct regexes are
|
|
both kept, and when several match one hostname the winner is unspecified.
|
|
|
|
These are deliberate scope choices for a connectivity-focused P2P tool, not
|
|
oversights; each is a small, well-isolated change away from being hardened.
|