fix(guest-agent): a container can make systemd SIGABRT the guest agent by asking for quotes - #1256
Conversation
get_quote in tdx-attest takes a process-global std::sync::Mutex and then blocks -- on configfs/TSM, or on a vsock connection to the host's QGS with no connect, read or write timeout. Until now exactly one caller treated it as blocking work: Attest on the v1 surface. GetQuote, the v0 Attest, Tappd.TdxQuote and RawQuote, Worker.GetAttestationForAppKey, issue_cert and get_info all ran it straight on the async executor. That parks a runtime worker for the whole quote, and the runtime is narrower than it looks: #[rocket::main] builds it before the configuration is read, so the workers = 8 in dstack.toml never reaches it and a CVM gets one worker per vCPU. Measured on TDX hardware with a 2-vCPU guest, one quote costs 1.02 s, and 24 concurrent GetQuote calls over the bind-mounted /var/run/dstack.sock stalled every other connection the agent served: 20 of 43 were dropped, and Version went from 0.4 ms to seconds. The stall reaches the agent's own systemd watchdog, whose heartbeat is a Worker.Version request to the external listener, so systemd hit WatchdogSec=30s, SIGABRTed the agent and restarted it. Any container in the CVM can do this -- the socket is bind-mounted into application containers by design -- and the same load over the external TCP listener does the same thing. Put the hop where the blocking is, not at each caller: quote_response and attest_cvm take it themselves, and issue_cert and get_info wrap their platform calls, so a new caller cannot forget. identity() stays on the executor, as its comment says: the retry throttle caps it at one quote per interval for the whole process. With the same load on the same guest, all 43 calls are answered, the Version probe stays at 2 ms p95 against a 1.0 s bound, and four consecutive 24-way loads leave the agent in the process it started as.
Full hardware measurement, four transports, before and afterLab Candidate agent, four transports — the drops are not the transport:
Unloaded Controls, same guest, same session:
The first rules out concurrency. The second shows two concurrent quotes already consume the whole runtime. With this branch, same guest, same four transports:
Same pid and start-time tick (4736 / 109370) across all four consecutive loads, Also run through the real acceptance harness against the leased guest: PASS with the fixed binary swapped in via a A regression test, |
End-to-end on a guest image built from this branchThe measurements above swapped the fixed binary into a running guest with a systemd drop-in. This is the same case, through the acceptance runner, against a guest image built by mkosi from this branch's source — so the agent under test is the shipped one, not a hot-patched one:
Against an image built from
|
|
Follow-up commit a193b5e fixes the residual startup-identity failure path: later anonymous Info retries are single-flight and run on spawn_blocking; cancellation retains the guard. |
Route every platform call through `AppStateInner::attest`, which takes an owned `tokio::sync::Mutex` guard and moves it into the `spawn_blocking` task. `dstack-attest` already serialises quotes under `QUOTE_LOCK`, so this costs no throughput, but a flood of callers now waits in the async queue instead of each parking one of up to 512 blocking-pool threads on that lock. The same lock gives `identity()` single-flight retries and cancellation safety for free, so the dedicated `identity_refresh` mutex goes. The throttle check moves ahead of the queue, so a throttled `Info` is rejected without waiting behind in-flight quotes. `decode_identity` takes the platform directly and `info_attestation` is dropped. The quote regression test reuses the probe gate instead of a separate `BlockingPlatform` fixture.
OnceCell::get_or_try_init already gives single-flight initialisation with retry on failure, so the hand-rolled cache check inside the attest closure and the cached_identity() split are no longer needed. The identity is written once and never changes, so it is returned by reference instead of through an Arc. Drop the single-flight test and the Info half of the probe gate; the behaviour it pinned is now OnceCell's.
Found by running the new
tc-gos-concurrency-002acceptance case on real Intel TDX hardware, and then chasing down what the drops actually were.What happens
tdx_attest::get_quotetakes a process-globalstd::sync::Mutexand then blocks — on configfs/TSM, or on a vsock connection to the host's QGS with no connect, read or write timeout. Exactly one caller treated it as blocking work:Atteston the v1 surface, whose comment names the problem precisely:GetQuote, the v0Attest,Tappd.TdxQuote,Tappd.RawQuote,Worker.GetAttestationForAppKey,issue_certandget_infoall ran it straight on the async executor.And the runtime is narrower than it looks.
#[rocket::main]builds the runtime before the configuration is read, so theworkers = 8indstack.tomlnever reaches it — a CVM gets one worker per vCPU. A 2-vCPU guest was measured running tworocket-workerthreads.Measured on TDX hardware
One quote costs 1.02 s on a 2-vCPU guest. 24 concurrent
GetQuotecalls over the bind-mounted/var/run/dstack.sock:Versionwent from 0.4 ms to seconds.And the stall reaches the agent's own watchdog. The systemd heartbeat is an HTTP
Worker.Versionrequest the agent sends to its own external listener, so a runtime parked on the quote mutex cannot answer it:systemd kills the agent and restarts it, and every open connection is reset. Any container in the CVM can do this — the socket is bind-mounted into application containers by design — and the same load over the external TCP listener does the same thing.
Ruling out the transport
The first observation came through a QEMU slirp forwarded host port, which has its own limits, so the same 24-way load was driven four ways against one lease-owned guest: over the forwarded port, over
/var/run/dstack.sockfrom inside the guest, over the agent's own TCP listener, and through the fixture's socket bridge. 15 to 21 of 43 dropped every time.Control: 24-way
Versionunder the same conditions answered 579 405 calls at a 2 ms p95. So it is the blocking quote, not the concurrency, and not the transport.The fix
Every platform call goes through
AppStateInner::attest: it takes an ownedtokio::sync::Mutexguard, moves it into aspawn_blockingtask, and runs the call there.quote_response,attest_cvm,issue_cert,get_infoandidentity()all use it, so a new caller cannot forget.dstack-attestalready serialises quotes underQUOTE_LOCK, so the async lock costs no throughput. A flood of callers waits in the async queue instead of each holding a blocking-pool thread.tokio::sync::OnceCell, so a retry after a failed boot decode is single-flight: queued callers get the cached result, or the throttle error, without another quote.After
Same load, same guest: all 43 calls answered, the
Versionprobe stays at 2 ms p95 against a 1.0 s bound, and four consecutive 24-way loads leave the agent in the process it started as.cargo test -p dstack-guest-agent: 147 passed.cargo clippy -p dstack-guest-agent -- -D warnings --allow unused_variables: clean.Worth a separate look
#[rocket::main]building the runtime beforeload_config_figmentmeans[core] workersindstack.tomlis silently inert. This PR does not change that — it removes the reason it mattered — but the configuration key is currently a lie, and either it should take effect or it should go.