Track in-flight splices for failure reporting and crash recovery - #1080
Draft
jkczyz wants to merge 20 commits into
Draft
Track in-flight splices for failure reporting and crash recovery#1080jkczyz wants to merge 20 commits into
jkczyz wants to merge 20 commits into
Conversation
Wallet sync resolves a funding payment's id for any transaction linked to the record through its conflicting txids, and then adopted that transaction's txid and confirmation outright. A cooperative close conflicts with a pending splice in exactly that way: the splice record would report the close's txid and confirmation under its InteractiveFunding type and contribution figures and graduate as if the splice had confirmed, while the close's own record never received its confirmation. Adopt a transaction only when it is part of the payment's funding history — the record's current txid or a classified candidate. Anything else is recorded under its own txid-keyed id, which also delivers the close's confirmation to the close's own record. Generated with assistance from Claude Code. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A queued broadcast whose payment-record classification failed was dropped outright, on the theory that broadcasting a transaction we failed to record would leave it on-chain without a payment. For interactive funding that theory doesn't hold: the counterparty broadcasts the same transaction once the signature exchange completes, so dropping the package keeps nothing off-chain — it only guarantees the round is never recorded as a candidate on our side. The funding-status ownership gate then treats the round's confirmation as foreign to the funding record and re-keys it to a stray duplicate record, which shadows the funding record's txid lookups permanently: the splice payment stays Pending forever while an untyped duplicate holds the confirmation. Keep the package alive instead: requeue it after a short delay and retry classification until it succeeds, holding the broadcast back the whole time. Classification failures are persistence failures, so the retry is unbounded — a store that never recovers keeps the node from functioning anyway — and every failed round is logged. Generated with assistance from Claude Code. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The retry test slept a fixed three seconds and assumed classification had failed by then; if writes were re-enabled before the first attempt, the test would pass without any retry happening. Count failed writes in FailSwitchStore and wait for one before re-enabling writes. Also fix the test's store reads to use list_page: the payment store's cache is bounded, so list_filter is unavailable, and this commit did not compile its tests standalone (the conversion had landed in the following commit). Implemented with Claude Code. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The retry for a failed classification was a detached tokio::spawn that outlived the node. Its comment claimed a re-send after shutdown would fail because the queue had closed, but the queue receiver lives in the broadcaster and is only dropped with the node, so the re-send succeeded and a stale package would be classified and broadcast after a stop()/start() cycle. Queue failed packages inside the broadcast loop instead and retry them from a timer branch of the same select. New packages keep flowing while a retry waits, and pending retries are dropped when the loop stops. Implemented with Claude Code. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A queued classification can retry after a newer candidate of the same funding already classified. The retry carries the candidate history as of its own broadcast, so applying it rotated the record's txid back to the older candidate and shrank the stored candidate history — after which wallet sync could no longer map the newer transaction to the record and would file it as a foreign duplicate. A fresh interactive-funding classification always carries the record's current txid in its history, so one that doesn't is stale: ignore it, and never let a candidate-history update drop stored candidates. Implemented with Claude Code. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
LDK re-broadcasts pending claims every 30 seconds (and sweeps once per block) until they confirm, so while the payment store is unavailable, the list of pending retries accumulated a copy per rebroadcast — memory, retry load on the struggling store, and a duplicate broadcast burst on recovery all growing with the outage's duration. A package whose transactions already await a retry is not queued again, and the rest are bounded: at the bound, the oldest waiting non-funding package is dropped to make room — its transactions return with LDK's next periodic rebroadcast — but never a funding package, whose transaction would be left confirming without a recorded candidate. Fee-bumped rebroadcast variants carry new txids, so the bound, not the dedup, is what limits their accumulation. Implemented with Claude Code. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
👋 Hi! I see this is a draft PR. |
jkczyz
force-pushed
the
2026-08-splice-tracking
branch
from
September 4, 2026 04:16
aba7686 to
80a1b27
Compare
This was referenced Sep 4, 2026
Since declining to adopt a conflicting close's confirmation, a funding payment whose transaction was double-spent stayed Pending forever -- nothing wrote a terminal status for an on-chain record -- and the sync loop kept re-queueing the dead transaction for rebroadcast on every tip change. Mark such a record Failed once a conflict from outside its candidate history has confirmed through ANTI_REORG_DELAY while neither its own transaction nor any RBF candidate can still confirm, mirroring the anti-reorg finality the Succeeded transition already assumes. Removing the payment's pending entry then stops the re-queueing. Settling also removes the entry that maps candidate txids to the record, so a later wallet event for a dead candidate falls back to keying by that candidate's txid -- which, for the first candidate, is the record's own id. Skip such events rather than let the generic handling resurrect the settled record, and let a replayed replacement event finish an entry removal a crash interrupted instead of stamping the terminal status into the leftover entry. Implemented with Claude Code. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
jkczyz
force-pushed
the
2026-08-splice-tracking
branch
from
September 4, 2026 17:09
80a1b27 to
ca4e5fd
Compare
Wallet sync can learn of a splice transaction before broadcast-time classification records it: once tx_signatures are exchanged, the counterparty may broadcast first, and sync then files the round under a duplicate record keyed by its txid, which shadows the funding record's txid lookups from then on. Retrying a failed classification only narrows that window: a round the counterparty broadcasts is still observed before our record exists. Record the funding payment while handling FundingTransactionReadyForSigning, before funding_transaction_signed hands our signatures to LDK. The counterparty cannot broadcast without them, so the record precedes anything wallet sync can observe, and every later observer resolves to it. The record is written from the channel's pending splice history -- the same history LDK later hands the broadcaster, under the same id -- so the round's broadcast-time classification has nothing left to write but the fact of the broadcast. If the record cannot be written, the event is replayed rather than proceeding unrecorded: LDK re-offers it in-session and regenerates it across restarts while the transaction remains unsigned. A failed write leaves no half-written record behind for the replayed event to build on. Should undoing it fail as well, the replayed event removes what was left of a first round once the round is gone from the channel's history; the leftovers of a bump live under an earlier round's record, which wallet sync moves on as that round confirms or fails. Recording before the round is negotiated means a recorded round can still be abandoned: the counterparty may abort after we sign but before its commitment_signed, or the channel may close, and until LDK has released our signatures nothing can ever broadcast the transaction. Left in place, the record would wait forever on a payment nothing can confirm. The signed round is therefore marked as awaiting broadcast until its classification clears the mark, and a marked round is dropped once LDK no longer holds it, unless the wallet has seen its transaction: the counterparty may broadcast a round it received our signatures for while LDK still waits on its own. A round whose classification has run keeps its place whether or not wallet sync has seen it yet, and so does the channel's current funding: a zero-conf splice becomes the funding as soon as splice_locked is exchanged, before its transaction confirms or its classification has necessarily run. Dropping a round leaves the record on the last remaining round this node contributed to, moving it there if it still names the dropped round, or removes the record when none remains. LDK's view is consulted when it reports the failed negotiation of a channel it still lists, when the channel closes -- a round awaiting the counterparty's signatures gets no failure report then, and a failure reported once the channel is gone is left to this report, which carries the channel's last funding -- and at startup, before any background task runs: LDK reports the loss of a negotiation its last channel manager write carried mid-way, but a round committed, negotiated and signed since that write gets no report if the node stops before the next one. A round already missing from the channel's history when the signing event is handled is not recorded at all. Rounds without a local contribution emit no signing event and are left to broadcast-time classification, as before. Developed with assistance from Claude Code. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Keep, at ChannelClosed, the splice rounds the channel's monitor still watches. The channel manager forgets a pending round with the channel and reports no failed negotiation for one awaiting the counterparty's signatures, so the handler took back every recorded round but the last funding. The monitor, however, watches every round from the counterparty's commitment_signed on, and our signatures cannot have left the node before that message: the counterparty may hold a fully signed transaction and broadcast it, in which case the dropped record resurfaced as an untyped payment under the transaction's id, or a dropped bump marked the splice Failed when it confirmed. Reachable when this node's contribution is the smaller one -- a queued contribution merged into a counterparty-initiated round -- and on LDK's own force-close after a tx_abort that follows its commitment, as well as at reload when the monitor is ahead of the channel manager. The kept set is every round the monitor watches, not only those our signatures left for: the monitor cannot tell them apart, and a watched round still marked at close is in either case one whose counterparty signatures never arrived, since with both signature sets held LDK broadcasts the round and its classification clears the mark. Such a round stays a Pending record until the close spend matures and the monitor's DiscardFunding arrives; that handler only reclaims addresses today, so nothing terminates the record yet (pre-existing; the following commit adds that). A round is marked at signing, which LDK triggers at tx_complete, before the counterparty's commitment_signed, so a marked round that message never reached is still dropped; our signatures cannot have left for it. After a zero-conf lock the monitor stops watching the rounds the lock superseded -- an RBF sibling of the locked round and the previous funding scope alike -- so a later close still drops those, which matters only while their classification is queued. Two integration tests drive the handler into each case by holding back the peers' store writes, which precede their signing: one lets this node sign only once the counterparty's commitment_signed has arrived and keeps the counterparty's tx_signatures from ever arriving, and expects the record kept; the other keeps the counterparty from signing at all and expects the record dropped. When squashing, the base message's "which carries the channel's last funding" should read "which carries the channel's last funding, joined by the rounds its monitor still watches". Developed with assistance from Claude Code. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A splice round this node signed is kept at `ChannelClosed` when the channel's monitor watches it: the counterparty committed to it, so our signatures may have left the node, and the counterparty may broadcast the round and see it confirm. A close the wallet sees as a conflict -- a cooperative close spending an input the round shares -- fails the payment once it confirms beyond the reorg depth, but nothing resolved such a record when a commitment transaction, which pays no wallet script, won instead. Once the close matures -- after the reorg delay for a counterparty's commitment transaction, and once the to_self_delay on our balance has passed for one of our own -- the monitor stops watching the rounds it kept and queues a `DiscardFunding` event for each, and the handler only reclaimed the contribution's addresses: the funding payment stayed `Pending` forever. Likewise for a round of ours that a sibling round this node did not contribute to replaced on an open channel: LDK discards our round as the sibling locks, and the payment stayed `Pending` for a transaction that can no longer confirm. Resolve the funding payment the event names. A round nothing ever broadcast is dropped first, as at `ChannelClosed`, and with it a record no broadcast round of ours remains under. The payment is then left alone if a round of ours that LDK still holds remains in its record -- the round that locked, or one still pending -- or one LDK promoted to the funding before, and failed otherwise: no round of ours can confirm anymore, whether the channel closed on a commitment transaction or a round we did not contribute to locked. The rounds LDK holds are the channel's pending rounds and funding while the manager lists the channel, and once it does not, the funding its monitor settled on plus whatever the monitor still watches. The monitor is left out for a listed channel: its updates land after the manager's, deferred to the background processor's flush, so it may still watch a round the manager let go, and it learns a round only after the manager lists it. The event names a round by its transaction only when this node did not contribute to it; otherwise it describes what LDK returns of the contribution: the inputs and output scripts the round that replaced it does not reuse. Record each candidate's contributed inputs and output scripts so the event can be matched to the round, exactly or as the one recorded contribution with more parts. A round recorded before this carries no parts, and an event describing its contribution changes nothing while the channel is listed, as before. So does an event describing the channel's current funding: LDK also returns a contribution it refused before building a round from it, whole when the channel had no pending splice to check it against -- a fee bump adjusted from a round that locked as the bump was built, queued until the channel goes quiescent for it and returned once the node restarts, the channel force-closes or begins a cooperative close while no stfu is outstanding on it, the user cancels it, or the negotiation begun from it is refused, fails or is cut off by a disconnect -- and such a bump describes the locked round, while no round LDK discards can be the funding. A zero-conf splice is promoted to the funding as `splice_locked` is exchanged, before its transaction confirms, and a later splice moves the funding on again: at the close neither the manager nor the monitor holds the earlier round, although it can still confirm, the later round descending from it. So the funding payment records each promotion LDK reports through `ChannelReady`, and a round promoted once counts as one that can confirm wherever the rounds LDK holds decide: when LDK discards a sibling round, and when the channel closes. The monitor's events can reach the handler ahead of the channel's `ChannelClosed` when one sync delivers the close and its maturity: the channel manager polls the monitor's report of the close at the start of each event pass and on peer traffic, and the monitor's own events are handled right after the manager's. Each event then finds the channel still listed with every round held and leaves the payment. So `ChannelClosed` now fails every payment of the channel left with no round of ours the monitor watches and none promoted before, and a `DiscardFunding` event for a channel the manager no longer lists resolves each record by the rounds the monitor holds alone, without matching the event to a round: the close has settled what remains, and the held rounds decide for a record written without its parts or for records signed under different first-candidate ids that share one contribution, which a match cannot tell apart. Developed with assistance from Claude Code. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Funding records were keyed by a PaymentId derived from a funding txid: the broadcast txid in the generic classification path, the first negotiated candidate's txid in the interactive path. A txid is no identity for a replaceable transaction — the record deliberately outlives RBF rounds of its funding, so its key carried the txid of whichever round happened to come first, and code could be tempted to re-derive the id from a txid instead of resolving it. Generate the id from the OS entropy source when the record is created, and resolve existing records through their transaction history (find_payment_by_txid) everywhere. RBF stability now comes from resolution instead of derivation. Resolution must share one lock acquisition with the record writes: resolved outside it, the id could go stale against a record wallet sync creates for the same transaction, producing a divergent record — so classification acquires the cross-store lock itself and the write helper now takes the guard. The funding-record surface (classification, candidates, stable ids) debuts in the upcoming release — v0.7.0 shipped splice_in with no record machinery — so changing the scheme now costs nothing, while one release later it would break payment(&PaymentId(funding_txid)) lookups for new records. Generated with assistance from Claude Code. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A user-initiated splice dropped before LDK persists it leaves no trace in LDK. Recovering whatever the splice reserved and describing later events about it in terms of the original request both require persisting the splice intent before handing it to LDK, which happens before negotiation and therefore before any funding transaction exists. The pending-payment record was built around an on-chain PaymentDetails carrying a txid, which cannot represent a splice that has not been broadcast yet. Reshape PendingPaymentDetails into an enum: a PendingSplice variant that holds only the generated PaymentId and the splice intent, and a Tracked variant that is the previous record plus an optional intent retained until the splice locks. Add the SpliceIntent and SpliceKind types that record what was handed to LDK and the API call that produced it. Wallet writes to the pending store go through DataStore::mutate, replacing racy read-then-write pairs. They share one helper whose closure re-reads the payment's status inside the critical section — only Pending payments belong in the pending store, and a status read taken outside it can go stale against graduation — and promotes a bare PendingSplice to a Tracked record once a payment exists under its id: a plain payment-tracking merge would silently no-op against the variant, leaving the splice invisible to txid lookups. This is groundwork; nothing constructs a PendingSplice yet. A later commit adds the classification that reads the variant; the entry points that persist splice intents land with the splice tracking built on this. Generated with assistance from Claude Code. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A user-initiated splice will be keyed by a PaymentId generated at splice time rather than derived from a candidate's txid, so its splice intent, funding payment, and candidate history all share one record. Teach the classifier to find a pre-broadcast splice intent by its channel and reuse that id for a splice no record tracks yet, promoting the intent record to a tracked funding payment while preserving the intent until the splice locks. A round already on record keeps its record, whatever id it is under: the id of the first round of the history any record tracks is adopted before the channel's intent is consulted, and a fresh id is generated only when neither yields one. The intent identifies the channel, not a round: after a zero-conf lock, the channel may carry the intent of a newer splice while the locked round's classification is still queued, and consulting the intent first would file that round under the newer splice as a second record, with wallet sync then graduating whichever of the two it finds first. Every splice round this node contributes to that the wallet records is recorded when it is signed, before our signatures are released, so a round of ours is always on record by the time it is classified, and the intent only ever decides the id of a splice's first signed round. Splices we did not originate (counterparty-initiated or V2 dual-funded opens) have no intent, and a round wallet sync recorded first converges on the record sync created. An intent submitted for a channel whose history is already on record under another id is never promoted and stays bare until the splice locks or fails. A splice under a generated id is no longer found by the txid-derived lookup, so it leans on find_payment_by_txid's candidate probe to map its txids back to the record. The generic funding classification already resolves an existing record the same way before generating a fresh id: LDK re-broadcasts a promoted-but-unconfirmed 0conf funding transaction through that path, and a test added here covers the rebroadcast merging into the record classification already created rather than creating a duplicate. Promotion of a pre-broadcast intent in persist_funding_payment_locked is gated on the payment still being Pending, read inside the pending store's critical section like the rest of the write's decision: a payment that confirmed through ANTI_REORG_DELAY before classification must not re-enter the pending store, which graduation and rebroadcast assume holds only Pending payments. No splice intents are created yet; the splice entry points that persist them land in a follow-up — on this branch the intent probe stays dormant. Generated with assistance from Claude Code. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Wallet sync can observe a funding round before it is recorded as a candidate: the counterparty broadcasts an interactively funded transaction on its own, so a classification failure being retried here — or a sync poll racing the broadcast queue — leaves the round unrecorded while its events arrive. The funding-status gate rightly reports such a round foreign, and sync re-keys the event to the round's txid-derived id, creating an untyped duplicate record whose pending entry from then on shadows the funding record in txid resolution: even after the round's classification lands, every later event routes to the duplicate, the confirmation strands there, and the funding record never confirms or graduates. Fold the duplicate back in when its round becomes a recorded candidate: adopt its confirmation onto the funding record — through the same status-update path wallet sync uses, so the confirmed candidate's figures land — and remove the duplicate along with its pending entry. A duplicate for a round that never confirmed is dropped without adopting anything; the actively-broadcast candidate stays the record's current txid. The merge runs whenever a round is recorded — at its broadcast-time classification, or when this node signs a later round and records the channel's history with it — under the writer's cross-store lock acquisition, so sync cannot interleave, and is idempotent, so the broadcast queue's classification retry can re-run it after a partial failure. At signing time the merge is a courtesy and a failure is only logged: the signed round can have no duplicate yet, as our signatures have not left the node, its own classification re-runs the merge with the retry behind it, and failing the signing would replay it against a record whose two-store write already completed, which the write's rollback does not cover. The pending entry is removed before the payment record: a retry rediscovers the duplicate through the record, so a failure between the two removals can still be cleaned up, instead of orphaning a pending entry that would shadow txid resolution all over again. Generated with assistance from Claude Code. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
LDK only persists a splice once its negotiation reaches AwaitingSignatures, so a splice in flight when the node stops can leave no trace in LDK. Persist each user-initiated splice as an intent record before its contribution is handed to LDK, so such a splice can be recognized at the next startup — releasing whatever the wallet still holds for it, which a later commit adds — and so events about the splice can be described in terms of the original request. Each splice gets a record of its own, so that its failure is described from its own intent and a restart recognizes it whatever became of the channel's other splices: a splice queued behind a pending one negotiates as a splice of its own once the pending one locks, and its rounds must not be filed under the pending splice's payment. Only a fee bump joins an existing record, that of the round it replaces. A splice is refused while the channel carries an intent anchored at another funding — one the lock that superseded it failed to settle or to re-anchor — rather than recorded beside it. A submission reads the channel's funding under the lock that serializes splice submissions and anchors its intent there, not at the funding the caller read before building the contribution: a splice locking in between moves the funding, and an intent anchored at the old one would never be settled by the lock that superseded it. A funding that moved refuses a fee bump, whose round has locked, and a splice-in, whose inputs the locked round may have spent; a splice-out carries no wallet inputs and proceeds. A splice submitted after the previous one locked with zero confirmations settles that splice's intent first, as the lock's event would have: LDK promotes the funding as soon as splice_locked is exchanged but only queues the event. The lock and close event handlers settle intents under the same lock, so a lock handled mid-submission cannot settle the new intent before its contribution reaches LDK. The record is undone when LDK rejects the hand-off synchronously and settled once the splice locks, its failure is surfaced, or its channel closes. A failure event settles the intent only after the event is durably queued — a crash in between leaves the intent for the replayed event to settle, erring toward a duplicate report over a lost one — and only when the event's contribution identifies the recorded splice: a mismatch means the failure concerns an older, superseded attempt with no record of its own. Taking back the funding record of a signed round the failure abandoned leaves its intent behind as a bare intent, so the report can still describe the splice. A splice queued behind another pending splice survives the pending splice's lock, so its intent is re-anchored to the new funding rather than settled. Wallet state staged on a splice's behalf is flushed only after the intent record persists, so nothing the wallet reserves for a splice can outlive the record through which a later startup would release it. A splice that fails before the hand-off immediately releases what the wallet holds for it and no other round uses — a fee bump built by adjusting the fee of the round it replaces shares that round's inputs and change address, which stay reserved while the round can confirm; one LDK rejects has it returned through the DiscardFunding event instead. A lock settles an intent without releasing anything: what the locked round did not spend, LDK returns through the DiscardFunding events it queues at the promotion. Once a splice funding payment is classified, the intent is carried on the payment's record until the splice locks; a payment that already graduated instead removes the leftover intent record. The funding payment recorded when this node signs a splice round is filed under the record of the intent carrying the round's contribution, written while holding the lock that serializes splice submissions, so neither a fee bump replacing the intent nor a failure settling it can interleave with the write. A signing write cut short after the payment store leaves that payment under a bare intent; it records a round whose signatures never left the node, so it is dropped — when the replayed signing finds the round gone, or with the intent once the splice settles — rather than promoted into a record nothing could ever drive. Developed with assistance from Claude Code. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The signing handler previously logged and dropped both failure paths (with TODOs to abort once LDK supported it), leaving the negotiation dangling until a peer disconnect abandons it. Cancel the contributed funding instead. LDK then emits DiscardFunding, releasing whatever the wallet holds for the contribution, and SpliceNegotiationFailed, which surfaces the failure and settles the persisted intent. Cancel errors are only logged: every error case means the splice is already beyond canceling. When LDK refuses the already-signed transaction, the failure report that cancelling produces also takes back the payment recorded at signing time: the round is gone from the channel's history and nothing can ever broadcast it, so left in place the record would wait forever on a payment nothing can confirm. Developed with assistance from Claude Code. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
An application handling SpliceNegotiationFailed had nothing to act on: the event did not say why the splice failed, nor what the failed call had attempted. Both matter for deciding what to do next — a fee bump lost to a disconnect can simply be re-issued, while the splice it meant to bump may still confirm at the prior feerate. Attach a reason, mapped from LDK's NegotiationFailureReason onto an ldk-node-owned enum so the event's serialization and bindings do not change with LDK's, and the parameters of the originating API call, taken from the persisted splice intent when the failure identifies it. Both fields are optional and serialized as odd TLVs: events written by LDK Node v0.7 read back as None, and v0.7 readers ignore the new fields. Developed with assistance from Claude Code. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
LDK only persists a splice once its negotiation reaches AwaitingSignatures, so a splice in flight when the node stops can leave no trace in LDK's channel state, and no event of LDK's ever returns what the wallet reserved for it — today the addresses its outputs pay; once lightningdevkit#1037 locks a contribution's inputs in the wallet, those too, forever. At startup, reconcile each persisted splice intent against live channel state: release the reservations of a splice LDK no longer holds and drop its record, re-anchor a queued splice whose predecessor locked while the node was down, and keep — minus any inputs no surviving round still claims — those LDK resumes on its own. A splice whose channel closed meanwhile is released only if no round of it reached signing: a signed round is one the channel's monitor watches until the close matures, and what it reserved is spent by it or returned through DiscardFunding then. Reconciliation holds the lock that serializes splice submissions, as the event handlers settling intents do. Recovery fabricates no failure event for a splice lost this way: the initiating call already returned, and the channel simply no longer shows a pending splice. LDK itself reports the loss of a contribution it was still queueing or negotiating when it was last persisted — it fails the contribution as it is written and replays the failure at startup. The replay runs after reconciliation, so that report carries the splice's parameters only where reconciliation kept the intent: for a splice queued behind a pending one of ours, or a fee bump of one, but not for a channel's only splice, whose intent reconciliation settled. Reconciliation runs before background syncing and broadcasting start, so nothing can act on the stale reservations first. Events LDK replays from its last persisted state (e.g. a DiscardFunding for a splice that died before the node stopped) are likewise consumed before the node is running, so they cannot act on state a new user operation set up since. Developed with assistance from Claude Code. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A disconnect during the interactive negotiation fails the splice with PeerDisconnected. The first test asserts that exactly one SpliceNegotiationFailed reaches the user — carrying the reason and the originating request's parameters — and that a new splice initiated afterwards completes with a single funding payment. The window only exists mid-negotiation: a contribution still queued at disconnect is resumed by LDK itself on reconnect, and one awaiting signatures survives re-establishment. The test therefore synchronizes on the counterparty's splice_ack — logged by LDK's peer handler — and stretches the negotiation by funding the splice from many small UTXOs, each of which adds an interactive-tx round trip. A splice dropped by a restart is recovered silently: startup reconciliation releases what the wallet reserved and drops the record without fabricating a failure event. What does reach the user is the failure LDK persisted at shutdown and replays at startup — once, with parameters only when it still matches a kept record. The restart tests cover both cases: a dropped splice-out surfaces without parameters and a further restart stays silent, while a dropped fee bump — whose record reconciliation keeps, since LDK still holds the negotiated splice — surfaces with the bump's parameters. In both, the application re-initiates and the splice completes. A splice confirmed while its node was offline keeps exactly one payment record under its splice-time id regardless of whether wallet sync or classification sees the confirmation first. Three more cases: a second splice submitted right after a zero-conf lock gets a record of its own rather than being folded into the record of the splice that just locked; a queued splice the node stopped on, which LDK fails as it shuts down, is reported at startup with its parameters — its record, an intent that never became a payment, outlives the pending splice's graduation, and reconciliation keeps it while LDK still holds that splice; and a funding record left half-written by a stop between the signing write's two stores is dropped at the next startup instead of lingering as a payment nothing indexes. Developed with assistance from Claude Code. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
jkczyz
force-pushed
the
2026-08-splice-tracking
branch
from
September 8, 2026 16:33
ca4e5fd to
d961644
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Persist every application-initiated splice until LDK is guaranteed to remember it, so that a startup pass can release wallet inputs held for splices lost to a crash, and failure events can say which operation failed and why.
LDK persists a splice only once negotiation reaches
AwaitingSignatures; rounds short of that are failed on reload through theSpliceNegotiationFailed/DiscardFundingevents aChannelManagerwrite records alongside itself. But a splice initiated after the last manager write leaves no trace at all — no event ever comes. Once #1037 lands and splice contributions lock wallet inputs, a splice lost in that window would leave its inputs reserved forever, with nothing left running that knows to release them.What changes
Intent record. The splice is written into its pending payment record (from #1079) before the contribution is handed to LDK, and settled once the splice locks, its failure surfaces, or its channel closes. A fee bump replaces the intent of the splice it bumps — at most one record per channel.
Funding recorded at signing. Each negotiated round is recorded into the payment while handling
Event::FundingTransactionReadyForSigning, beforefunding_transaction_signedreleases ourtx_signatures. The counterparty can't broadcast without them, so the record precedes anything wallet sync can see, closing #1079's duplicate-record window for every round this node signs; the merge there remains the backstop for rounds we never sign. A failed write replays the event, which LDK re-offers and regenerates across restarts while unsigned. A round lost to a crash after signing keeps its candidate on record, so a counterparty broadcast after restart resolves to the existing payment.Abort on signing failure. If signing fails or LDK refuses the signed transaction, the splice is aborted through
cancel_funding_contributed, retracting the signing-time record in the latter case. LDK'sDiscardFundingandSpliceNegotiationFailedthen release the contribution and settle the intent. Previously the handler logged and left the negotiation dangling.Failure events.
Event::SpliceNegotiationFailedgains the failure'sreason(mirroring LDK's negotiation-failure reasons) and the initiating request'sparameters(In/Out/FeeBump). Both areNonefor splices this node didn't initiate; they're new odd TLVs, so events written by v0.7.0 still read and v0.7.0 readers skip them.Startup reconciliation. Runs before syncing and event processing, checking each record against the reloaded channel:
AwaitingSignatures, resumed on reconnect, orNegotiatedwith our contribution — keeps its record. Only inputs no surviving candidate spends are released; a fee bump lost with the restart may have reserved extras.Recovery is silent — the initiating call already returned and the channel shows no pending splice, so no failure event is fabricated. The failures LDK replays on reload are consumed before the node is marked running.
Notes for reviewers
DiscardFundinghandler that releases a discarded contribution's inputs, and a companion commit proposed there makes coin selection stage its locks so they reach disk only with the intent record.funding_contributedalso queues aSpliceNegotiationFailed, so a rejected call both returns an error and surfaces a failure event. Benign — the event no longer matches a recorded intent by then.Last in the PR stack replacing #930 for this release, per the discussion there; stacked on #1057 → #1079 (funding payment model).
Developed with assistance from Claude Code.