Today we expose 12 metrics: payment counts (total/successful/pending/failed), channel counts (total/public/private), four balance gauges and a peer count.
Health and liveness
| Metric |
Source |
Why |
| Sync freshness (lightning wallet, on-chain wallet, fee-rate cache, RGS, scores) |
NodeStatus, already in GetNodeInfo |
Alert on a stalled sync; today this needs a gRPC client rather than a scrape |
| Current block height |
NodeStatus.current_best_block |
Catches a stuck chain backend |
| Build info (version, commit) as a labelled gauge |
added in #229 |
Tells dashboards which build is running |
| Connected peers, separate from known peers |
Peer.is_connected |
update_peer_count currently stores list_peers().len(), which includes disconnected peers |
Channels and funds
| Metric |
Source |
Why |
| Usable / ready / unusable channels |
ChannelDetails.is_usable, is_channel_ready |
"Channels I have" vs "channels I can route through"; also covers the "offline channels" half of the LSP risk item in #121 |
| Total inbound / outbound capacity |
sum over ChannelDetails |
LSP capacity planning: can we still serve JIT channels? |
| Funds pending from channel closures, by state |
BalanceDetails.pending_balances_from_channel_closures |
How much is stuck in closes; current balance gauges hide this |
| Anchor reserve headroom (reserve vs spendable) |
existing balance fields |
Warns before the node starts rejecting inbound channels |
Payments and forwarding
| Metric |
Source |
Why |
| Payment failures by reason |
reason is on the PaymentFailed event (#251) |
Distinguishes routing failures from recipient rejections; bounded cardinality |
| Forwarding totals (forwards, fees earned, skimmed fees) |
ldk-node analytics via #283 |
LSP revenue; aggregate only, per-channel labels would be unbounded |
Event stream
| Metric |
Source |
Why |
| Dropped events |
service.rs already catches RecvError::Lagged(n) and discards the count |
A slow subscriber loses events silently today; relates to #245 |
| Active subscribers |
broadcast::Sender::receiver_count() |
|
API
No instrumentation exists in service.rs today: no request counts, error counts or latency for any RPC, and no counter for auth failures. I'd suggest aggregates first, with per-method labels as a separate decision because of cardinality.
Two related observations
payment_status_counts paginates through every payment on each poll (default 60s) to recount
statuses. Channel counts are already event-driven; doing the same for payments would remove
the scan and make additional metrics cheaper.
- There's no health/readiness endpoint, only systemd notifications. The sync values above plus
is_running would make one straightforward, if that's wanted.
If this looks reasonable I'd start with the health and liveness group, since those are the ones that make alerting possible. Happy to split the rest into separate issues.
Today we expose 12 metrics: payment counts (total/successful/pending/failed), channel counts (total/public/private), four balance gauges and a peer count.
Health and liveness
NodeStatus, already inGetNodeInfoNodeStatus.current_best_blockPeer.is_connectedupdate_peer_countcurrently storeslist_peers().len(), which includes disconnected peersChannels and funds
ChannelDetails.is_usable,is_channel_readyChannelDetailsBalanceDetails.pending_balances_from_channel_closuresPayments and forwarding
PaymentFailedevent (#251)Event stream
service.rsalready catchesRecvError::Lagged(n)and discards the countbroadcast::Sender::receiver_count()API
No instrumentation exists in
service.rstoday: no request counts, error counts or latency for any RPC, and no counter for auth failures. I'd suggest aggregates first, with per-method labels as a separate decision because of cardinality.Two related observations
payment_status_countspaginates through every payment on each poll (default 60s) to recountstatuses. Channel counts are already event-driven; doing the same for payments would remove
the scan and make additional metrics cheaper.
is_runningwould make one straightforward, if that's wanted.If this looks reasonable I'd start with the health and liveness group, since those are the ones that make alerting possible. Happy to split the rest into separate issues.