Skip to content

Candidate metrics to expand /metrics` #302

Description

@f3r10

Today we expose 12 metrics: payment counts (total/successful/pending/failed), channel counts (total/public/private), four balance gauges and a peer count.

Health and liveness

Metric Source Why
Sync freshness (lightning wallet, on-chain wallet, fee-rate cache, RGS, scores) NodeStatus, already in GetNodeInfo Alert on a stalled sync; today this needs a gRPC client rather than a scrape
Current block height NodeStatus.current_best_block Catches a stuck chain backend
Build info (version, commit) as a labelled gauge added in #229 Tells dashboards which build is running
Connected peers, separate from known peers Peer.is_connected update_peer_count currently stores list_peers().len(), which includes disconnected peers

Channels and funds

Metric Source Why
Usable / ready / unusable channels ChannelDetails.is_usable, is_channel_ready "Channels I have" vs "channels I can route through"; also covers the "offline channels" half of the LSP risk item in #121
Total inbound / outbound capacity sum over ChannelDetails LSP capacity planning: can we still serve JIT channels?
Funds pending from channel closures, by state BalanceDetails.pending_balances_from_channel_closures How much is stuck in closes; current balance gauges hide this
Anchor reserve headroom (reserve vs spendable) existing balance fields Warns before the node starts rejecting inbound channels

Payments and forwarding

Metric Source Why
Payment failures by reason reason is on the PaymentFailed event (#251) Distinguishes routing failures from recipient rejections; bounded cardinality
Forwarding totals (forwards, fees earned, skimmed fees) ldk-node analytics via #283 LSP revenue; aggregate only, per-channel labels would be unbounded

Event stream

Metric Source Why
Dropped events service.rs already catches RecvError::Lagged(n) and discards the count A slow subscriber loses events silently today; relates to #245
Active subscribers broadcast::Sender::receiver_count()

API

No instrumentation exists in service.rs today: no request counts, error counts or latency for any RPC, and no counter for auth failures. I'd suggest aggregates first, with per-method labels as a separate decision because of cardinality.

Two related observations

  • payment_status_counts paginates through every payment on each poll (default 60s) to recount
    statuses. Channel counts are already event-driven; doing the same for payments would remove
    the scan and make additional metrics cheaper.
  • There's no health/readiness endpoint, only systemd notifications. The sync values above plus
    is_running would make one straightforward, if that's wanted.

If this looks reasonable I'd start with the health and liveness group, since those are the ones that make alerting possible. Happy to split the rest into separate issues.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions