Skip to content

Fix scan limits mempool hardening - #264

Merged
philippem merged 28 commits into
Blockstream:new-indexfrom
agoodminute:fix/scan-limits-mempool-hardening
Oct 5, 2026
Merged

philippem merged 28 commits into
Blockstream:new-indexfrom
agoodminute:fix/scan-limits-mempool-hardening

Conversation

@agoodminute

Copy link
Copy Markdown
Collaborator

Summary

This branch hardens electrs against heavy addresses, slow or non-reading Electrum clients, and mempool churn. Each fix bounds a cost that used to grow without limit:

  • Scripthash scans. Expensive history and UTXO lookups are now bounded. Their progress is checkpointed, so repeated requests keep moving forward.
  • REST and mempool. REST handlers read the mempool from one consistent snapshot, so transactions that are evicted or confirmed mid-request no longer cause panics or inconsistent responses.
  • Electrum responses. Response sizes and write times are capped, so one client can no longer pin memory or a writer thread.
  • Mempool sync and daemon reconnect. These recover from transient failures instead of killing the process.
  • Checkpoint Merkle proofs. Proofs are served from a memory-bounded tree cache instead of being rebuilt from genesis on every request.

Changes

Scan limits (history / UTXO / stats)

  • New --history-scan-limit (default 100000, must be ≥ 1). It limits how many history rows one utxo/stats lookup reads. Once the limit is hit, the scan finishes the current height and then stops with TooBigHistory ("too popular"). It never stops in the middle of a height, which used to make it stall forever.
  • New --utxos-checkpoint-limit (default 10 × --utxos-limit, must be ≥ --utxos-limit). It limits the size of a partially scanned UTXO set and its saved checkpoint, as well as how many unconfirmed history entries are scanned per address.
  • --utxos-limit is now checked against the final UTXO set, not against intermediate scan state.
  • A checkpoint is saved when the scan limit is reached, so the next request resumes from there. This fixes stats_delta double-counting on resume.
  • Fixed: the mempool scan bound, history pagination truncation (truncated pages are now rejected), and cursor scans continuing past the cursor height.
  • Precache retries on TooBigHistory.
  • REST returns TooBigHistory as HTTP 400.

REST mempool consistency

  • Each handler takes a single mempool read snapshot for transaction selection and prevout resolution. It captures that snapshot before chain reads and releases it before building the response, so the lock is not held during serialization.

  • Confirmed history entries are deduplicated, and confirmation status is preserved when a transaction confirms during a request.

  • Error classification:

    • 400 for an invalid block page index or a too-large history;
    • 503 for storage failures and confirmed-prevout lookup failures;
    • 404 for an unconfirmed mempool prevout that has disappeared.

    Lookup failures are logged at warn.

  • Query::mempool() recovers from a poisoned lock on both the read and write paths.

Electrum response limits

  • --electrum-rpc-max-response-num-bytes limits the size of one reply line. Defaults: 8 MiB on Bitcoin, 32 MiB on Liquid. A reply that overflows gets a correlated error (code 1), and the remaining batch elements are not executed.
  • --electrum-rpc-global-response-budget-bytes limits reply buffer memory across all connections combined. Defaults: 64 MiB on Bitcoin, 256 MiB on Liquid.
    • Requests over the budget are rejected with code 2, and the connection stays usable.
    • Replies under 16 KiB per connection don't count toward the budget.
  • --electrum-rpc-write-timeout (default 30s) is a deadline for transmitting a complete reply. It is not extended by partial progress. When it expires, the connection is closed and its buffers are released.
  • Validation:
    • the budget must be ≥ the per-line cap;
    • a finite budget requires a non-zero write timeout or --electrum-rpc-conn-max-age.
  • New metrics: electrum_per_line_response_overflows, electrum_global_response_budget_rejections, electrum_client_write_timeouts_total.
  • Fixed a header range overflow.

Checkpoint Merkle proofs

  • The previous concurrency cap is replaced by a tree cache keyed by cp_height. A cached height is served without a rebuild.
  • New --electrum-checkpoint-merkle-cache-mb (default 256). Cached trees and in-flight rebuilds share this one budget. Cached trees use at most half of it, and rebuild memory is reserved up front.
  • --electrum-checkpoint-proof-concurrency-limit now applies only to rebuilds. Its default is half the CPU cores (minimum 1) instead of 2.
  • Fixed races between cache eviction and reorgs.

Mempool sync and daemon

  • Missing transactions are fetched through bounded RPCs outside the mempool write lock. Presence is rechecked before insertion.
  • Transient sync errors no longer kill the process. The main sync loop no longer pauses when adding transactions fails.
  • Unresolvable transactions are isolated from the rest of the batch. Sync stops after repeated failures. Fixed an off-by-one in the retry cap.
  • The boolean sync result is replaced with an enum. A partial sync is reported as FailedToIndex.
  • New metric: mempool_sync_failed_to_index.
  • retry_reconnect backs off between attempts.
  • After a broadcast, the transaction is indexed using the witness the daemon actually accepted.

DB compatibility

  • index_unspendables and address_search are now included in the compatibility check. They are stored under a new VF key, and the V record keeps its old format so older binaries can still read it.
  • If the stored flags don't match the current config, startup fails with a reindex message. Existing DBs without VF record the current flags and log a warning.
  • README: --address-search is documented as best-effort with respect to reorgs.

Liquid

  • Fixed MissingTxo txid parsing.
  • Larger Electrum response defaults, sized for a 2016-header window during a dynafed vote.

Behavior changes for operators

  • New limits are enabled by default:

    • history scan limit: 100k rows;
    • Electrum per-line cap: 8/32 MiB;
    • global response budget: 64/256 MiB;
    • write timeout: 30s.

    To restore the old unbounded behavior, set the response caps and the timeout to 0. --history-scan-limit can only be raised.

  • The default for --electrum-checkpoint-proof-concurrency-limit changes from 2 to half the CPU cores.

  • REST error responses: some lookups that used to return 404 or panic now return 400 or 503.

  • A DB built with different --index-unspendables / --address-search values now refuses to start.

EddieHouston and others added 28 commits September 30, 2026 11:07
- Query::mempool() recovers from a poisoned lock via
  unwrap_or_else(|e| e.into_inner()) rather than unwrapping. Under
  panic=abort the process terminates before a poison marker is
  observable, but serving stale reads is strictly better than
  aborting the front-end for anyone who turns that off later.
- Remove with_mempool() and with_read_snapshot() along with the
  test that only covered the helper. The five REST call sites now
  bind the guard directly via let mempool = query.mempool(); with
  an explicit drop(mempool); at the end of the atomic section. The
  atomicity guarantee comes from scope, and the codebase keeps a
  single mempool-access idiom.
- Delete Query::lookup_txos. With prepare_txs taking a lookup
  closure instead of a HashMap, the wrapper had no non-test
  callers.
- HttpError::lookup emits a warn log before returning so operators
  can see what is driving REST 404 and 503 responses.
- Note the pre-existing race between chain.history() and the
  mempool snapshot in the address and asset history handlers.
- Adjust tests/mempool_eviction_race.rs to call
  query.mempool().lookup_txos(...) directly instead of the deleted
  Query::lookup_txos wrapper. Two separate mempool snapshots is
  exactly the pre-fix code path the test needs to exercise. Also
  add query, mempool, and daemon accessors on TestRunner in
  tests/common.rs.
  Keep one mempool read guard across transaction selection and prevout
  resolution in the eviction regression test. Assert that a writer cannot
  acquire the lock during this atomic section and can acquire it after the
  snapshot is released.

  Retain coverage that an evicted prevout returns MissingTxo instead of
  panicking.
Capture mempool transactions and prevouts before chain reads, then release the lock before response preparation. Deduplicate confirmed history
  entries and preserve transaction confirmation status.

  Exercise REST preparation in regression tests and fix Liquid test compatibility.

  Validation: 111 Bitcoin tests and 112 Liquid tests passed.
Check chain storage before capturing mempool transactions. Return 400 for invalid block page indices, 503 for storage and confirmed-prevout failures, and 404 for vanished unconfirmed mempool prevouts.
Fetch missing transactions through bounded proxied RPCs outside the mempool lock. Recheck presence before insertion, log unavailable txids, and cover lock behavior with a regression test.
- stop killing the process on transient mempool sync errors
- back off between retry_reconnect attempts
- replace the overloaded bool mempool sync result with an enum
- stop pausing the main sync loop on mempool add failures
- fix off-by-one in the mempool sync retry cap
- add mempool_sync_failed_to_index metric
- cover reconnect backoff and interruption in tests
…hot merge

The prepare_history path from fix/rest-mempool-snapshot-race wraps chain
lookups in HttpError::lookup (503); the scan-limit branch expects
TooBigHistory to reach clients as 400. Classify it explicitly.
…nse limits

Bound utxo checkpoints with --utxos-checkpoint-limit, retry TooBigHistory in precache,
stop cursor scans past the cursor height, and reject a zero history scan limit.
Keep the V record readable by older binaries and store index flags under VF.
Index broadcasts from the submitted hex, isolate unresolvable mempool txs and stop
after repeated sync failures. Reserve rebuild memory in the checkpoint Merkle cache,
require a write timeout or max age with a finite response budget, do not abort
completed writes at the deadline, and fix the header range overflow and the Liquid
MissingTxo txid parsing.
@philippem
philippem merged commit dd09701 into Blockstream:new-index Oct 5, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants