Skip to content

CL: SIGTERM during startup while waiting for the EL forces a full redb repair on next start #466

Description

@ZanCorDX

App::start() opens the redb store before connecting to the EL. During the EL
retry window (up to 30 s) no SIGTERM handler is installed yet, so a
systemctl stop there kills the process with the store open and without the
quick-repair commit redb writes on Drop. The next start runs a full repair:
1h04 on a 103 GB testnet store. Seen on arc-node v0.8.0, testnet, follow mode.

Reproduce (EL stopped):

  1. systemctl start arc-malachite, wait for the Failed to connect to Ethereum node retries.
  2. systemctl stop arc-malachite during the retries.
  3. Start the EL, then systemctl start arc-malachite: logs
    Database repair in progress: 0.00% instead of Database opened.

An orderly startup failure (execution engine unreachable ... Exiting...) does
not trigger it, since the store is dropped cleanly; only the unhandled signal
does. Opening the store after connect_and_resolve_chain_identity (first use is
State::builder) fixes it; verified on testnet. PR to follow.

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions