Skip to content

Latest commit

 

History

History
211 lines (159 loc) · 14 KB

File metadata and controls

211 lines (159 loc) · 14 KB

Hyperion Full-History Setup for XPR Network

Guide for standing up a Hyperion v4 full-history node on XPR mainnet, written from a production full-history build on XPR mainnet (Hyperion 4.0.8, built July 2026) that serves history at head with all /v2/health services OK as of September 2026. Every caveat here was hit in practice or confirmed by other XPR Network operators.

Read hyperion-operations-caveats.md before your first backfill. It covers the failure modes that cost days: the composable-template trap, Redis bloat, disk-full stalls, queue purges that silently lose data, and how to prove the index is actually complete.

Architecture

Hyperion is a pipeline, not a single service:

nodeos (SHIP) ──ws──> Hyperion indexer ──> RabbitMQ queues ──> Elasticsearch (+ MongoDB in v4)
                                                                      │
                             Hyperion API (7000) <────────────────────┘
                             Streaming API (1234, socket.io)

All components can run on one box for XPR's load. Redis is used for caching/IPC; MongoDB is mandatory in v4.

Version matrix (verified July 2026)

Component Required Notes
Hyperion v4.0.8+ (4.1.0 current, Sept 2026) v4 added MongoDB; v4 has exploit fixes over v3 — don't deploy v3
nodeos Leap 5.0.x (or Spring 1.2.2+) Match the network: check server_version_string from get_info on a trusted mainnet API before choosing (mainnet BPs report v5.0.0–v5.0.3 as of Sept 2026)
Elasticsearch 9.x
RabbitMQ 4.x (min 3.12) Ubuntu 24.04's apt ships 3.12 (EOL) — use upstream repo
MongoDB 8.x
Node.js 22+ NodeSource apt is simpler than fnm for root services
OS Ubuntu 24.04 LTS NOT the newest Ubuntu: MongoDB/ES repos lag new releases; Leap debs target 22.04/24.04. 22.04 EOL is too close to build on

Hardware sizing (XPR mainnet, ~392M blocks, July 2026)

  • Full history needs ~2TB today, growing slowly. Measured: state-history ~1.2TB + blocks ~220GB + ES indices on top. Plan 3.5–4TB raw for years of headroom.
  • Fast NVMe with good 4K random reads is non-negotiable for Elasticsearch — on slow disk both indexing and query latency collapse. State-history can live on slower disk; ES cannot.
  • 64GB RAM works; 128GB is comfortable. Never give ES more than ~31GB heap (compressed-oops threshold; bigger heaps hurt).
  • 8 modern cores suffice for XPR's current load; 16 makes the initial index much faster.
  • Reference build that runs it comfortably: one dedicated server with a current 16-core desktop-class CPU, 128GB ECC RAM, and 2×1.92TB datacenter NVMe, roughly $260–300/mo at mid-2026 dedicated-hosting prices. Entry-level models turn out poor value once per-drive upgrade fees (~$50/mo per 1.92TB NVMe) are added.

Disk layout caveat (dedicated hosts with two NVMe drives)

Most dedicated-host installers default to RAID1 across both drives — that halves capacity and full history won't fit in 1.92TB. Split it: keep OS + ES on drive 0, dedicate drive 1 to nodeos data. Breaking the mirror live (no reinstall/rescue needed):

mdadm /dev/md2 --fail /dev/nvme1n1p4          # may need: echo idle > /sys/block/md2/md/sync_action first
mdadm /dev/md2 --remove /dev/nvme1n1p4         # "Device or resource busy" → wait for the fail to settle, retry
mdadm --grow /dev/md2 --raid-devices=1 --force # clean [1/1] state, not degraded
mdadm --zero-superblock /dev/nvme1n1p4
mkfs.xfs -f /dev/nvme1n1p4
# fstab by UUID, mount /data, then: mdadm --detail --scan → /etc/mdadm/mdadm.conf, update-initramfs -u -k all

Leave /boot/EFI/swap partitions mirrored — costs nothing. Reboot-test before building anything on top. No redundancy is acceptable: a history node is rebuildable.

OS prep

# sysctl (/etc/sysctl.d/90-hyperion.conf)
vm.max_map_count=1048575    # ES requirement
vm.swappiness=1
fs.file-max=2097152

# limits (/etc/security/limits.d/90-hyperion.conf) — SHIP nodes die with
# "too many open files" under Hyperion load without this (seen on mainnet)
* soft nofile 262144
* hard nofile 262144

Firewall: 22 (rate-limited), 80/443 (nginx), 9876 (p2p if serving as seed). Everything else (ES 9200, RabbitMQ 5672/15672, Redis, MongoDB, SHIP 8080, nodeos 8888, Hyperion 7000/1234) binds localhost only.

Dependency install caveats

Elasticsearch 9: apt repo https://artifacts.elastic.co/packages/9.x/apt. Security autoconfig prints the elastic password during install AND enables TLS on the HTTP layer — plain curl http://127.0.0.1:9200 silently fails. For localhost-only, disable it (keep auth!):

# /etc/elasticsearch/elasticsearch.yml — in the autogenerated block:
xpack.security.http.ssl:
  enabled: false
network.host: 127.0.0.1
discovery.type: single-node   # and comment out cluster.initial_master_nodes

Heap via /etc/elasticsearch/jvm.options.d/heap.options: -Xms31g -Xmx31g.

RabbitMQ: the official apt mirrors moved from ppa1/ppa2.rabbitmq.com to deb1/deb2.rabbitmq.com (single team key 0A9AF2115F4687BD29803A206B73A36E6026DFCA via keys.openpgp.org). If DNS fails on ppa hosts, that's why. Upgrade trap: 4.x refuses to start on a 3.12/3.13 data dir (disabled_required_feature_flag: message_containers). On a fresh box just rm -rf /var/lib/rabbitmq/mnesia and recreate vhost/user:

rabbitmq-plugins enable rabbitmq_management
rabbitmqctl add_vhost hyperion
read -rs -p "hyperion rabbitmq password: " RMQ_PASS   # keeps it out of argv/ps and shell history
rabbitmqctl add_user hyperion "$RMQ_PASS"; unset RMQ_PASS
rabbitmqctl set_user_tags hyperion management          # not administrator — Hyperion needs no cluster admin
rabbitmqctl set_permissions -p hyperion hyperion ".*" ".*" ".*"

MongoDB 8: standard repo (repo.mongodb.org/apt/ubuntu noble/mongodb-org/8.0). No auth needed for localhost (Hyperion default), but don't expose it.

nodeos SHIP node

Install the Leap .deb matching the network (https://github.com/AntelopeIO/leap/releases/download/v5.0.3/leap_5.0.3_amd64.deb — installs fine on 24.04). Clone https://github.com/XPRNetwork/xpr.start for genesis.json (chain id 384da888112027f0321850a169f737c33e53b388aad48b5adace4bab97f437e0).

Key config.ini for a SHIP-only node (not public API):

http-server-address = 127.0.0.1:8888        # local only; Hyperion + ops
p2p-listen-endpoint = 0.0.0.0:9876          # optional public seed
chain-state-db-size-mb = 32768
wasm-runtime = eos-vm-jit
eos-vm-oc-enable = 1                        # big replay speedup
eos-vm-oc-compile-threads = 8
contracts-console = false                   # console output is NOT indexed by Hyperion anyway

plugin = eosio::chain_api_plugin
plugin = eosio::http_plugin
plugin = eosio::net_plugin
plugin = eosio::net_api_plugin
plugin = eosio::state_history_plugin
state-history-endpoint = 127.0.0.1:8080     # NEVER expose SHIP publicly
trace-history = true
chain-state-history = true

Point --data-dir at the dedicated data drive. Live peer list: https://danemarkbp.com/apis/get_json_mainnet.php?status=active&type=p2p (danemarkbp probes endpoints continuously — far better than any static list in docs).

Full history: the only two honest paths

A snapshot cannot produce full history — SHIP data must exist from block 1, and snapshots carry state only. Either:

  1. blocks.log + local replay (recommended): download a full block log, put it in the data dir, start with --genesis-json. nodeos re-executes every block locally and writes state-history from block 1. Identical trust model to p2p sync — the log is just input data; all state is derived locally.
    • Source (still live Sept 2026): https://snapshots.saltant.io/mainnet/blocks.tar.zst (~127GB compressed). Mirrors come and go — ask in the official Telegram (https://t.me/XPRNetwork) if dead.
  2. Pure p2p from genesis: same local re-execution, plus days of network fetch, and depends on peers serving 6-year-old blocks.

Replay runs for days. eos-vm-oc-enable = 1 matters here. The 0.5s block time means XPR has ~2× the blocks of most chains per year of history.

Hyperion install & config caveats

cd /opt && git clone https://github.com/eosrio/hyperion-history-api.git
cd hyperion-history-api && npm ci    # if it exits nonzero at "Updating permissions...", run npm ci again
  • Config lives in config/ (config/connections.json, config/chains/<chain>.config.json).
  • ./hyp-config connections init crashes without a TTY (ERR_USE_AFTER_CLOSE: readline was closed) — its MongoDB section prompts unconditionally, ignoring CLI flags. Workaround: copy references/connections.ref.json to config/connections.json, fill in elasticsearch/amqp/redis/mongodb blocks, set "chains": {}, then verify with ./hyp-config connections test (must show all four true).
  • ./hyp-config chains new proton --http http://127.0.0.1:8888 --ship ws://127.0.0.1:8080 requires nodeos serving — it fetches the chain id and does a live SHIP handshake. A live-syncing node works; a replaying node does not (no listeners until replay ends — see the ordering caveat below).
  • Generated chain config fails its own schema: experimental key is missing and validate --fix can't add it. Set "experimental": {} manually, then ./hyp-config chains validate proton passes.
  • Chain short name proton is recognized (pre-rebrand name) and pulls correct chain metadata.

Scaling for a 16-core box (defaults are all 1):

"scaling": { "readers": 4, "ds_queues": 4, "ds_threads": 8, "ds_pool_size": 8,
             "indexing_queues": 4, "ad_idx_queues": 2, "dyn_idx_queues": 2 }

Indexing flow: first run with "abi_scan_mode": true (fast ABI discovery pass), then flip to false for the full index. Do not run the full index as one open-ended pass from genesis — backfill in bounded 10M-block start_on/stop_on ranges (caveats §5.5).

Ordering caveat (verified, contradicts common advice): you cannot run the indexer while nodeos is replaying. During replay, nodeos opens neither the SHIP websocket (8080) nor the HTTP endpoint (8888) — plugins only start listening once the chain finishes loading. Hyperion has nothing to connect to, and hyp-config chains new fails for the same reason. Sequence must be:

  1. Start the replay, wait for it to complete (days for full history — controller.cpp:528 replay ] N of M in the log tracks it).
  2. Once nodeos serves get_info and 8080 is listening, run hyp-config chains new.
  3. Then ABI scan, then the full index.

Indexing concurrently with a live-syncing node (i.e. post-replay, catching up over p2p) is fine — that's the case the general Hyperion docs describe.

Enable streaming if you want the websocket API: "features": { "streaming": { "enable": true, "traces": true, "deltas": true } } — the v4 stream client is the dev branch of eosrio/hyperion-stream-client, published on npm as @eosrio/hyperion-stream-client@4.0.0-rc.3 (the latest tag).

Operations (hard-won BP wisdom)

  • Stopping the indexer: use trigger stop, not pm2 stop — trigger stop lets the queues drain; a hard stop leaves messages in RabbitMQ and can corrupt progress. (./stop <chain>-indexer in v4; it reads control_port from connections.json → chains.<chain>.control_port, default 7002.)
  • Watch disk on BOTH filesystems. Running out of space is the #1 Hyperion killer among operators (it has taken out a BP's pair of instances simultaneously). ES stops allocating shards (new partitions fail) at the high watermark and goes read-only at flood stage — on large disks both are capped by max_headroom, so check the real values — see caveats §3, plus §2 for the Redis temp-file failure mode.
  • Missing blocks: chain microforks cause occasional gaps; there's a repair script in Hyperion's scripts/ folder to find and fix them. EOSphere's blog also covers it.
  • Queue backups: RabbitMQ queues randomly backing up is a known Hyperion quirk (operators running dozens of instances see it) — watch queue depth in the mgmt UI (127.0.0.1:15672). Revive procedure in caveats §5.5.
  • Health check: /v2/health must show StateHistory/RabbitMq/NodeosRPC/Elasticsearch all OK — this is what danemarkbp's public checker polls. Get listed: https://danemarkbp.com/bp-stats.
  • Internal 429s come from Hyperion's own rate limiter (protects ES), not the proxy.

nginx public endpoint

limit_req_zone $binary_remote_addr zone=hyperion:20m rate=50r/s;

server {                                   # plain HTTP redirects to TLS
    listen 80;
    server_name hyperion-xpr-mainnet.example.com;
    return 301 https://$host$request_uri;
}

server {
    listen 443 ssl http2;
    server_name hyperion-xpr-mainnet.example.com;

    ssl_certificate /etc/letsencrypt/live/hyperion-xpr-mainnet.example.com/fullchain.pem;
    ssl_certificate_key /etc/letsencrypt/live/hyperion-xpr-mainnet.example.com/privkey.pem;

    location / {
        limit_req zone=hyperion burst=100 nodelay;
        proxy_pass http://127.0.0.1:7000;
        proxy_http_version 1.1;
        proxy_set_header Host $host;
        proxy_read_timeout 120s;          # deep history queries are slow
    }
    location /stream/ {                    # socket.io lives on the separate stream port
        proxy_pass http://127.0.0.1:1234/socket.io/;
        proxy_http_version 1.1;
        proxy_set_header Upgrade $http_upgrade;
        proxy_set_header Connection "upgrade";
        proxy_read_timeout 3600s;          # long-lived websockets
    }
}

If DNS is on Cloudflare: grey cloud (DNS only). The orange-cloud proxy times out long history queries and interferes with streaming websockets.

bp.json

Advertise the endpoint under your producer account's bp.json nodes array with "features": ["hyperion-v2"] (and "chain-api" if you also proxy /v1/chain). Hyperion forwards non-history calls to its configured chain_api upstream, so a Hyperion endpoint doubles as a chain API unless you point it elsewhere.