In four days fukuii’s Sepolia node went from crashing every ten hours to a steady 0.5 GB heap, its storage peers went from 10 of 11 penalised to none, and we started splitting the 6,800-line SNAP controller into small, testable phase modules so the next round of speedups can land safely.
All of it was measured on a deliberately modest machine: 4 cores, 16 GB of RAM, a consumer SATA SSD and Wi-Fi. We tune for hardware a home operator already owns, not for cloud instances.
Where we started
We run a fukuii node on Sepolia as the proving ground for a full SNAP sync, the same path an ETH mainnet node will take. Sepolia activated Amsterdam on Oct 6, so this node is also our first real check that the new fork works end to end.
A quick note on what SNAP sync is. Instead of replaying every block since genesis, a new node picks a recent block (the pivot) and downloads that block’s state directly from peers: every account, every contract’s bytecode, and every contract’s storage slots, each range checked against a Merkle proof. A healing pass then patches whatever changed while the download ran. Storage is the heavy phase because it is by far the largest part of the state: tens of millions of slots spread across contracts, each one written into a trie on local disk.
By Oct 8 the sync was making progress but kept tripping over itself:
- It ran out of memory every ~10 hours. The 4 GB heap filled at 0.3–0.4 GB an hour, then the JVM exited and the node restarted.
- It lost hours of work on each restart. Storage tasks already finished were downloaded again, and an earlier build threw away account progress entirely.
- Its storage peers were almost all penalised. At one point 10 of 11 were benched, leaving one or two to serve the whole storage phase.
- The SNAP controller was hard to change safely. One 6,800-line class held peer handling, pivot selection, downloads, healing, validation, resume and finalization, with 82 mutable fields shared between them.
We worked from measurements, not guesses: a heap dump read with a streaming parser, per-peer penalty counts from the logs, GC logs, and the controller’s own checkpoint lines for sync progress.
A deliberately small box
All of this runs on purpose on modest hardware: one machine with 4 cores, 16 GB of RAM, a consumer SATA SSD (a Crucial BX500), a 4 GB Java heap, and a Wi-Fi link. That is not a stopgap until we get a bigger server. We want fukuii to sync well on hardware an individual operator already owns, without assuming cloud instances, NVMe arrays or a 32 GB heap. A constrained box also makes problems loud. A memory leak that a 64 GB server would hide for weeks kills this node in ten hours, and a write pattern a fast NVMe drive would absorb shows up here as a stalled disk. Every bottleneck in this post was found because the hardware refused to hide it.
What we fixed
The heap dump settled the memory question quickly. The SNAP data was not what filled the heap. The space went to things the node should not have been holding at all during a sync:
- 2.6 GB of blob-transaction sidecars from the transaction pool, which a syncing node can’t validate or use yet
- 630 MB of idle discovery (discv4) server channels that were never closed
- 313 MB of storage-task payloads that were kept in full when a task was requeued
Each of these became its own small PR. We then added circuit breakers, so the next unknown leak gets throttled instead of crashing the node.
| PR | What it does |
|---|---|
| #1501 | SNAP intake caps and a heap watchdog that pauses intake when old-gen use stays high after GC |
| #1500 | A fixed memory budget for RocksDB block cache and memtables |
| #1522 | TxGossipGate: no transaction gossip while syncing; blob sidecars get a lifecycle and a 256 MiB budget |
| #1521 | discv4 idle-channel timeout, supervised consumers, and a cap of 2,048 server channels |
| #1520 / #1523 | A requeued storage task keeps only its key range, not the downloaded slots and proofs |
| #1524 | Caps on the request-tracker maps |
| #1525 | One requester per announced transaction, with alternates and a 5 s timeout (matches geth’s tx fetcher) |
| #1531 | The request tracker’s owner decides timeouts in mailbox order, so late replies no longer penalise good peers |
| #1499 | Fixed reconnect races in the peer crawler |
#1531 fixed the peer penalties. Storage replies that arrived just after their timer fired were counted as failures, even though the data was valid and already queued. Under load that happened all the time, and peers were benched for being slightly slow.
Results on the live node
The Sepolia node has run these fixes without an OOM since they landed. Measured before and after:
| Metric | Before | After |
|---|---|---|
| Old-gen heap after GC | 3.8 GB and climbing | ~0.5–0.7 GB, flat |
| Max heap | 6 GB, then OOM at 4 GB | 4 GB with room to spare |
| Storage peers penalised | 10 of 11 | 0 |
| Storage-slot download rate | ~0.4M entries/h (just before #1533) | ~1.7M entries/h, sustained |
| Progress lost on restart | hours | resumes from the last checkpoint |
The rate took two steps. The memory and peer fixes lifted it to a 1.49M/h peak, but it did not hold. With memory and peers healthy, the sync hit a new limit: one storage coordinator actor that verified every proof, built every slot trie and wrote every trie node to RocksDB itself. Once the backlog grew, throughput fell to about 0.4M entries/h, with peers idle and that one actor saturated. That was issue #1533.
In order, the storage rate went like this:
| When | What changed | Storage rate |
|---|---|---|
| Oct 8 | Starting point: OOM every ~10 hours, 10 of 11 storage peers penalised | ~0.69M entries/h |
| Oct 9 | Memory caps, txpool and discovery leak fixes, peer-penalty fix (#1531) | peak 1.49M/h |
| Oct 9–10 | Storage backlog builds up; one coordinator actor saturates | falls to ~0.4M/h |
| Oct 10, 11:09 | #1533 hotfix: batched trie writes, off-actor verification | ~1.7M/h sustained |
The #1533 fix (PR #1548) shipped to the node on Oct 10 as hotfix 0.9.24-sepolia.2. It writes each response’s trie nodes as one RocksDB batch instead of one write per node, and it moves proof verification to a small worker pool while keeping commits in arrival order on the actor. Over the first hour the node went from 41.71M to 43.70M storage entries, about 1.7M/h. That is roughly 4× the rate before the fix, and above the earlier peak.
The fix also logs where each response’s time goes ([STORAGE-PERF]), and the numbers changed our picture of the problem. Proof verification turned out to be cheap: 14 ms median, 24 s in total over the first 973 responses. The RocksDB write was 95% of the actor’s time. Batching those writes is where the gain came from, and the disk is now the limit.
That reading points straight at the next levers. If RocksDB writes are the cost, the options are to write less per entry (lighter write-ahead logging while SNAP runs, since done-markers already make it restartable), to stop other work competing for the same disk (throttling chain backfill during state sync), or to give the writes a faster place to land. Each is covered under What’s next, and the same [STORAGE-PERF] line will tell us which one paid off.
At the 12:19 checkpoint on Oct 10 the node held 43.7M storage entries.
The plan of attack: splitting the SNAP controller
Every fix above touched SNAPSyncController, and every one was harder than it needed to be. In that class, a change to pivot refresh can break healing, and a change to resume logic can break finalization. Spec 016 (issue #1401) breaks it into modules, one per SNAP phase or concern, so later work can target a single phase.
We started by mapping the class: which of its 82 fields each method reads and writes. That map drives the order of the split. The work then runs in four stages:
- Lock down current behaviour (S0). Pin tests, golden-byte tests for persisted state, and characterization tests that record what each phase does today, including its odd edge cases.
- Prerequisites (P1–P4). Group the child-actor references into
CoordinatorHandles. GivePhaseFlagsareset(kind)so phase resets are explicit. Split peer-event handling into arms, and dispatch messages through a per-phase table,phaseArms(currentPhase). - Module moves (M1–M11). Move each concern into a self-typed trait in two commits. The first is byte-identical and is checked by a script. The second, required narrowing commit cuts the module down to what it needs: a small
<Module>Stateinterface plus the few shared capabilities it actually uses. - Coordinators. Apply the same treatment to the account, storage, bytecode and healing coordinators, then finish with a final review pass.
[embedded content: SNAP controller split · today vs. target modules]
The highlighted modules are already merged to staging. Each module states the few fields and capabilities it uses, so the core that remains is just message dispatch and shared state.
The plan has 47 PRs over about 9–12 weeks. Done so far: all of S0 and P1–P4, plus M1–M6a. That covers the policy objects, state validation, the peer pool, finalization and the resume planner. The healing module (M6b) is in progress. Stagnation, pivot selection, lifecycle, pivot refresh and the download supervisor come next.
Safety rails
The split rewrites code that a live sync depends on, so each step has to be cheap to check:
- Moves are verified, not trusted. Every move commit carries
# moved:markers. A script (scripts/snap-split/verify.py, run in CI assnap-split-verify) checks that the moved lines match the originals byte for byte. A reviewer of a move commit only has to confirm the script passed. - Behaviour changes are kept apart. Narrowing commits change interfaces, not logic. Anything that changes behaviour goes in its own PR with its own tests.
- The tests come first. Golden bytes protect persisted state (checkpoints, done-markers). Characterization tests and dispatch tables record which phase handles which message, so a lost message arm fails a test instead of stalling a sync.
- Every PR gets a second reviewer. A different specialist from the author reviews each one before merge. Consensus code is reviewed by
beacon(ETH) orforge(ETC). - The live node is isolated from the split. Sepolia runs from a hotfix line (
hotfix/sepolia-0.9.24), cut from a tagged base taken before the moves. Urgent fixes such as #1531 are cherry-picked onto it. The module moves stay onstaginguntil they have had a full sync run.
Expected benefits
- Phases we can target. A change to healing, pivot refresh or finalization goes into one module with its own tests, instead of a 6,800-line class where any edit can affect every phase. Issues about one phase can name the module and the interface it may touch.
- Coupling we can measure. Each module states what it needs: its
<Module>Statefields and the shared capability traits it uses (SnapSharedState,SnapControllerEnv,CoordinatorHandles,PhaseFlags). Those counts are numbers we can track. If one grows, review sees it.SnapControllerEnvis already on our watch list. - Safer performance work. The next speedups mostly mean moving work off an actor or reordering it: #1533’s off-thread processing (now live), parallel healing and snapshot serving. Those changes are only safe when the actor’s state and message handling are small enough to reason about. The split is the groundwork for them.
- Readiness for mainnet. ETH mainnet state is roughly 10–20× Sepolia’s. Every bottleneck we see on Sepolia will be larger there, and a restart that costs an hour on Sepolia would cost a day on mainnet. The memory caps, checkpoints and peer fixes are prerequisites for a mainnet sync, and the split lets us keep tuning safely once that sync is running.
What’s next
- Easing the disk. With #1533 live, RocksDB writes on the node’s SSD are the bottleneck. The options include lighter write-ahead logging during SNAP (the done-markers already make it restartable), throttling chain backfill while state sync runs, and moving the datadir to faster storage.
- Spec 016 continues. Next come M6b (healing), then stagnation, pivot selection, lifecycle (reset, dormant mode, heap watchdog), pivot refresh and the download supervisor. After the controller, the coordinators get the same treatment.
- Snapshots (spec 013). We plan to publish state snapshots so new nodes can start from a recent state instead of syncing from scratch.
- History off the SSD (#1483). Block history moves to slower storage so the SSD is left for state.
The Sepolia node is still syncing. We’ll post again when it reaches head and the first full sync on the split controller has run.

