glamsterdam-devnet-9 · incident report · index of nine

Gloas fork transition failure

P0 — Network not finalizing
Component whole network Observed 2026-09-03 08:30–09:15 UTC Network glamsterdam-devnet-9 (chain 7013099983) Fork Gloas @ epoch 225

glamsterdam-devnet-9 has not finalized since epoch 223. Participation was 95% at epoch 224 and collapsed to 56% at epoch 225 — the Gloas fork epoch — and has been flat at 55–63% for over 160 epochs. This is not the devnet's planned non-finality test. Two consensus clients failed the fork transition outright, ~34% of the active stake has been offline since, and the network cannot reach the 2/3 threshold until Prysm is fixed.

Nine reports

P0 — Consensus split Prysm / consensus layer

Prysm fails the Gloas state transition

307/310 Prysm nodes plus all 4 bootnodes frozen at slot 7199. A full config-spec diff rules out misconfiguration — only EIP-55 casing differs — so this is a genuine state-transition bug, bounded to process_slots(7199→7203) + process_block(7203). The 312 stuck nodes span all seven ELs, ruling out the execution layer. 31% of stake offline.

P0 — Client cannot follow the chain Grandine / consensus layer

Grandine cannot cross the fork — sync livelock and OOM

Grandine requests zero payload envelopes for the empty fork-boundary slot and retries forever, while every one of its 10 hosts is kernel-OOM-killed at 30–31 GB RSS on 32 GB boxes.

P0 — Devnet validity fleet image configuration

The fleet is not running the images the spec sheet says it is

Grandine was left behind by the "bump images" commit - an omission, with the spec sheet checklist still unticked. Lodestar is different: it was deliberately reverted from :unstable to a devnet-8 tag 3 minutes 40 seconds before the fork, undocumented, with all 31 hosts redeployed in a 10-minute window. Reth and Nimbus-EL get a clean bill of health.

P1 — Degraded duties + P0 finding Teku / consensus layer

Teku rejects canonical blocks, and its attestation subnets are starved

Twelve Teku hosts transiently mark canonical blocks as failed validation and reject every attestation voting for them — 75,380 gossip rejections in 3 h, with seven rejected roots all confirmed canonical against a healthy node. Separately, the attestation-subnet starvation is independent of heap pressure and is caused by peer-table poisoning.

P1 — Headline feature untested buildoor / EIP-7732 ePBS builder path

All four builder nodes are dead — the ePBS path has never worked

A one-word omission in a terraform allow-list put the four builders on 16 GB hosts while all 1,000 validator nodes got 32 GB. They have never submitted a single bid, from genesis onward: 20 sampled post-fork slots give builder_built=0, every block carrying builder_index=UINT64_MAX.

P1 — Measurement integrity fleet operations / watchtower

Watchtower swapped client images across the fleet, mid-fork

Watchtower replaced client containers 4,694 times across 972 of 1,008 hosts, including 763 hosts in the 75 minutes before the fork. Prysm's image did not change, so its bug is genuine; Grandine's changed twice, so the binary that failed the fork and the one now crash-looping are different images.

P1 — Capacity fleet capacity / all consensus clients

Fleet memory exhaustion at a 4M-entry validator registry

Prysm acquires 11.16 GiB in the single fork hour and never releases it, sitting at 23-27 GiB for 17 h; Lighthouse on identical hosts took +2.54 GiB. Nodes restarted after the fork stay flat at 12-15 GiB, so it is the fork transition, not the retry loop. A rolling restart of the 61 hosts above 22 GiB reclaims ~10 GiB each and stops the next OOM wave today.

P2 — Observability and tooling checkpointz, forky, erpc, dora, slashoor, tracoor

Service tier: four independent bugs behind the noise

An error-regex over log bodies ranks this tier exactly backwards: the loudest workloads are healthy and the quietest are down. Six failures are consequences of the chain split and self-resolve; four are independent bugs — missing TLS certs on healthy rpc endpoints, checkpointz pinned to dead upstreams, forky unable to parse Lighthouse fork-choice, and forkmon 503.

P2 — Investigated, largely negative p2p gossip / all consensus clients

Gossip degradation is CPU exhaustion, not peer-table poisoning

The poisoning hypothesis is falsified, not merely unproven: r(stuck-peer share, publish failures) = -0.271 while r(CPU%, publish failures) = +0.931. Peer scoring works across all four CLs. The fork digest is derived cryptographically and is correct, so discovery-level isolation of the stuck nodes was structurally impossible - not a misconfiguration.

Why the network cannot finalize

Active set ~1,001,087 validators (eligible ether 32,034,782 ETH), 1000 per node.

CohortValidatorsShare
Prysm frozen at slot 7199 (307 nodes)307,00030.7%
Grandine (1 stuck + 9 OOM-looping)10,0001.0%
Prysm down / unreachable (3)3,0000.3%
Teku beacons unresponsive (13)13,0001.3%
Lighthouse (7) + Lodestar (3) down10,0001.0%
Offline outright343,00034.3%
Gossip/CPU degradation on live hosts~2.4%

That gives a practical ceiling of ~63.3% against a 66.7% threshold, versus an observed mean of 56.4%. Block-inclusion capacity explains part of the remaining gap (r(blocks proposed, participation) = +0.546 over 16 epochs, ~30% of variance); the rest is reported as unresolved rather than assigned. Fixing Prysm is the only lever on finality — peer scoring already isolates the stuck nodes, so stopping them would not measurably help the healthy population.

Fleet snapshot

1008 hosts, swept 2026-09-03 08:30–09:00 UTC.

Consensus clientAt tipStuck @ 7199 Beacon downHosts OOM-killed
lighthouse423073 / 430
prysm03073255 / 308
teku1260141 / 140
nimbus80001 / 80
lodestar27030 / 30
grandine01910 / 10
bootnode (runs Prysm)0401 / 4
buildoor0024 / 4

275 of 1007 hosts have kernel OOM kills. Victims: beacon-chain (Prysm) 257, grandine 10, lighthouse 4, java (Teku) 2, nimbus_beacon_n 1. Note docker inspect reports every one of these as OOMKilled:false ExitCode:0 — only dmesg reveals them.

What is healthy

Three corrections we made to our own findings

Recorded because each one looked convincing before it was tested.

Order of action

  1. Prysm — run the offline process_slots(7199→7200) test. This alone blocks finality.
  2. Rolling-restart the 61 Prysm hosts ≥22 GiB — reclaims ~10 GiB each and stops the next OOM wave, independent of the state-root fix.
  3. Bootnodes — they run Prysm and sit on the dead fork.
  4. Image tags — put Grandine on trunk; establish why Lodestar was reverted 3m40s before the fork.
  5. Teku — the canonical-block rejection is a P0 for the Teku team.
  6. Pin images for the remainder of the devnet.
  7. Buildoors — add buildoor to the terraform supernode regex.
  8. checkpointz — repoint at healthy nodes and add failover.