glamsterdam-devnet-9 has not finalized since epoch 223. Participation was 95% at epoch 224 and collapsed to 56% at epoch 225 — the Gloas fork epoch — and has been flat at 55–63% for over 160 epochs. This is not the devnet's planned non-finality test. Two consensus clients failed the fork transition outright, ~34% of the active stake has been offline since, and the network cannot reach the 2/3 threshold until Prysm is fixed.
307/310 Prysm nodes plus all 4 bootnodes frozen at slot 7199. A full config-spec diff rules out misconfiguration — only EIP-55 casing differs — so this is a genuine state-transition bug, bounded to process_slots(7199→7203) + process_block(7203). The 312 stuck nodes span all seven ELs, ruling out the execution layer. 31% of stake offline.
Grandine requests zero payload envelopes for the empty fork-boundary slot and retries forever, while every one of its 10 hosts is kernel-OOM-killed at 30–31 GB RSS on 32 GB boxes.
Grandine was left behind by the "bump images" commit - an omission, with the spec sheet checklist still unticked. Lodestar is different: it was deliberately reverted from :unstable to a devnet-8 tag 3 minutes 40 seconds before the fork, undocumented, with all 31 hosts redeployed in a 10-minute window. Reth and Nimbus-EL get a clean bill of health.
Twelve Teku hosts transiently mark canonical blocks as failed validation and reject every attestation voting for them — 75,380 gossip rejections in 3 h, with seven rejected roots all confirmed canonical against a healthy node. Separately, the attestation-subnet starvation is independent of heap pressure and is caused by peer-table poisoning.
A one-word omission in a terraform allow-list put the four builders on 16 GB hosts while all 1,000 validator nodes got 32 GB. They have never submitted a single bid, from genesis onward: 20 sampled post-fork slots give builder_built=0, every block carrying builder_index=UINT64_MAX.
Watchtower replaced client containers 4,694 times across 972 of 1,008 hosts, including 763 hosts in the 75 minutes before the fork. Prysm's image did not change, so its bug is genuine; Grandine's changed twice, so the binary that failed the fork and the one now crash-looping are different images.
Prysm acquires 11.16 GiB in the single fork hour and never releases it, sitting at 23-27 GiB for 17 h; Lighthouse on identical hosts took +2.54 GiB. Nodes restarted after the fork stay flat at 12-15 GiB, so it is the fork transition, not the retry loop. A rolling restart of the 61 hosts above 22 GiB reclaims ~10 GiB each and stops the next OOM wave today.
An error-regex over log bodies ranks this tier exactly backwards: the loudest workloads are healthy and the quietest are down. Six failures are consequences of the chain split and self-resolve; four are independent bugs — missing TLS certs on healthy rpc endpoints, checkpointz pinned to dead upstreams, forky unable to parse Lighthouse fork-choice, and forkmon 503.
The poisoning hypothesis is falsified, not merely unproven: r(stuck-peer share, publish failures) = -0.271 while r(CPU%, publish failures) = +0.931. Peer scoring works across all four CLs. The fork digest is derived cryptographically and is correct, so discovery-level isolation of the stuck nodes was structurally impossible - not a misconfiguration.
Active set ~1,001,087 validators (eligible ether 32,034,782 ETH), 1000 per node.
| Cohort | Validators | Share |
|---|---|---|
| Prysm frozen at slot 7199 (307 nodes) | 307,000 | 30.7% |
| Grandine (1 stuck + 9 OOM-looping) | 10,000 | 1.0% |
| Prysm down / unreachable (3) | 3,000 | 0.3% |
| Teku beacons unresponsive (13) | 13,000 | 1.3% |
| Lighthouse (7) + Lodestar (3) down | 10,000 | 1.0% |
| Offline outright | 343,000 | 34.3% |
| Gossip/CPU degradation on live hosts | — | ~2.4% |
That gives a practical ceiling of ~63.3% against a 66.7% threshold, versus an observed mean of 56.4%. Block-inclusion capacity explains part of the remaining gap (r(blocks proposed, participation) = +0.546 over 16 epochs, ~30% of variance); the rest is reported as unresolved rather than assigned. Fixing Prysm is the only lever on finality — peer scoring already isolates the stuck nodes, so stopping them would not measurably help the healthy population.
1008 hosts, swept 2026-09-03 08:30–09:00 UTC.
| Consensus client | At tip | Stuck @ 7199 | Beacon down | Hosts OOM-killed |
|---|---|---|---|---|
| lighthouse | 423 | 0 | 7 | 3 / 430 |
| prysm | 0 | 307 | 3 | 255 / 308 |
| teku | 126 | 0 | 14 | 1 / 140 |
| nimbus | 80 | 0 | 0 | 1 / 80 |
| lodestar | 27 | 0 | 3 | 0 / 30 |
| grandine | 0 | 1 | 9 | 10 / 10 |
| bootnode (runs Prysm) | 0 | 4 | 0 | 1 / 4 |
| buildoor | 0 | 0 | 2 | 4 / 4 |
275 of 1007 hosts have kernel OOM kills. Victims: beacon-chain (Prysm) 257,
grandine 10, lighthouse 4, java (Teku) 2,
nimbus_beacon_n 1. Note docker inspect reports every one of these as
OOMKilled:false ExitCode:0 — only dmesg reveals them.
Recorded because each one looked convincing before it was tested.
ansible sweep appeared to show nethermind and reth running
5–16 slots behind geth, in a clean split by EL. It was an artifact of alphabetical probe
ordering. A simultaneous probe shows no difference.process_slots(7199→7200) test. This alone
blocks finality.buildoor to the terraform supernode
regex.