Data 2026-09-30 10:47 → 2026-10-03 14:06 GST (UTC+4) · 4327 one-minute samples · 33 failures (deduped, onset time) · generated 2026-10-03T14:07:57+04:00.
Lost = no reply to 1 ping in 2 s; slow = RTT > 100 ms; DNS fail = router DNS failures counted in that minute; load = 1-min loadavg > 25.
Lift = signal rate in the window before failures ÷ rate across all minutes. p_perm = permutation test (5000 random circular shifts of all failure times together); Fisher p assumes minutes are independent (too optimistic for bursty signals).
Conclusion
Router update (3 Oct, read-only admin check): the router's connection table (nf_conntrack) is full. Every router-log sample shows "nf_conntrack: table full, dropping packet", about 1–2.5 logged drops per second plus more hidden by log rate limiting, and the PPP link has discarded 36.8 M inbound packets. That makes connection-table exhaustion the leading router-stall cause. Reboots, double NAT, old firmware and a bad Studio cable are ruled out. The router↔fibre-box port (11.2 M transmit errors) is a secondary suspect. Two new monitors (router-watch every 5 min and a 5 Hz stall detector) will show whether these drops spike before Grok Bot failures. See the table near the end of the page.
Does speed line up with failures? No. Transmission download speed is slightly lower in the 15 minutes before failures (0.92× normal, p=0.54), upload is 0.88× (p=0.45), and peer count and active torrents are flat (1.00×, p≥0.3). Day by day it runs the other way: Sep 30–Oct 1 had the highest speeds (13–25 MB/s, ~300 peers) and fewer failures, while Oct 2–3 had lower speeds (~5–7 MB/s, ~110 peers) and more failures. qBittorrent on the Mini, Studio-wide socket counts and en0 throughput had no history, so they only start now (see below).
The signals that best precede failures are Tailscale path problems: home DERP relay switches (2.2× normal, p=0.004) and "UDP is blocked" events (1.9×, p=0.005). Router DNS failures (2.25×, p=0.06) and router ping RTT (1.4×, p=0.10) come next but are borderline, and lost router pings are weak (1.4×, p=0.19). Load, swap, memory pressure, OrbStack CPU and torrent traffic show no relationship. With ~20 signals tested, p≈0.005 is suggestive rather than proof, and none of these signals predicts a drop reliably on its own.
Tracked going forward (com.anthony.cause-tracker, every 30 s, on Studio): (1) router ping, 5 pings per sample, loss and max RTT; (2) router DNS and 1.1.1.1 DNS times plus router DNS-fail count; (3) internet ping loss/RTT to 1.1.1.1 and 8.8.8.8; (4) Transmission down/up; (5) Transmission peers and active torrents; (6) qBittorrent on the Mini down/up, active torrents and peers; (7) Studio TCP total/ESTABLISHED/SYN_SENT, UDP sockets and top-3 processes by socket count; (8) en0 in/out MB/s and errors; (9) load, total CPU, swap, memory pressure, top CPU processes, OrbStack/Plex CPU and Plex transcoders; (10) Tailscale state (relay, box link direct or relayed, netcheck UDP every ~2 min) plus Grok Bot worker count. The ranking will firm up as new failures build up.
| # | signal | source | minutes with data | failures covered (15m) | note |
| 1 | Router loss % (5 pings/30 s) | cause-tracker | 30 | 0 | |
| 1 | Router max RTT (5 pings) | cause-tracker | 30 | 0 | |
| 1 | Router conntrack-full drops/s (router log) | router-watch | 4 | 0 | |
| 1 | Router stall seconds/min (5 Hz) | router-stall | 13 | 0 | |
| 1 | Estimated router flows (Studio+Mini, lower bound) | cause-tracker | 11 | 0 | |
| 1 | Studio internet TCP flows | cause-tracker | 11 | 0 | |
| 2 | 1.1.1.1 DNS lookup time | cause-tracker | 30 | 0 | |
| 3 | Internet ping loss % (1.1.1.1+8.8.8.8) | cause-tracker | 30 | 0 | |
| 3 | Internet ping RTT 1.1.1.1 | cause-tracker | 28 | 0 | |
| 3 | Router memory % | router-watch | 2 | 0 | |
| 6 | qBittorrent (Mini) download | cause-tracker | 30 | 0 | |
| 6 | qBittorrent (Mini) upload | cause-tracker | 30 | 0 | |
| 6 | qBittorrent peers (active torrents) | cause-tracker | 30 | 0 | |
| 6 | qBittorrent active torrents | cause-tracker | 30 | 0 | |
| 7 | Studio TCP sockets (all) | cause-tracker | 30 | 0 | |
| 7 | Studio TCP ESTABLISHED | cause-tracker | 30 | 0 | |
| 7 | Studio TCP SYN_SENT | cause-tracker | 30 | 0 | |
| 7 | Studio UDP sockets | cause-tracker | 30 | 0 | |
| 8 | en0 in | cause-tracker | 26 | 0 | |
| 8 | en0 out | cause-tracker | 26 | 0 | |
| 8 | en0 errors/s | cause-tracker | 26 | 0 | |
| 9 | Studio total CPU % | cause-tracker | 30 | 0 | |
| 9 | Plex transcoders running | cause-tracker | 30 | 0 | |
| 9 | ISP-hop stall seconds/min | router-stall | 13 | 0 | |
| 9 | 1.1.1.1 stall seconds/min | router-stall | 13 | 0 | |
| 10 | Box link relayed (not direct) | hop | 3433 | 33 | confounded: the box's tailnet link is driven by agents' own SSH sessions (incl. heals) (lift 0.00×, p=0.0004) |
| 10 | Tailscale status timeout (cause-tracker) | cause-tracker | 30 | 0 | |
| 10 | Box link relayed (cause-tracker) | cause-tracker | 4 | 0 | confounded (same as above) |
| 10 | Tailscale netcheck UDP=false | cause-tracker | 1 | 0 | |
| # | Candidate cause | Verdict | Evidence (read-only, 3 Oct 13:40–13:56 GST) |
| 1 | NAT / connection-table (nf_conntrack) exhaustion | LIKELY: table is full now, and a router stall was seen during a flow burst | Router kernel log shows "nf_conntrack: table full, dropping packet" in every Log Viewer sample, about 0.7–2.5 logged drops/s, plus up to 135 more hidden by log rate limiting per 5 s. The PPP link has 36.8 M inbound packets discarded (~75/s). The GUI shows no session count or maximum. Visible flows: Studio 400–4,200 internet TCP (bursts of up to 2,800 half-open SYN_SENT, mostly to ports 50300/2234 = Soulseek/slskd in OrbStack) + ~220 UDP; Transmission 100–160 peers, with DHT/PEX/LPD on 309 torrents (many UDP flows we can't see); qBittorrent on the Mini only 3–7 peers / 52 TCP; Mac Pro runs another torrent client (port 6881 mapping). Typical maximum for this class of router is 4k–16k entries. At 13:57:21 Studio's flow estimate jumped from ~450 to 1,908, and the router itself stopped answering (router stalls of 1.0–1.8 s, ISP hop and 1.1.1.1 stalls of 2–4.5 s, and a 5-ping burst losing 80–100%) from 13:56:55 to 13:57:42. Not yet matched against Grok Bot failures: the router log only holds ~25 s, so router-watch now samples it every 5 min. |
| 2 | Weak / overheating hardware | CAN'T TELL (no temperature or CPU readout) | The GUI shows no CPU % or temperature. Pings from Studio to the router itself sometimes take 300–600 ms (normally ~1 ms), so the router CPU is busy at times. That fits #1 (connection-table churn) as well as weak hardware. |
| 3 | Memory leak / old firmware (worse with uptime) | UNLIKELY | Firmware built April 2026. Memory 59% and steady. Failures don't grow with uptime: the router has been up the whole 12.6 days and failures run 5–10 a day. |
| 4 | Spontaneous reboots / crashes | RULED OUT for Sep 30 – Oct 3 | Device uptime is 12 d 15 h and the WAN connection has been up 1,093,186 s (same boot, ~Sep 20 GST). None of the 33 failures came with a reboot. |
| 5 | Router DNS forwarder (dnsmasq) choking | POSSIBLE, probably a symptom | Router DNS failures are 2.25× normal before failures (p=0.06). No dnsmasq errors appear in the log samples. DNS over UDP gets dropped first when the connection table is full. |
| 6 | CPU-heavy features (inspection / parental / QoS / IPv6) | UNLIKELY | Firewall level Medium (default, stateful). Parental: access control on, but 0 site rules and 0 time-of-day rules. No QoS page. IPv6 dual-stack on. EasyMesh controller and Samba on. Nothing unusual. |
| 7 | UPnP / NAT-PMP port-mapping churn | POSSIBLE, minor | UPnP and NAT-PMP are on with 20 mappings: Plex ×11 from 7 devices, Transmission, Tailscale ×3, Mac Pro 6881. The log shows "unsupported NAT-PMP version: 2" every few seconds (Tailscale probing). It's churn, but low volume. |
| 8 | Wi-Fi client load | UNLIKELY for Studio | Studio is wired (1 GbE). 5 Wi-Fi clients. The Wi-Fi radios do show high error counts (5 GHz: 5.8 M rx / 12 M tx errors) and mesh steering, which share the router CPU. |
| 9 | WAN / ISP resyncs or congestion | Resyncs RULED OUT; WAN-side gaps POSSIBLE | PPPoE has been up since boot, last connection error NONE, PPP errors 0. Before failures, internet latency is only 1.11× normal (not significant). But the WAN port to the fibre box (Port 4, ae_wan) has 11.2 M transmit errors out of 625 M packets (1.8%). It is still happening: 4,439 more tx errors out of 348k packets (1.3%) between 13:52 and 14:07. The new stall detector saw a 2.4 s gap at the ISP hop and 1.1.1.1 at 13:55:01 with no gap to the router. |
| 10 | Bad cable / port / switch | Studio↔router RULED OUT; router↔fibre-box link POSSIBLE | Studio en0: 0 errors, 0 collisions, 1000baseT full duplex. LAN ports 1–2: 0 errors. WAN Port 4: 11.2 M tx errors (see #9): reseat or replace the router↔fibre-box cable. |
| 11 | Power issues | UNLIKELY recently | No reboot in 12.6 days. The last reboot cause was "Power" (~Sep 20). Brownouts that don't reboot it can't be seen. |
| 12 | Broadcast storm / misbehaving device / loop | POSSIBLE (odd LAN traffic) | 22 devices, 40 ARP entries (normal), no storm messages in the log. But over 15 min the router sent ~10,500 packets/s out of LAN Port 2 while only ~600 packets/s came in from the internet. That much traffic going through the router's CPU but not coming from the internet suggests hairpin NAT (LAN devices reaching Plex/Transmission/Tailscale via the public IP), multicast flooding, or Studio's second (Wi-Fi) IP. Worth a look; it also uses up connection-table slots. |
| 13 | Double NAT | RULED OUT (for Studio) | The router's WAN IP 2.49.151.241 is public, matches the public IPv4 from ifconfig.me, and PPPoE ends on the router. (Only OrbStack containers sit behind a second, local NAT.) |