Mac Studio → router ping vs Grok Bot connection drops

Data 2026-09-30 10:47 → 2026-10-03 14:06 GST (UTC+4) · 4327 one-minute samples · 33 failures (deduped, onset time) · generated 2026-10-03T14:07:57+04:00. Lost = no reply to 1 ping in 2 s; slow = RTT > 100 ms; DNS fail = router DNS failures counted in that minute; load = 1-min loadavg > 25. Lift = signal rate in the window before failures ÷ rate across all minutes. p_perm = permutation test (5000 random circular shifts of all failure times together); Fisher p assumes minutes are independent (too optimistic for bursty signals).

Conclusion

Router update (3 Oct, read-only admin check): the router's connection table (nf_conntrack) is full. Every router-log sample shows "nf_conntrack: table full, dropping packet", about 1–2.5 logged drops per second plus more hidden by log rate limiting, and the PPP link has discarded 36.8 M inbound packets. That makes connection-table exhaustion the leading router-stall cause. Reboots, double NAT, old firmware and a bad Studio cable are ruled out. The router↔fibre-box port (11.2 M transmit errors) is a secondary suspect. Two new monitors (router-watch every 5 min and a 5 Hz stall detector) will show whether these drops spike before Grok Bot failures. See the table near the end of the page.

Does speed line up with failures? No. Transmission download speed is slightly lower in the 15 minutes before failures (0.92× normal, p=0.54), upload is 0.88× (p=0.45), and peer count and active torrents are flat (1.00×, p≥0.3). Day by day it runs the other way: Sep 30–Oct 1 had the highest speeds (13–25 MB/s, ~300 peers) and fewer failures, while Oct 2–3 had lower speeds (~5–7 MB/s, ~110 peers) and more failures. qBittorrent on the Mini, Studio-wide socket counts and en0 throughput had no history, so they only start now (see below).

The signals that best precede failures are Tailscale path problems: home DERP relay switches (2.2× normal, p=0.004) and "UDP is blocked" events (1.9×, p=0.005). Router DNS failures (2.25×, p=0.06) and router ping RTT (1.4×, p=0.10) come next but are borderline, and lost router pings are weak (1.4×, p=0.19). Load, swap, memory pressure, OrbStack CPU and torrent traffic show no relationship. With ~20 signals tested, p≈0.005 is suggestive rather than proof, and none of these signals predicts a drop reliably on its own.

Tracked going forward (com.anthony.cause-tracker, every 30 s, on Studio): (1) router ping, 5 pings per sample, loss and max RTT; (2) router DNS and 1.1.1.1 DNS times plus router DNS-fail count; (3) internet ping loss/RTT to 1.1.1.1 and 8.8.8.8; (4) Transmission down/up; (5) Transmission peers and active torrents; (6) qBittorrent on the Mini down/up, active torrents and peers; (7) Studio TCP total/ESTABLISHED/SYN_SENT, UDP sockets and top-3 processes by socket count; (8) en0 in/out MB/s and errors; (9) load, total CPU, swap, memory pressure, top CPU processes, OrbStack/Plex CPU and Plex transcoders; (10) Tailscale state (relay, box link direct or relayed, netcheck UDP every ~2 min) plus Grok Bot worker count. The ranking will firm up as new failures build up.

cause ranking causes timeline pre-failure signal rates per-day ping timeline

Per day

daysampleslostslow>100msDNS-fail minload>25 minfailures
2026-09-307702157331458
2026-10-0113762561144195
2026-10-021377511284551410
2026-10-0380418453039810

Pre-failure lift (rate in window / baseline rate), permutation p (circular shift) and Fisher p

signalbaselineWpre-fail rateliftp_permFisher pfailures w/ ≥1 (exp.)
Lost ping2.66%5m3.18%1.20×0.32070.40625 (4.0)
Lost ping2.66%15m3.77%1.42×0.10060.100414 (10.6)
Lost ping2.66%30m3.74%1.41×0.04820.025620 (17.5)
Slow ping >100ms6.73%5m14.01%2.08×0.00140.000714 (8.0)
Slow ping >100ms6.73%15m10.48%1.56×0.01320.009122 (16.8)
Slow ping >100ms6.73%30m11.02%1.64×0.00040.000030 (23.0)
Router DNS fail2.82%5m4.52%1.60×0.16540.14653 (3.1)
Router DNS fail2.82%15m6.30%2.23×0.05320.000011 (7.4)
Router DNS fail2.82%30m5.88%2.08×0.06640.000018 (12.2)
Load >2534.14%5m39.49%1.16×0.21020.088918 (18.8)
Load >2534.14%15m40.25%1.18×0.10300.001027 (28.0)
Load >2534.14%30m38.82%1.14×0.13080.011428 (29.9)

Reverse: signal minutes followed by a failure within 15 min

signalminutesfollowedfracbase fraclift
Lost ping1151714.8%10.7%1.39×
Slow ping >100ms2914415.1%10.7%1.42×
Router DNS fail1222923.8%10.6%2.23×
Load >25147618812.7%10.7%1.19×

Failures (onset GST)

Candidate causes ranked by how well they precede failures (15-min window, n failures with data ≥ 8)

rank#signalsourcepre-fail mean (15m)normal meanlift 5mlift 15mlift 30mp 15m (two-sided)directiontop-quartile→fail ≤15m lift
110Tailscale home DERP switchedhop0.0780.0352.63×2.21×1.75×0.0048higher2.02×
210Tailscale UDP blockedstream log0.1520.0802.35×1.90×1.72×0.0056higher1.72×
37Grok Bot open TCP connshop9.1378.6001.09×1.06×1.06×0.0284higher1.55×
42Router DNS failureshop0.0630.0281.61×2.24×2.10×0.0532higher2.23×
51Router ping RTThop46.60532.7031.96×1.43×1.53×0.0980higher1.14×
69Plex CPU % (top-5 only)hop40.61525.3501.99×1.60×1.73×0.1340higher1.47×
73Internet HTTPS latencyhop631.224568.7761.30×1.11×1.08×0.1564higher1.08×
81Router ping loss (1 ping)hop0.0380.0261.21×1.43×1.42×0.2000higher1.39×
910Tailscale self offline/unknownhop0.0050.0010.00×4.33×2.20×0.2288higher4.52×
1010Tailscale status call timed out/failedhop0.1660.2010.88×0.82×0.82×0.2599lower0.85×
115Transmission active torrentstr-metrics.db30.08030.2031.00×1.00×1.00×0.3311lower0.73×
129OrbStack CPU %hop174.281153.7711.18×1.13×1.22×0.3503higher1.22×
133Internet HTTPS check failedhop0.0630.0511.23×1.23×1.49×0.4147higher1.11×
144Transmission uploadtr-metrics.db0.4230.4770.88×0.89×0.87×0.4775lower0.79×
152Router DNS lookup timehop488.028453.0281.12×1.08×1.06×0.4819higher1.07×
164Transmission downloadtr-metrics.db9.0989.7990.97×0.93×0.94×0.5699lower0.73×
179Swap usedhop2209.2442358.1430.94×0.94×0.92×0.7770lower0.74×
189Load average (1 min)hop19.07619.1220.98×1.00×0.99×0.8882lower1.28×
199Memory pressure warn/criticalhop0.0020.0170.00×0.12×0.37×0.8886lower0.13×
205Transmission peers connectedtr-metrics.db159.900159.3390.99×1.00×1.01×0.9490higher1.18×

Tracked but not enough failures yet (new cause-tracker / no history)

#signalsourceminutes with datafailures covered (15m)note
1Router loss % (5 pings/30 s)cause-tracker300
1Router max RTT (5 pings)cause-tracker300
1Router conntrack-full drops/s (router log)router-watch40
1Router stall seconds/min (5 Hz)router-stall130
1Estimated router flows (Studio+Mini, lower bound)cause-tracker110
1Studio internet TCP flowscause-tracker110
21.1.1.1 DNS lookup timecause-tracker300
3Internet ping loss % (1.1.1.1+8.8.8.8)cause-tracker300
3Internet ping RTT 1.1.1.1cause-tracker280
3Router memory %router-watch20
6qBittorrent (Mini) downloadcause-tracker300
6qBittorrent (Mini) uploadcause-tracker300
6qBittorrent peers (active torrents)cause-tracker300
6qBittorrent active torrentscause-tracker300
7Studio TCP sockets (all)cause-tracker300
7Studio TCP ESTABLISHEDcause-tracker300
7Studio TCP SYN_SENTcause-tracker300
7Studio UDP socketscause-tracker300
8en0 incause-tracker260
8en0 outcause-tracker260
8en0 errors/scause-tracker260
9Studio total CPU %cause-tracker300
9Plex transcoders runningcause-tracker300
9ISP-hop stall seconds/minrouter-stall130
91.1.1.1 stall seconds/minrouter-stall130
10Box link relayed (not direct)hop343333confounded: the box's tailnet link is driven by agents' own SSH sessions (incl. heals) (lift 0.00×, p=0.0004)
10Tailscale status timeout (cause-tracker)cause-tracker300
10Box link relayed (cause-tracker)cause-tracker40confounded (same as above)
10Tailscale netcheck UDP=falsecause-tracker10

Router stall causes: rule in / rule out (router: Etisalat/Vantiva FGA228BETI, firmware 23.2.a.0655 built 2026-04-18, up 12 d 15 h, last reboot cause "Power", memory 59%)

#Candidate causeVerdictEvidence (read-only, 3 Oct 13:40–13:56 GST)
1NAT / connection-table (nf_conntrack) exhaustionLIKELY: table is full now, and a router stall was seen during a flow burstRouter kernel log shows "nf_conntrack: table full, dropping packet" in every Log Viewer sample, about 0.7–2.5 logged drops/s, plus up to 135 more hidden by log rate limiting per 5 s. The PPP link has 36.8 M inbound packets discarded (~75/s). The GUI shows no session count or maximum. Visible flows: Studio 400–4,200 internet TCP (bursts of up to 2,800 half-open SYN_SENT, mostly to ports 50300/2234 = Soulseek/slskd in OrbStack) + ~220 UDP; Transmission 100–160 peers, with DHT/PEX/LPD on 309 torrents (many UDP flows we can't see); qBittorrent on the Mini only 3–7 peers / 52 TCP; Mac Pro runs another torrent client (port 6881 mapping). Typical maximum for this class of router is 4k–16k entries. At 13:57:21 Studio's flow estimate jumped from ~450 to 1,908, and the router itself stopped answering (router stalls of 1.0–1.8 s, ISP hop and 1.1.1.1 stalls of 2–4.5 s, and a 5-ping burst losing 80–100%) from 13:56:55 to 13:57:42. Not yet matched against Grok Bot failures: the router log only holds ~25 s, so router-watch now samples it every 5 min.
2Weak / overheating hardwareCAN'T TELL (no temperature or CPU readout)The GUI shows no CPU % or temperature. Pings from Studio to the router itself sometimes take 300–600 ms (normally ~1 ms), so the router CPU is busy at times. That fits #1 (connection-table churn) as well as weak hardware.
3Memory leak / old firmware (worse with uptime)UNLIKELYFirmware built April 2026. Memory 59% and steady. Failures don't grow with uptime: the router has been up the whole 12.6 days and failures run 5–10 a day.
4Spontaneous reboots / crashesRULED OUT for Sep 30 – Oct 3Device uptime is 12 d 15 h and the WAN connection has been up 1,093,186 s (same boot, ~Sep 20 GST). None of the 33 failures came with a reboot.
5Router DNS forwarder (dnsmasq) chokingPOSSIBLE, probably a symptomRouter DNS failures are 2.25× normal before failures (p=0.06). No dnsmasq errors appear in the log samples. DNS over UDP gets dropped first when the connection table is full.
6CPU-heavy features (inspection / parental / QoS / IPv6)UNLIKELYFirewall level Medium (default, stateful). Parental: access control on, but 0 site rules and 0 time-of-day rules. No QoS page. IPv6 dual-stack on. EasyMesh controller and Samba on. Nothing unusual.
7UPnP / NAT-PMP port-mapping churnPOSSIBLE, minorUPnP and NAT-PMP are on with 20 mappings: Plex ×11 from 7 devices, Transmission, Tailscale ×3, Mac Pro 6881. The log shows "unsupported NAT-PMP version: 2" every few seconds (Tailscale probing). It's churn, but low volume.
8Wi-Fi client loadUNLIKELY for StudioStudio is wired (1 GbE). 5 Wi-Fi clients. The Wi-Fi radios do show high error counts (5 GHz: 5.8 M rx / 12 M tx errors) and mesh steering, which share the router CPU.
9WAN / ISP resyncs or congestionResyncs RULED OUT; WAN-side gaps POSSIBLEPPPoE has been up since boot, last connection error NONE, PPP errors 0. Before failures, internet latency is only 1.11× normal (not significant). But the WAN port to the fibre box (Port 4, ae_wan) has 11.2 M transmit errors out of 625 M packets (1.8%). It is still happening: 4,439 more tx errors out of 348k packets (1.3%) between 13:52 and 14:07. The new stall detector saw a 2.4 s gap at the ISP hop and 1.1.1.1 at 13:55:01 with no gap to the router.
10Bad cable / port / switchStudio↔router RULED OUT; router↔fibre-box link POSSIBLEStudio en0: 0 errors, 0 collisions, 1000baseT full duplex. LAN ports 1–2: 0 errors. WAN Port 4: 11.2 M tx errors (see #9): reseat or replace the router↔fibre-box cable.
11Power issuesUNLIKELY recentlyNo reboot in 12.6 days. The last reboot cause was "Power" (~Sep 20). Brownouts that don't reboot it can't be seen.
12Broadcast storm / misbehaving device / loopPOSSIBLE (odd LAN traffic)22 devices, 40 ARP entries (normal), no storm messages in the log. But over 15 min the router sent ~10,500 packets/s out of LAN Port 2 while only ~600 packets/s came in from the internet. That much traffic going through the router's CPU but not coming from the internet suggests hairpin NAT (LAN devices reaching Plex/Transmission/Tailscale via the public IP), multicast flooding, or Studio's second (Wi-Fi) IP. Worth a look; it also uses up connection-table slots.
13Double NATRULED OUT (for Studio)The router's WAN IP 2.49.151.241 is public, matches the public IPv4 from ifconfig.me, and PPPoE ends on the router. (Only OrbStack containers sit behind a second, local NAT.)

Live router monitoring so far (router-watch every 5 min, router-stall continuous)

metricvalue
router log samples6 (2026-10-03T13:52 → 2026-10-03T14:07)
samples showing 'conntrack table full'6 of 6 (100%)
estimated conntrack drops/s (median / max)2.19 / 3.24
stalls ≥1 s: router13 (total 17.8 s, longest 2.0 s)
stalls ≥1 s: isp_hop31 (total 95.0 s, longest 7.7 s)
stalls ≥1 s: cloudflare21 (total 70.9 s, longest 10.5 s)