Fix relay watchdog false-positive: sudo the pkill

restart_relay() ran pkill without sudo, but relay processes are
root-owned and the timer runs as User=super. The kill failed EPERM
(silently swallowed by || true), the old relay kept running, and
SO_REUSEADDR let the replacement double-bind the same port. The
post-restart health check then passed and logged "restarted OK"
when nothing was actually restarted.

Adding sudo -n to the pkill, matching the sudo -n ip netns exec
already used to start the relay.

Session: sidechat/chromebox-ops
This commit is contained in:
operator-main
2026-10-04 18:22:10 +00:00
parent c07802dfa2
commit ae1bf17de5
+4 -1
View File
@@ -84,7 +84,10 @@ restart_relay() {
port="${target##*:}" port="${target##*:}"
netns="warp-$node" netns="warp-$node"
# Kill any existing relay for this node's port (correct or not) # Kill any existing relay for this node's port (correct or not)
pkill -f "netvm-cdp-relay.py .* $port 127.0.0.1 $port" 2>/dev/null || true # Relays are root-owned (started via sudo ip netns exec); the timer runs as
# super, so the kill needs sudo too. Without it pkill fails EPERM silently
# and the "restart" false-positives via SO_REUSEADDR double-bind.
sudo -n pkill -f "netvm-cdp-relay.py .* $port 127.0.0.1 $port" 2>/dev/null || true
sleep 2 sleep 2
# Launch inside the netns, listening on the veth IP (host-reachable) # Launch inside the netns, listening on the veth IP (host-reachable)
sudo -n ip netns exec "$netns" setsid nohup python3 \ sudo -n ip netns exec "$netns" setsid nohup python3 \