[03:04:12] FIRING: SystemdUnitFailed: wmf_auto_restart_prometheus-mysqld-exporter.service on db2209:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [12:57:43] I'm gonna test sre.mysql.multiinstance_reboot that I've just merged, doing a roll-reboot of clouddb* [12:57:57] * cezmunsta hides [12:58:01] :) [13:02:34] failed on the first one :D [13:02:49] What did it fail with? [13:06:29] "Can't connect to MySQL server on 'clouddb1013.eqiad.wmnet' (timed out)" [13:06:46] I'm looking at the stack trace [13:11:13] I cannot resist the fun of handling mysql servers, but this is the end of my work week, have a nice day! [13:11:48] have a good weekend! [13:12:21] cezmunsta: it's the new check for number of replicas, I should've tested in clouddbs :) [13:12:26] jynus: enjoy the final on Sunday, if you watch it that is :) [13:13:01] "A tope con Argentina", I guess? [13:13:58] dhinus: ah, well you found it now :D [13:14:02] jynus: +1 [13:17:20] cezmunsta: https://mende.io/img/import/2021-12-11-testing-in-production.jpg, I guess :D [13:22:40] :) [13:27:10] I think that instance.cursor() tries a TCP connection from cumin to the host, which fails for multi-instance hosts [13:27:19] I'm preparing a fix [13:39:46] https://gerrit.wikimedia.org/r/c/operations/cookbooks/+/1311857 [13:40:24] I'm testing it on a single clouddb