[05:16:25] FIRING: SystemdUnitFailed: gitlab-package-puller.service on apt-staging2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [05:21:25] RESOLVED: SystemdUnitFailed: gitlab-package-puller.service on apt-staging2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [07:45:06] 10netops, 10Cloud-VPS, 06Data-Platform-SRE, 10Data-Services, and 4 others: Plan to make clouddumps more resilient and easier to operate - https://phabricator.wikimedia.org/T411248#12118080 (10fgiunchedi) I forgot to attach a task number, however: 07:44 -stashbot:#wikimedia-cloud- Logged the message at htt... [07:56:25] FIRING: [23x] SystemdUnitFailed: cfssl-ocsprefresh-Wikimedia_Internal_Root_CA.service on pki2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [08:01:25] FIRING: [23x] SystemdUnitFailed: cfssl-ocsprefresh-Wikimedia_Internal_Root_CA.service on pki2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [08:06:25] RESOLVED: [23x] SystemdUnitFailed: cfssl-ocsprefresh-Wikimedia_Internal_Root_CA.service on pki2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:00:36] 10netops, 10Cloud-Services, 06collaboration-services, 06Data-Persistence, and 7 others: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918#12118300 (10CWilliams-WMF) [09:05:17] 10netops, 10Cloud-VPS, 06Data-Platform-SRE, 10Data-Services, and 4 others: Plan to make clouddumps more resilient and easier to operate - https://phabricator.wikimedia.org/T411248#12118309 (10fgiunchedi) As of now `dumps_use_nfs_lb: true` is deployed to toolsbeta, unfortunately for k8s workers that means a... [09:05:28] 10netops, 10Cloud-Services, 06collaboration-services, 06Data-Persistence, and 7 others: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918#12118310 (10CWilliams-WMF) [11:08:20] 10netops, 10Cloud-Services, 06collaboration-services, 06Data-Persistence, and 7 others: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918#12118704 (10cmooney) [11:08:59] 10netops, 10Cloud-VPS, 06collaboration-services, 06Data-Persistence, and 8 others: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918#12118705 (10taavi) [11:44:51] dear foundations, there is a server marked as failed, and we are geting this diff (for an irrelevant change) [11:44:51] https://phabricator.wikimedia.org/T421711#12118737 [11:45:37] 10netops, 10Cloud-VPS, 06collaboration-services, 06Data-Persistence, and 8 others: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918#12118824 (10cmooney) [11:45:52] I am asking because I wouldn't expect the mgmt enrty to be removed from there [11:49:00] 10netops, 10Cloud-VPS, 06collaboration-services, 06Data-Persistence, and 8 others: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918#12118846 (10BTullis) I have fully drained dse-k8s-worker2003 and depooled cirrussearch2110. Good to go from our side. [12:07:25] FIRING: SystemdUnitFailed: dump_proxy_ranges.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [12:44:45] btullis, brouberol o/ - I'd need to reboot the krb hosts :) [12:45:36] It has been a while since I've done it, how did Moritz do it? In theory it should be sufficient to just reboot them, it shouldn't cause a big issue [12:46:50] ahhh see Luca from the past: https://gerrit.wikimedia.org/r/c/operations/puppet/+/542112 [12:47:05] always a smart person (present Luca a lot less) [12:47:14] do you prefer to do a full failover? [12:55:38] 10netops, 10Cloud-VPS, 06collaboration-services, 06Data-Persistence, and 8 others: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918#12119177 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=8b130fc6-8cf0-455b-b0d3-cfde4861384c) set by cmooney@cumin1003 for 2:00:00 o... [12:57:32] 10netops, 10Cloud-VPS, 06collaboration-services, 06Data-Persistence, and 8 others: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918#12119202 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=bdeaaae5-a950-451d-a6e0-365a469fd764) set by cmooney@cumin1003 for 2:00:00 o... [13:49:46] FIRING: [2x] ProbeDown: Service idp2005:443 has failed probes (http_idp_wikimedia_org_ip4) - https://wikitech.wikimedia.org/wiki/CAS-SSO#Alerting - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [13:51:49] 10netops, 10Cloud-VPS, 06collaboration-services, 06Data-Persistence, and 8 others: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918#12119568 (10cmooney) Upgrade is now complete, all looks healthy. Thanks everyone for their assistance. [13:54:41] RESOLVED: [4x] ProbeDown: Service idm2001:443 has failed probes (http_idm_wikimedia_org_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [15:12:21] 10netops, 06Infrastructure-Foundations, 10ops-drmrs, 06SRE: Move TATA IP Transit circuit 1243318 cross-connect from rack B12 to rack B13 - https://phabricator.wikimedia.org/T431605#12120098 (10cmooney) 05Open→03Declined As it turns out dc-ops report that the panel in B13 is also full. So we have o... [16:07:40] FIRING: SystemdUnitFailed: dump_proxy_ranges.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [16:42:25] RESOLVED: SystemdUnitFailed: dump_proxy_ranges.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed