[06:38:15] 10netops, 10Cloud-VPS, 06Data-Platform-SRE, 10Data-Services, and 4 others: Plan to make clouddumps more resilient and easier to operate - https://phabricator.wikimedia.org/T411248#12122473 (10fgiunchedi) A side effect of the toolsbeta deployment from yesterday was {T432115} i.e. pods getting stuck while st... [07:26:43] elukey: sorry I was ooo yesterday. I think we do need to perform a failover yes [07:28:32] ah, I just say ben's message on slack. nvm [07:29:16] brouberol: o/ [07:29:39] as you prefer! Probably a little downtime shouldn't be a big deal, but if it causes issues we can failover [07:29:50] nah I think a reboot should be fine [10:50:25] FIRING: SystemdUnitFailed: fetch-rings-eqiad.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [10:55:24] ^-- that's because I'm reimaging the ring manager ATM [11:05:40] 10netops, 06Infrastructure-Foundations, 06SRE, 13Patch-For-Review: Don't announce OSPF routes in unicast BGP on Nokia SR-Linux - https://phabricator.wikimedia.org/T423430#12123447 (10cmooney) Ok all merged and looking good. We now have no hidden routes in the table on the codfw spines ` cmooney@ssw1-e... [11:05:44] 10netops, 06Infrastructure-Foundations, 06SRE, 13Patch-For-Review: Don't announce OSPF routes in unicast BGP on Nokia SR-Linux - https://phabricator.wikimedia.org/T423430#12123450 (10cmooney) 05Open→03Resolved [11:10:25] RESOLVED: SystemdUnitFailed: fetch-rings-eqiad.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [11:10:45] 10netops, 06Infrastructure-Foundations, 06SRE, 13Patch-For-Review: Don't announce OSPF routes in unicast BGP on Nokia SR-Linux - https://phabricator.wikimedia.org/T423430#12123477 (10cmooney) 05Resolved→03Open [11:12:25] FIRING: SystemdUnitFailed: fetch-rings-eqiad.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [11:13:59] ^-- should be fixed now [11:17:25] RESOLVED: SystemdUnitFailed: fetch-rings-eqiad.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [12:32:25] FIRING: SystemdUnitFailed: fetch-rings-codfw.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [12:37:12] ^-- that's because I'm reimaging the ring manager ATM [13:02:09] 10netops, 10Cloud-VPS, 06collaboration-services, 06Data-Persistence, and 8 others: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918#12124086 (10cmooney) 05Open→03Resolved [13:12:25] RESOLVED: SystemdUnitFailed: fetch-rings-codfw.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:17:54] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-eqiad, 06SRE: Install new MPC10E-10C line cards on cr1-eqiad and cr2-eqiad slot 0. - https://phabricator.wikimedia.org/T426343#12124189 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=121d96b4-7ab1-4cc6-9b24-75a50e15a132) set by... [13:19:32] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-eqiad, 06SRE: Install new MPC10E-10C line cards on cr1-eqiad and cr2-eqiad slot 0. - https://phabricator.wikimedia.org/T426343#12124196 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=84179922-a401-44a7-95de-38b063d38fb1) set by... [14:00:19] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-eqiad, 06SRE: Install new MPC10E-10C line cards on cr1-eqiad and cr2-eqiad slot 0. - https://phabricator.wikimedia.org/T426343#12124411 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=7bebb0cd-fd97-4261-8062-22ac9765f376) set by... [14:57:52] 10netops, 06Infrastructure-Foundations: eqiad: upgrade routers (2026) - https://phabricator.wikimedia.org/T417873#12124790 (10cmooney) [18:23:37] 10netops, 06Infrastructure-Foundations, 06SRE: GSHUT (and other?) community matching/actions not working on SR-Linux - https://phabricator.wikimedia.org/T430810#12125726 (10cmooney) Going to close this one, with 'graceful-shutdown sender' configured on cr1-eqiad I can see that the community is matches on ssw... [18:23:49] 10netops, 06Infrastructure-Foundations, 06SRE: GSHUT (and other?) community matching/actions not working on SR-Linux - https://phabricator.wikimedia.org/T430810#12125727 (10cmooney) 05Open→03Resolved a:03cmooney [18:27:42] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-eqiad, 06SRE: Install new MPC10E-10C line cards on cr1-eqiad and cr2-eqiad slot 0. - https://phabricator.wikimedia.org/T426343#12125735 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=b3a498bf-e4dc-461c-9cbe-4c3255e0019c) set by... [18:29:26] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-eqiad, 06SRE: Install new MPC10E-10C line cards on cr1-eqiad and cr2-eqiad slot 0. - https://phabricator.wikimedia.org/T426343#12125749 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=944ebf34-ac03-4b88-9838-2666b665680e) set by... [19:21:39] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-eqiad, 06SRE: Install new MPC10E-10C line cards on cr1-eqiad and cr2-eqiad slot 0. - https://phabricator.wikimedia.org/T426343#12125906 (10cmooney) We have successfully installed the new switch-control boards and and MPC10E in cr2-eqiad. All is wor...