[07:42:07] Starting switchover for x3 codfw shortly [08:55:15] federico3: can you remember if any hosts were swapped out of msX sections in the past few months? [08:55:58] cezmunsta: you mean deprovisioned? I'd have to look at the puppet git history [08:56:23] but there was a section split, wait [08:56:43] I think that some were decommissioned [08:56:57] T424028 [08:56:58] T424028: Decommission db2141-db2152 - https://phabricator.wikimedia.org/T424028 [08:57:52] That seems to have undone some of the Trixie upgrades for ms as codfw is showing as Bookworm: db[2251-2253].codfw.wmnet [08:57:59] also some hosts moved around, see git log hieradata/hosts/ [10:22:04] Hi folks, could I get a +1 to https://gerrit.wikimedia.org/r/c/operations/puppet/+/1311416 please? Drain the final 3 old-style storage nodes still on bullseye (once drained we can conver to new-style and reimage to trixie, and that will be easier) [10:22:40] looking [10:22:49] TY :) [10:24:42] Emperor: are the 3 hostnames listed somewhere that i can crossreference? [10:25:36] federico3: I mean, that file is where you can observe they are old-style storage; if you wanted to confirm that directly, you'd log onto the node and see things like /srv/swift-storage-sdl1 rather than /srv/swift-storage-objects0 [10:28:50] (sorry, I know that's a bit unsatisfactory) [10:52:41] ok, FWIW I checked /srv [10:56:36] thanks :) [11:16:29] federico3: what is the reason that a major-upgrade should fail with "Unsupported section kind" when depooling allows those to be chosen to continue? [11:26:31] major-upgrade is not implemented for all type of sections [11:26:58] It is executed against a single host though isn't it? [11:27:05] in general depooling is broader and meant for all SREs [11:27:33] I am not sure that explains things [11:28:50] If the cookbook only works against a single host then surely it should be usable without failing because a host is not in dbctl? [11:29:38] The code checks that too and only issues depool when not in dbctl [11:33:33] where did you get the Unsupported section kind from? [12:04:06] extract_section_kind_and_method gets called [12:04:29] via: imeta = fetch_host_instance_from_zarcillo(args.hostname) [14:45:48] FIRING: MysqlReplicationLagPtHeartbeat: MySQL instance db1152:9104 has too large replication lag (11m 8s) - https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting#Depooling_a_replica - https://grafana.wikimedia.org/d/000000273/mysql?orgId=1&refresh=1m&var-job=All&var-server=db1152&var-port=9104 - https://alerts.wikimedia.org/?q=alertname%3DMysqlReplicationLagPtHeartbeat [14:50:18] ^ relates to rack maintenance - no lag seen in orchestrator for ms1 [14:50:48] RESOLVED: MysqlReplicationLagPtHeartbeat: MySQL instance db1152:9104 has too large replication lag (12m 48s) - https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting#Depooling_a_replica - https://grafana.wikimedia.org/d/000000273/mysql?orgId=1&refresh=1m&var-job=All&var-server=db1152&var-port=9104 - https://alerts.wikimedia.org/?q=alertname%3DMysqlReplicationLagPtHeartbeat [23:01:56] FIRING: SystemdUnitFailed: wmf_auto_restart_prometheus-mysqld-exporter.service on db2209:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed