[03:04:11] FIRING: SystemdUnitFailed: wmf_auto_restart_prometheus-mysqld-exporter.service on db2209:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [05:34:48] FIRING: MysqlReplicationLag: MySQL instance db2209:9104@s3 has too large replication lag (12h 2m 2s). Its replication source is db2205.codfw.wmnet. - https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting#Depooling_a_replica - https://grafana.wikimedia.org/d/000000273/mysql?orgId=1&refresh=1m&var-job=All&var-server=db2209&var-port=9104 - https://alerts.wikimedia.org/?q=alertname%3DMysqlReplicationLag [05:34:48] FIRING: [2x] MysqlReplicationLagPtHeartbeat: MySQL instance db2209:9104 has too large replication lag (12h 2m 2s) - https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting#Depooling_a_replica - https://grafana.wikimedia.org/d/000000273/mysql?orgId=1&refresh=1m&var-job=All&var-server=db2209&var-port=9104 - https://alerts.wikimedia.org/?q=alertname%3DMysqlReplicationLagPtHeartbeat [05:41:56] RESOLVED: SystemdUnitFailed: wmf_auto_restart_prometheus-mysqld-exporter.service on db2209:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [06:14:48] RESOLVED: MysqlReplicationLag: MySQL instance db2209:9104@s3 has too large replication lag (15m 38s). Its replication source is db2205.codfw.wmnet. - https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting#Depooling_a_replica - https://grafana.wikimedia.org/d/000000273/mysql?orgId=1&refresh=1m&var-job=All&var-server=db2209&var-port=9104 - https://alerts.wikimedia.org/?q=alertname%3DMysqlReplicationLag [06:14:48] RESOLVED: [2x] MysqlReplicationLagPtHeartbeat: MySQL instance db2209:9104 has too large replication lag (15m 37s) - https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting#Depooling_a_replica - https://grafana.wikimedia.org/d/000000273/mysql?orgId=1&refresh=1m&var-job=All&var-server=db2209&var-port=9104 - https://alerts.wikimedia.org/?q=alertname%3DMysqlReplicationLagPtHeartbeat [08:30:58] federico3: https://gerrit.wikimedia.org/r/c/operations/cookbooks/+/1277076/comments/9e281f43_98c772f2 [08:34:05] marostegui: ? [08:36:51] federico3: what is unclear? [08:38:28] AFAICT it's pending review, or is there something else we want to implement before merge? [08:40:44] federico3: Could you coordinate with cezmunsta to see if it can be reviewed/approved/discussed this week so we can ship it and close the incident follow up task? [08:41:46] starting s7 codfw switchover shortly ahead of rack B5 maintenance [08:43:39] ok [08:44:35] cezmunsta: can I merge https://gerrit.wikimedia.org/r/c/operations/cookbooks/+/1277076 ? I can put you as reviewer on the CR instead of CC for better clarity if you prefer [08:54:08] If the other reviewers' comments were all addressed then fine with me [09:00:21] ok, if anyone wants to +1 it I can merge it [09:10:54] federico3: has this been documented? [09:11:00] I think that may be the only thing pending [09:14:38] maybe it could be an opportunity to update https://wikitech.wikimedia.org/wiki/SRE/Data_Persistence/Databases/Runbooks adding a dedicated section? [09:16:27] go for it yeah [09:16:31] or maybe https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting [09:17:23] the latter is linked as "Incident Response" from https://wikitech.wikimedia.org/wiki/SRE/Data_Persistence/Documentation so it could be a better candidate [09:18:59] maybe embrace the power of "and"? :) [also: any relevant p.aging alerts?] [09:20:55] It would be quite rare that'd need to use this cookbook but in any case probably https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting but also liking from https://wikitech.wikimedia.org/wiki/SRE/Data_Persistence/Databases/Runbooks [09:49:38] ah, probably it should be here or link from here actually https://wikitech.wikimedia.org/wiki/Dbctl#Setting_all_sections_read-only [09:57:41] federico3: whatever you think works best, but also make a link from the link we use on paging [10:40:56] dhinus: any clue why we have this alias: db-clouddb-sanitization: P{an-redacteddb1001.eqiad.wmnet or clouddb1016.eqiad.wmnet or clouddb1020.eqiad.wmnet} [10:41:00] and just those hosts? [10:41:21] Ah I think I know, that's used by the sanitization cookbook [10:41:22] marostegui: I think I've seen it before [10:41:43] dhinus: I will add there 1029 as that's replacing 1016 "soon" [10:42:27] https://gerrit.wikimedia.org/r/c/operations/puppet/+/1090859 [10:43:02] dhinus: yeah I think it for the cookbook [10:43:22] dhinus: https://gerrit.wikimedia.org/r/c/operations/puppet/+/1310542 [10:44:37] but why only those ones? [10:45:48] because they belong to s5, which is the default wiki for new wikis so the only ones that get the "sanitization" by the cookbook [10:45:58] s/default wiki/default section [10:47:00] ah I see, can you add a note while you're changing the alias? I didn't immediately think that only s5 needs to be checked [10:48:27] another way would be to connect to all clouddb*, then in the cookbook check if it contains the s5 section, otherwise skip [10:49:05] dhinus: I will add a note, give me a sec [10:49:19] thx [10:50:13] dhinus: sent a new patch [10:51:50] marostegui: +1d [10:51:54] thank you [11:29:14] marostegui: is there any reason that es1046 is still running bookworm? [11:29:56] cezmunsta: nope! Other than it slipped through the cracks [11:30:04] I'll reimage it [11:30:13] Cool, thanks! [11:31:02] Thanks for flagging it [11:38:18] federico3: does the rolling restart script need manual checks for rack maintenance? [11:40:03] e.g. I don't see locks and it doesn't use the depool cookbook [12:04:18] cezmunsta: the rolling restart "core" was developed way before cookbooks and lock mechanism afaik and discovers the replication topology directly from the hosts, depools them using dbctl and is not aware of other activities [12:05:55] dhinus: you okay with this https://gerrit.wikimedia.org/r/c/operations/puppet/+/1310547 ? [12:09:45] federico3: I will take that as a, "Yes, care should be taken". Any reason that it has not been made aware? [12:11:17] cezmunsta: there is T423841 for adding locking to it but we haven't got around to do it yet [12:11:18] T423841: Add section-level locking to scripts and cookbooks - https://phabricator.wikimedia.org/T423841 [12:11:37] federico3: rolling_restart_pc_ms.py uses the cookbook though, hence the question [12:12:31] rolling_restart_pc_ms is much much newer (see the git history) and it's using pool,depool and also spicerack abstractions [12:24:30] The one for es has spicerack in too though and uses dbctl instead of the cookbook [12:31:42] cezmunsta: es1046 reimaged [12:32:29] *\o/* [13:27:44] I'm gonna optimize ruwikinews tables on the dbstore machine so I can answer a volunteer's question [13:29:55] FYI: if you forgot how to check RAID rebuilding progress on command line with perccli: https://phabricator.wikimedia.org/T431367#12119381 [13:31:06] I will add it to wikitech [14:32:47] any maint happening on s3? if not, I'm going to run a optimize table to save up some space [14:38:31] Amir1: I don't think that there is - noting that it had issues at the weekend though [14:40:19] thanks [15:31:16] "But since MariaDB 11.2, ALTER TABLE can do most operations with ALGORITHM=COPY, LOCK=NONE;" https://mariadb.org/mariadb-hidden-gem-online-schema-change-without-pt-osc/ [16:16:56] FIRING: SystemdUnitFailed: swift_ring_manager.service on ms-fe2009:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [16:19:13] ^-- (I think just upset by a roll-restart) [16:21:56] RESOLVED: SystemdUnitFailed: swift_ring_manager.service on ms-fe2009:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [17:31:27] inflatador: hey, might you be familiar with the wcqs nodes? [17:31:52] wcqs2001 is in codfw rack B6 where we are doing switch maintenance Thursday [17:32:08] I was wondering do we / how to depool before the works [17:50:42] topranks ACK, sorry for not responding earlier. It just needs a manual depool/repool. I'll get a patch up to add that hieradata with the instructions [17:51:37] inflatador: super thank you [17:51:53] happy to do as much as we can ourselves if there are instructions :)