[08:10:00] marostegui, cezmunsta I can start the rolling reboots in pc/ms/s* after rebooting test-s4 [08:10:25] federico3: I added a discussion point for the reboots in today's meeting, can you hold that? [08:10:33] ok [08:10:36] thanks [08:37:11] marostegui: are you testing anything with db2209 at the moment? [08:37:26] cezmunsta: not yet [08:37:42] Have you noticed the change in pattern in InnoDB checkpoints? [08:38:17] after the crash? [08:38:22] yes [08:38:32] no, I haven't touched the host since it crashed [08:40:41] cezmunsta: I've not checked the graphs, but I would expect changes after it stopped being a master and now it is a replica [08:46:09] Its peak/trough gap is about 1.2B .. its sibling is ~0.2B gap [08:48:49] cezmunsta: I haven't had time to check, but I do expect different behaviour since it stopped being a master and even more, it is not receiving any traffic other than the replication thread, so i do expect it to be different from the other replicas [08:49:35] At the moment I am waiting for dcops to check on their side, but I don't think graphs are very realistic at the moment as since the crash it didn't have any traffic [10:28:24] Amir1: I am going to push x4 codfw [10:28:46] cool cool [10:29:02] Amir1: can i do one at the time or both should go at the ssame time? eqiad and codfw I mean [10:29:39] marostegui: ps: remember to take some TOIL after you worked on the weekend [and I think ideally also fill out a post-shift form so its recorded] [10:30:01] Emperor: will do, thanks for the reminder. I will do the oncall post shift indeed yeah [10:30:21] 👍 [10:33:46] Amir1: pushed both sections [10:34:05] I missed the previous ping sorry [10:34:23] yeah, as long as we haven't set the virtual domain, everything shoudl be noop [13:26:51] "everything shoudl be noop" That aged like a fine milk [13:29:05] Emperor: gonna vacuum ms-be1070 now [13:30:14] federico3 cezmunsta https://phabricator.wikimedia.org/T430769 I've edited the task and added the 4 different cases (master, replica for RO and master,replica for RW) [13:30:19] Let me know if that's clearer [13:30:29] thanks [13:31:34] marostegui: yep, all clear ty [13:43:43] https://grafana.wikimedia.org/d/000000378/ladsgroup-test?orgId=1&from=now-1h&to=now&timezone=utc&viewPanel=panel-29 [13:44:21] https://usercontent.irccloud-cdn.com/file/GWxD9dH8/image.png [13:59:11] FIRING: SystemdUnitFailed: swift_dispersion_stats.service on ms-fe1009:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [14:03:27] that's swift-dispersion-report failing, though it runs fine now. [14:04:11] RESOLVED: SystemdUnitFailed: swift_dispersion_stats.service on ms-fe1009:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [14:52:06] o/ has something changed in swift access patterns over the last month or so that would cause an increase in requests (and logs of them, specifically?) [14:52:35] centrallog2002's syslog disk is full and it's mostly rotated swift logs [14:57:36] more directly, could we drop messages for AUTH_search-update-pipeline from being sent to centrallog? they're the vast majority of messages [14:59:04] not that I've noticed, but then I've not been staring at them very much [15:00:23] sorry, that drop is only for the thanos-be* hosts. ms-be hosts are unsurprisingly mostly thumbs requests which is normal enough. we'll take other measures anyway [15:01:32] I have been reimaging all the thanos-be hosts to trixie in the last 2 weeks. [15:04:39] I doubt it's that - might just be an us problem [15:21:55] Scraper maybe? [16:28:23] vacumming ms-be1065 now [17:12:51] next: ms-be1071 [23:01:56] FIRING: SystemdUnitFailed: wmf_auto_restart_prometheus-mysqld-exporter.service on db2209:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed