[00:14:27] FIRING: SystemdUnitFailed: swift_rclone_sync.service on ms-be1069:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [04:14:27] FIRING: SystemdUnitFailed: swift_rclone_sync.service on ms-be1069:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [06:25:35] cezmunsta: The timeout question was regarding those from icinga? [06:48:56] I've switched m5 master [06:50:13] dhinus: there are some wmcs tools using that master, so fyi [06:51:37] marostegui: ack thanks! [07:07:30] marostegui: yes, i.e. waiting to go green in a cookbook run [07:12:17] yeah, I have not seen those timeouts recently [07:26:17] Maybe I have just been unlucky :D [08:14:27] FIRING: SystemdUnitFailed: swift_rclone_sync.service on ms-be1069:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [08:31:15] [I'll get to that in a bit] [09:10:21] https://grafana.wikimedia.org/d/75a174f3-44b6-4416-a8b8-201ad5a0c09f/swift-krinkle-copy?orgId=1&from=now-7d&to=now&timezone=utc&var-site=$__all&viewPanel=panel-35 [09:10:22] Emperor fyi, I'm deleting VP8 transcodes. Some decent drop in transcode containers is expected (still ongoing). Nothing to be alarmed about [09:10:25] https://usercontent.irccloud-cdn.com/file/dakDKTxE/grafik.png [09:11:11] then I add MPEG4 Part II, then drop MJPEG. It'll take a while [09:12:02] thanks for the update :) [10:46:32] rclone> lost a single delete race https://commons.wikimedia.org/wiki/File:Marie_claire_2019_pensive_cropped.jpg [10:48:58] RESOLVED: SystemdUnitFailed: swift_rclone_sync.service on ms-be1069:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [11:01:27] federico3: have you had a chance to look at https://gerrit.wikimedia.org/r/c/operations/cookbooks/+/1315796 yet? [11:07:26] no not yet... [12:34:37] OK. Do you have an idea of when you might have a chance? I left it intentionally so that the test failed for now, but I can undo that and it could be merged without affecting any other cookbooks. I see merge conflicts already appearing and it would be better to avoid as many of those as possible, perhaps even trying to use the approach for the new cookbook for replication(?) [12:35:17] i.e. I could put the changes to __init__ in a new MR on their own [12:36:53] marostegui: short term hack for the rolling restarts: runing the cookbook up to 3 times before the exit, already 1 ticket avoided :D [12:37:07] hahaha [12:37:18] Tomorrow we can just remove the check once the absent has rolled out everywhere [12:40:34] Emperor: o/ if you have time for a quick review https://gerrit.wikimedia.org/r/c/operations/puppet/+/1318701 [12:41:01] sretest2010 seems almost ready, I'll need to reimage again to properly install the OS on ssds and clean the hdds (the current reimage targeted those for the os) [12:41:14] after that, we'll be able to test the new config j specs [12:49:05] elukey: I think that change isn't quite right - I've commented, and linked to https://gerrit.wikimedia.org/r/c/operations/puppet/+/1185973 where I last set this up [12:50:05] Emperor: my aim is just to reimage without swift-related settings, with the standard recipes etc.. do you think that it won't work? [12:50:14] I can add the ms-be specific bits as well [12:51:57] but probably this kind of node doesn't really work with the standard recipe anyway [12:55:16] Emperor: updated the code change :) [13:00:19] Ta, have meeting now, will look after [14:53:26] federico3: I see quite a few "No zarcillo data" when running the security_reboots_updater script; 87 matches in the log to be precise. That seems a little concerning, is something not working properly or is this expected? [14:54:21] it's probably hosts that are not tracked [14:55:04] 87 of them seems high though, given the total number of DB hosts, doesn't it? [14:55:26] I can give you the log if you want to check anything in particular [14:55:47] some tasks also had a lot of other hosts, just sample which hosts are not found [14:59:50] Added as a comment on the MR [15:01:15] odd, it worked before, so it must be something else that broke [15:07:10] No errors, warnings all-but-2 for the "no data" and the 2 for "No kernel version in Prometheus". The only other notable part from what I see is that it ignored the entries where I had previously marked the instances as reimaged and left the previous release name there with strikethrough [15:09:13] yet I see https://phabricator.wikimedia.org/transactions/detail/PHID-XACT-TASK-lpwapzxzzkhrrr5/ - was this change done by the script? [15:09:56] "2026-07-28 14:44:12,544 [INFO] Task description updated" [15:10:29] uh? [15:10:45] That is what your script does [15:11:19] s/does/emits [15:12:21] i.e. check that the timestamp matches what you are looking at [15:12:42] so it's working as expected: some type of hosts are not supported e.g. dbproxy, test, hosts that have been deprovisioned or not deployed yet [15:13:43] Why are they not supported? They are in security tickets [15:14:40] Also, picking the first one: db1150 is in s3/s4 [15:14:54] because the script is in a very early stage and best-effort in trying to parse the task contents as we don't have a structured task format [15:17:01] I am still not following.. this comes from Zarcillo, right? Also, another is in s1 db2250 [15:18:36] cezmunsta: some hosts may be backup sources like I think db1150 (I'm on my phone but it sounds familiar) and we leave those to be done by jynus at his convenience [15:20:39] restarts? [15:20:55] Yes [15:21:12] I did those weeks ago and mentioned that on a meeting [15:22:19] jynus: this is about a script that reads the ticket and updates it. If the version was OK then the script that does the restarts would have skipped it [15:22:46] I can ensure the kernel version is ok [15:23:20] marostegui: certainly some are backup hosts, perhaps the log message is at fault, but I would expect the hosts to still be in Zarcillo else <> source of truth [15:23:55] didn't you deleted hosts from zarcillo a few weeks ago? [15:26:25] Definitely not backup hosts [15:26:43] cezmunsta: sorry yeah I thought you meant you were handling their restarts [15:27:34] I definitelly haven't touched zarcillo myself, but I remember someone from your team proposing to delete host entries [15:28:09] I can get you backups from the db from any time 3 months ago, though [15:34:32] I can see the backup sources on zarcillo: select * FROM instances JOIN servers on fqdn=server where `group` = 'dbstore'; [15:36:40] and I see them on grafana, so I think zarcillo && prometheus is sort of correct [15:37:01] issue must be somewhere else on the logic [15:37:53] for my automation, I personally use more puppet && cumin, as it tends to be more reliable [15:38:50] 41 seem to be in the instances table: select count(1) from instances where replace(replace(server, ".eqiad.wmnet", ""), "codfw.wmnet", "") in (...) [15:38:59] one for the morning [15:40:15] I saw 39 dbstores, but ofc backup sources are not all dbstores, that includes analytics and other hosts [15:40:40] instances, I mean [15:41:52] everything that I own I restarted, as I said here: https://phabricator.wikimedia.org/T431660#12108823 and marked on the task [15:44:10] https://phabricator.wikimedia.org/P95414 [15:46:21] except, of course https://phabricator.wikimedia.org/T431115 which is unoperable [15:50:40] sorry, I got pinged mid conversation, but if you need help with anything, just ask (tomorrow) [15:51:00] I think the context is wanting to automate all dbs, which I am ofc in favour, not only for backup sources, but for all dbs [15:51:13] so let me know if I can help in any way [15:52:18] same as ideally all hosts eventually handling its pool state on dbctl or at least etcd, all in favour of that [16:14:03] Emperor: sretest2010 ready for a test! [16:42:36] elukey: thanks, will have a v. quick look now, hopefully more tomorrow [16:51:57] elukey: I've fixed the sad fs, so puppet's a bit happier (and we'll stop getting email), looks reasonable at a first glance.