[06:06:25] FIRING: SystemdUnitFailed: dump_proxy_ranges.service on puppetserver2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [06:53:10] 10SRE-tools, 06Infrastructure-Foundations, 06SRE: Deploy wmflib 3.0.0 to production - https://phabricator.wikimedia.org/T430552#12097992 (10elukey) 05Open→03Resolved a:03elukey [06:53:23] pywmflib 3.1 rolled out everywhere [07:03:11] moritzm, XioNoX o/ should we do krb1002's reimage today/tomorrow? Before Moritz goes afk [07:03:40] it needs some coordination with Data Platform since they need to failover [07:04:22] I don't really have time for this, there's various other thing I need to wrap up before I leave [07:04:35] and there's no specific hurry, we can do it when I'm back? [07:04:46] okok yes no problem for me [07:34:17] XioNoX: I'm around (re rack maintenance) in case something is unclear or breaks [07:40:13] aux-k8s-etcd1* on bookworm [07:44:18] jayme: thanks! [08:15:23] 10netops, 06Data-Persistence, 06Infrastructure-Foundations, 06ServiceOps new, 06Traffic: codfw: rack B3 maintenance - https://phabricator.wikimedia.org/T430909#12098290 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=5e02ef3f-9732-4694-88a1-3e0f414cfe47) set by ayounsi@cumin1003 for 2... [08:31:25] FIRING: [2x] SystemdUnitFailed: fetch-rings-thanos.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [08:33:39] hello, I've been told you might have an opinion on that change suggesting to pull puppet changes via gerrit's discovery record: https://gerrit.wikimedia.org/r/c/operations/puppet/+/1296495 (https://phabricator.wikimedia.org/T420184) [08:35:21] jayme: all done [08:35:37] XioNoX: great, thanks [08:48:56] 10netops, 06Data-Persistence, 06Infrastructure-Foundations, 06ServiceOps new, 06Traffic: codfw: rack B3 maintenance - https://phabricator.wikimedia.org/T430909#12098486 (10ayounsi) 05Open→03Resolved alright, everything is done! [10:18:24] hi folks! I'm working on reimaging and re-iping several k8s nodes, and was wondering if it might be acceptable to run homer inline for our sre.k8s.renumber-node cookbook. At the moment, we prompt the person running the cookbook to run homer externally, and it would be nice to move that in-flow if it doesn't introduce too much risk. [10:21:17] bjensen: yep [10:21:32] XioNoX: right on, thanks :) [10:22:08] bjensen: https://github.com/wikimedia/operations-cookbooks/blob/master/cookbooks/sre/network/__init__.py#L500 [12:06:25] FIRING: [2x] SystemdUnitFailed: fetch-rings-thanos.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [12:11:48] Hi folks, anyone available to help with a system that won't UEFI HTTP Boot, please? [12:12:07] Watching it on console, it gets to the point of trying, declares there's no link available, and falls back to hdd booting [12:13:18] It looks to want to try and HTTP boot of Embedded NIC 1 Port 1 Partition 1, which I _assume_ is correct? The other options are Embeddec NIC 2 port 1 and Integrated Nic 1 Ports 1 and 2 [12:13:30] this is thanos-be1006 [12:16:03] Emperor: looking [12:17:11] TY <3 [12:18:52] Emperor: it was impacted by https://phabricator.wikimedia.org/T401966 but that should have fixed it [12:20:54] Emperor, looks like it's not fixed? https://www.irccloud.com/pastebin/dBnI43Ow/ [12:21:24] should be UEFI booting not PXE... [12:22:18] XioNoX: poking through the system setup UI, lists all 4 PXE devices as disabled under UEFI PXE Settings, and only HTTP Device1 enabled under UEFI HTTP Settings [12:23:12] but you know how user-friendly the system setup UI is... [12:23:41] Emperor: can you restart the process in a few ? [12:24:10] XioNoX: YM restart the reimage cookbook? Sure, LMK when to :) [12:26:35] yep, or just reboot the host if it's still configured to try network book [12:27:04] Emperor: I'm ready [12:27:43] but my guess is that it's trying to http boot on the wrong interface [12:27:43] XioNoX: OK, restarting cookbook [12:28:12] let me know once it's passed the DHCP step [12:28:22] and if you can screenshot the MAC it's trying to use [12:28:23] Host rebooted via Redfish [12:30:19] XioNoX: I'm afraid all it said was "Booting from HTTP Device 1: Embedded NIC 1 Port 1 Partition 1" before almost immediately saying something about no network connected and falling back to hdd [12:31:48] XioNoX: looking through the web iDRAC, that should be B4:E9:B8:BF:AC:49 [12:32:48] ok, so yeah it's trying the wrong nic, it should be 8c:84:74:21:53:b0 [12:33:14] Emperor: can you try to run the provision cookbook? see if it fixes it [12:33:36] that MAC is for Integrated NIC 1 Port 1 Partition 1 [12:34:00] XioNoX: YM sudo cookbook sre.hosts.provision --no-dhcp --no-users thanos-be1006 ? Will do [12:34:10] Emperor: something like that [12:34:49] I always got confused between the two but Embdeded is most liklely the 1G built-in NIC, and Integrated is the 10G add-on card [12:36:00] I'm not sure the cookbook is changing anything, but let's see (it's still running) [12:36:38] ah, no I do see 'Enabling UEFI HTTP boot on NIC NIC.Integrated.1-1-1' [12:36:55] great [12:36:58] Oh, no the cookbook has exploded [12:37:15] uh? [12:37:27] hang on, making a paste [12:38:10] XioNoX: https://phabricator.wikimedia.org/P94760 [12:38:18] cookbook is offering me retry, skip, abort [12:39:23] Emperor: try the retry, but that might need to be escalated to elukey [12:39:28] Ah, and the iDRAC web UI does show a scheduled job. [12:39:56] I've cleared the scheduled job and am retrying [12:43:16] [importsystemconfiguration job is now running, 20% done, host is rebooting] [12:52:02] that job ran to completion, but then the cookbook has failed 'Error: Unable to establish IPMI v2 / RMCP+ session' [12:52:45] going to try another reimage anyway, since it seems it can now talk redfish again... [12:53:09] I'm still getting `spicerack.redfish.RedfishError: Found more than 1 NIC with PXE enabled: ['NIC.Embedded.1-1-1', 'NIC.Integrated.1-1-1']` [12:53:24] so that might not be solved yet [12:53:33] if no luck I'll have a look at the BIOS screen [12:55:39] lemme know if I have to check the thanos hosts [12:57:17] it's about to try booting again, so we'll see if it gets the right NIC [12:58:45] Booting from HTTP Device 1: Integrated NIC 1 Port 1 Partition 1 [12:58:56] OK, looks to be at least trying to boot now [12:58:58] nice! [12:59:08] it got an IP from DHCP [12:59:25] XioNoX: you still getting the RedfishError about >1 NIC with this host? [13:00:14] yep [13:00:25] so maybe it's not relevant now with UEFI http [13:01:08] * Emperor has no idea... [13:02:03] Emperor: is it booting properly now? [13:02:59] XioNoX: currently in the Debian installer, so yes, thanks :) [13:03:17] awesome [13:05:07] hopefully the thanos-backends won't all take an hour of debugging to get to that point... [13:06:07] yeah... [13:09:39] XioNoX: thanks for all your help :) [13:09:49] np! [13:11:49] 10CAS-SSO, 06Infrastructure-Foundations: Migrate cloudinfra-idp-1 to Bookworm - https://phabricator.wikimedia.org/T379069#12099995 (10taavi) →14Duplicate dup:03T402006 [14:40:26] elukey: I've just had the exact same problem with thanos-be1007 :(, trying a re-run of the provision cookbook again [14:58:53] ack! [15:02:03] elukey: as with thanos-be1006, provision cookbook ends with Error: Unable to establish IPMI v2 / RMCP+ session (after skipping the admin users password change) when running ipmitool -I lanplus -H thanos-be1007.mgmt.eqiad.wmnet -U wm [15:02:03] froot -E chassis power status. But when I re-run the re-image cookbook, it does then boot into the installer [15:06:01] Emperor: interesting, lemme check [15:06:38] [I don't know what proportion of these nodes I'm going to have problems with, but it's 2/2 of the Dell backends I've tried so far], and it makes the reimage process take a lot longer (and then the post-install rebalance likewise) [15:11:09] okok I got what's happening, I think it is a little bug in the provision cookbook [15:11:47] the background story is that we are rolling out the wmfroot user on both dells and supermicros, that will become the standard user for all the BMC-related things from cookbooks/ipmi [15:12:42] on thanos-be we don't have that user set up yet, and you user --no-user when provisioning. So the cookbook must assume that wmfroot is present, and it fails [15:13:25] OK, it's a wrinkle not a biggie, but having to reprovision each host is more of a problem [15:17:53] that is probably a different issue, we have been having a stable provision/reimage for a while now but in the past some hosts had some config skew due to $reasons [15:18:12] reprovision should take some mins, why is it an issue? [15:18:18] I mean not ideal I know [15:18:53] Emperor: do you mind if I test provision on a thanos-be host? [15:19:01] maybe one that you are already working on it [15:21:20] jhathaway: o/ I have a couple of quick cookbook patches for you if you have time: https://gerrit.wikimedia.org/r/c/operations/cookbooks/+/1308668 and https://gerrit.wikimedia.org/r/c/operations/cookbooks/+/1308665 [15:21:32] yup [15:21:42] I also discovered other spicerack bugs when trying the new cookbook on other nodes [15:21:53] but hopefully nothing big [15:23:27] elukey: you can have thanos-be1007 once puppet is done setting it up. At least some of the time it does seem to be rebooting the host (and the problem is the extra time the whole "try reimage, observe it fails, try reprovision, then try reimage again" process takes (and that longer downtime means longer post-install backfilling before we can start on the next one) [15:23:41] elukey: I'll LYK once that's the case (first puppet run currently going) [15:26:24] 10netops, 06Data-Persistence, 06Infrastructure-Foundations, 06ServiceOps new, and 2 others: codfw: rack B3 maintenance - https://phabricator.wikimedia.org/T430909#12100808 (10Scott_French) The codfw etcd cluster is serving traffic as usual again. No issues encountered. I'll be refreshing our docs to re... [15:37:05] elukey: you can have at thanos-be1007 now; LMK when you're done, please? [15:38:54] ack doing it now [15:39:58] [time: thanos-be1005 (SM, no reprovision needed) was reimaged in 32 minutes; thanos-be1007 (Dell, needed reprovision) was reimaged in 65 minutes] [15:40:34] Emperor: done with 1007, the patch fixes the issue, merging it now [15:41:43] on the reimage timing, not really sure what could have been at fault, did you see where the cookbook spent the most of its time? [15:43:32] the process of downloading the config, modifying it, re-applying it (which seems to reboot the host at least once) I _think_ [15:43:50] 1007> ack, thanks. [15:54:42] 10netops, 06Infrastructure-Foundations, 06SRE: GSHUT (and other?) community matching/actions not working on SR-Linux - https://phabricator.wikimedia.org/T430810#12100957 (10cmooney) Found a few things testing this in container-lab: # We need to set "policy-result accept" in each statement or it will continu... [15:58:35] elukey: also, timing - that's the total time from starting the process to having the system re-imaged, so the larger total includes the first reimage cookbook running for long enough for it to be obvious that the host can't UEFI HTTP boot; then running the reprovision (which does take quite some time), then re-running the reimaging. [15:59:51] Emperor: okok got it, sadly this sometimes happens, we are trying to improve the reimage/provision experience as much as we could but the variety of vendor/server-types/firmwares combination is big :( [16:00:21] I appreciate it's not an easy problem! [16:06:40] FIRING: SystemdUnitFailed: fetch-rings-thanos.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [16:09:58] elukey: ...I was just trying to demonstrate why I cared about the need to run the provision cookbook :) [16:13:19] yes yes got it!! [16:16:34] sorry, I didn't mean to come across as grumpy [16:25:17] Emperor: I appreciated the thorough timing report :) [16:25:39] the more feedback we have the better for these tools [16:26:10] but I realized it is almost impossible to have something always consistent in BMC-land, so sad :( [16:26:27] firmware is a bit like the worst of both hardware and software :( [16:51:49] 10netops, 06Infrastructure-Foundations, 06SRE: GSHUT (and other?) community matching/actions not working on SR-Linux - https://phabricator.wikimedia.org/T430810#12101183 (10cmooney) Also discussed this on the SR Linux discord and got some good info. My observation in containerlab was that when we set the lo... [17:19:45] 10netops, 06Data-Persistence, 06Infrastructure-Foundations, 06ServiceOps new, 06Traffic: codfw: rack B3 maintenance - https://phabricator.wikimedia.org/T430909#12101299 (10Scott_French) Docs refresh: https://wikitech.wikimedia.org/wiki/Etcd/Main_cluster#Depool_etcd_client_traffic_from_a_cluster [17:27:39] jhathaway: hello! I got a PCC run failing with `WARNING: no nodes found for class: Class/Role::Jenkins` [17:27:39] I think it might be because facts from production needs to be synced up? [17:27:47] https://gerrit.wikimedia.org/r/c/operations/puppet/+/1308194/comments/812203a9_d61dc714 [17:28:28] no rush on that change, mutante has done a manual fix on the machine, that change is merely catching up with reality [17:32:07] thanks hashar, I'll take a look [17:45:36] jhathaway: thank you! :) [17:47:58] 10netops, 06Infrastructure-Foundations, 10ops-drmrs, 06SRE: Move TATA IP Transit circuit 1243318 cross-connect from rack B12 to rack B13 - https://phabricator.wikimedia.org/T431605 (10cmooney) 03NEW p:05Triage→03High [17:48:08] 10netops, 06Infrastructure-Foundations, 10ops-drmrs, 06SRE: Move TATA IP Transit circuit 1243318 cross-connect from rack B12 to rack B13 - https://phabricator.wikimedia.org/T431605#12101446 (10cmooney) [17:56:25] FIRING: [2x] SystemdUnitFailed: fetch-rings-thanos.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [17:56:34] 10netops, 06Infrastructure-Foundations, 10ops-drmrs, 06SRE: Move TATA IP Transit circuit 1243318 cross-connect from rack B12 to rack B13 - https://phabricator.wikimedia.org/T431605#12101471 (10cmooney) [18:24:39] 10netops, 06Infrastructure-Foundations, 10ops-drmrs, 06SRE: Move TATA IP Transit circuit 1243318 cross-connect from rack B12 to rack B13 - https://phabricator.wikimedia.org/T431605#12101591 (10cmooney) [21:56:40] FIRING: SystemdUnitFailed: fetch-rings-thanos.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed