[07:45:55] good morning folks [07:46:03] I am going to reboot the bastions for kernel upgrades [07:46:10] starting with 1004, I don't see folks connected [07:46:14] cc: godog, Emperor [07:48:24] elukey: I guess that is why I have just been disconnected :D [07:49:08] cezmunsta: I didn't see you in last though, did you just connect or was it something pre-existent? [07:49:15] sorry! [07:49:25] pre-existing [07:52:24] No worries, I am back in now [07:52:31] 1004 is up, proceeding with 2003 now [07:52:59] elukey: yeah, I think your check wasn't entirely accurate (I was running a reimage on cumin1003 via bast1004) [07:53:06] [not a problem, it's in a tmux] [07:55:32] Emperor: no it wasn't, probably I should've checked also `who -a`, indeed I see more stuff for 2003 [07:55:55] Filling Moritz's shoes is hard :D [07:58:57] ok doing 5005 no [07:58:59] *now [08:00:13] I don't see anybody in 3007, so I'll do that afterwards [08:10:56] 6003 is next, 2003 is used so I'll retry either later on or tomorrow [08:24:12] done, remaining one is 2003 [09:31:53] Depooled maps in eqiad to allow reboots [09:47:15] ugh, why do so many of my codfw swift frontends want VLAN moves? 's making the reimage process much slower [10:23:24] they are probably on all the rows that got affected by the move [10:23:57] it is a sizeable chunk of the infra, so the moves are bound to happen sooner or later [10:23:58] yeah, I know, I'm just grumpy :) [10:24:12] oh yeah that is totally fine if you want to rant :D [10:29:26] a note to oncall (which is partly me) that I'm about the reimage the stats reporter and ring manager for the eqiad ms cluster; so there will be a short gap in metrics and so on. [11:01:52] repooled maps after the reboots [12:14:25] a note to oncall that I'm about the reimage the stats reporter and ring manager for the codfw ms cluster; so there will be a short gap in metrics and so on. [12:49:44] Emperor: ok thanks, had missed your message but nothing has fired [12:50:49] I'm not done yet ;) [12:50:56] haha ok :) [12:51:13] fwiw I am about to start draining traffic on cr2-eqiad ahead of JunOS upgrade / line card install [12:51:34] we'll be extra careful that the correct router is unplugged this time [12:58:32] :-) [14:58:43] heads-up that I just pushed a change to cumin aliases (1309288). Everything seems OK in my limited tests but please LMK (and feel free to roll back) if your aliases aren't working [22:01:05] Has anyone reimaged a host to bookworm in the last few days or so? I've 2 R440 hosts that aren't detecting their disks in the bookworm installer, no problems in trixie. Error is at https://phabricator.wikimedia.org/T430879#12126230 [22:08:05] inflatador: I haven't seen that, but these seems strange https://packages.debian.org/bookworm/linux-image-6.1.0-42-amd64 [22:09:31] jhathaway ACK, I'm trying to reimage a VM to bookworm just to see if I get different results [22:09:45] but of course I ran into a different issue ;P [22:09:47] `sudo gnt-instance console datahubsearch1001.eqiad.wmnet` [22:10:04] `Host key for ganeti01.svc.eqiad.wmnet has changed and you have requested strict checking.` [22:11:46] I can start a ticket for the Ganeti host key stuff if you like, I was running from `ganeti1048` [22:12:22] linux-image-6.1.0-50-amd64 is the latest for bookworm, I assume the other version was removed, since it is a security update [22:13:17] but I don't know where 6.1.0-42 is referenced offhand [22:14:12] my best guess is that it's somewhere in the install server PXE config? [22:16:10] which install server were you using inflatador? [22:17:41] this doc looks promising, https://wikitech.wikimedia.org/wiki/SRE/Infrastructure_Foundations/Debian-installer [22:17:45] trying that now... [22:17:53] jhathaway is it typically in the same DC as the host that's reimaging? If so, I'm currently reimaging datahubsearch1001 in eqiad. [22:18:09] yes it is [22:18:31] Nice find [22:19:00] Sounds like we'd need to run that command on `puppetserver1001`? [22:20:53] yeah, doing that now. [22:27:02] inflatador: okay, I believe I followed the instructions correctly, give it another go! [22:27:13] jhathaway excellent, thanks for your help! [22:27:36] I think this is one of those **many things** m.oritz takes care of :) [22:28:21] side note: one of these days we should sit down and talk https://fai-project.org/ , I've got it doing VM and baremetal reimages in my homelab [22:28:53] fai is great, though I haven't used it in many years, I failed to get my last shop to adopt it as well, but hope dies last! [22:29:00] so I would love to chat about it! [22:29:39] Nice, I'll try and scrape together a demo that will undoubtedly fail miserably, maybe for next week if that works for you? ;) [22:30:20] or really any time in the future, no rush [22:30:22] I love a demo, but don't spend too much time, I deployed it at a job around 2004, and it worked great [22:30:39] its speed is the nicest aspect [22:32:33] I just started learning it. I came from a Red Hat shop, and Anaconda is pretty slick compared to partman/deb installer. I started looking for alternatives and I guess it's still a thing [22:34:31] anyway, just restarted the reimage for https://fai-project.org/ [22:34:39] err...for datahubsearch1001 [22:36:40] great [22:38:31] also created T432300 for the ganeti console stuff, no rush on that one [22:38:31] T432300: Low priority: `sudo gnt-instance console` fails on EQIAD due to SSH host key changes - https://phabricator.wikimedia.org/T432300 [22:39:06] thanks [22:49:31] sorry need to go afk inflatador, hopefully it re-images, fingers-crossed [22:55:45] np jhathaway . I can confirm is it working now on `wcqs1002` [22:57:37] thanks again for the assist