[00:13:30] Oh sorry, I've only just seen this. Moritz is correct, this is about some ml-cache hosts that were decommissioned by k.lausman recently. It's perfectly safe for you to deploy to wikikube and aux, if you wish. [00:18:00] ah perfect thanks! I'm just winding down for the day but will do tomorrow without fear, appreciate it [00:22:07] Added a comment to the ticket, to that effect. Thanks for checking. [07:46:19] jynus: hello, did you depool ms-backup2004 for https://phabricator.wikimedia.org/T430909 ? [08:01:23] starting rack B3 depools [08:03:23] FAO oncallers - I'm about to start re-imaging thanos-fe1004, which is the stats reporter host for the thanos-swift cluster; so metrics from that cluster won't update for a little while. [08:06:44] Emperor: just fyi, there is thanos-fe2006 in rack B3 that will go down soon [08:08:52] OK [08:19:00] alright, rebooting codfw B3 switch [08:35:32] all done! [08:35:44] I'll slowly repool the services [08:35:48] XioNoX: yay \o/ [08:49:41] nice one :) [08:51:30] I guess too late now [08:57:42] jynus: let me know if that caused an issue with ms-backup2004 [08:57:55] absolutelly not [08:58:01] great :) [08:58:04] just some errors which are not fatal [08:58:19] I was relaxed preciselly because of that, as it is the client in this case, not the server [08:58:31] in which also are non faltal errors, but more annoying [08:58:39] cool, it's nice when services are resilient [08:58:47] I just didn't check the time properly, tottally my fault [08:59:07] I thought it was the usual time so was taking my time [09:02:06] I moved it to the morning while I'm oncall [09:02:12] another one tomorrow morning [09:02:43] yeah, I think it was the only one I should have depooled, no others from me on the followup [09:02:54] I am rechecking now [09:03:52] just to be clear, anything important is redundant, ms-backup2003 kept working in case a recovery was needed, so no bigie [10:22:48] XioNoX: heads up both our 100G Transports from eqiad<->codfw are down [10:23:00] uh? [10:25:39] don't see anything related to them in the maintenance calendar, I'll hold off on any draining of cr2 for now [10:26:50] traffic currently routing via eqord [10:27:07] https://www.irccloud.com/pastebin/QtutCcB3/ [10:27:10] and nothing in the email notif [10:28:54] writing an email to arelion support [10:29:01] they broke at different times at least :) [10:29:13] is a ticket via their portal not better? I can look at Lumen one [10:29:45] Arelion is quite reactive, but I can look at their portal [10:30:33] no probs either way [10:30:40] this is the log / time of it going down: [10:30:41] Jul 8 06:52:08 re0.cr1-eqiad mib2d[41416]: SNMP_TRAP_LINK_DOWN: ifIndex 910, ifAdminStatus up(1), ifOperStatus down(2), ifName et-1/1/2 [10:33:50] ticket opened [10:38:13] maybe we should look at adding an HE circuit between Dallas and Ashburn before we decom eqord [10:40:22] "We have identified several network elements in Dallas that are currently unreachable, limiting our ability to complete a full diagnosis of the issue." [10:40:41] aka. Network is too down to troubleshot the network. [10:40:59] "A request has been submitted for a field technician to attend the site and perform further troubleshooting. At this time, no estimated time of arrival (ETA) is available." [10:46:42] hmm ok [10:47:06] Lumen's automatic check suggests the problem with them is in Ashburn (and with us, but I presume they just code the site to always blame the customer) [10:48:05] this does rather focus the mind though, agree it might help to add another path over HE [10:49:24] if not our best option would be eqiad <-> ulsfo <-> codfw, however default cost of 500 everywhere would make esams, drmrs etc look as good [14:19:04] swfrench-wmf: hi. happy to do the dance again shortly, including the Liberica bits [14:21:17] sukhe: great, thank you very much! I need to check on a couple of things one last time, but otherwise ready to go. any preference on deferring until after the staff meeting vs. not? [14:21:29] none from my end at least [14:21:55] I have another meeting right after that one though but we can do a few things async [14:25:46] sukhe: sounds good. if you have no concerns about getting started now, I'm happy to do so. everything looks good AFAICT and it would be nice to get this traffic off the WAN links :) [14:26:00] swfrench-wmf: now is better for sure for me yep :) [14:26:16] so I can get started on the pybals and then you can take care of the DNS one like yesterday? [14:26:21] will wait for your +1 [14:26:21] SGTM [14:26:24] ok [14:29:16] oh that's a curious bug in the way we insert the change-id into a revert commit message [14:30:16] :] [14:32:25] (i.e., when I created the revert, I had Hosts then Bug, expecting to have Change-Id appended after. somehow it permuted them as Change-Id, Hosts, Bug, which is puzzling) [14:42:11] it is weird indeed and unexpected [14:46:21] sukhe: TIL that _etcd-client-ssl._tcp.wikimedia.org is a thing! I was super confused for a second why certain confds in ulsfo were sticking to eqiad, until I noticed the SRV records they were using [14:47:10] not a problem here, but I'll make sure we add that to our procedure in the event we need to do the inverse of this operation (i.e., depool from eqiad) [14:47:15] swfrench-wmf: context is the CR yesterday? [14:47:53] yeah, so our docs only cover touching the wmnet zone template. no mention of wikimedia.org [14:48:09] totally fine here since we were shunting everything to eqiad anyway :) [14:48:35] but ... if we were doing it in the opposite direction, we'd need to make sure those are included [14:48:42] swfrench-wmf: I only thought of that because I was trying to see how this is consumed in MW for the upcoming urldownloader work [14:48:45] https://gerrit.wikimedia.org/g/operations/mediawiki-config/+/ca53835aac2245325e5594bba9348943be35754e/wmf-config/ProductionServices.php#177 [14:49:09] * swfrench-wmf thumbs up [14:49:18] swfrench-wmf: pybal's done, going to look at Liberica now [14:49:23] all good to go from your end I presume? [14:49:38] all good, yes - I'll proceed with the rest of the confd restarts now [14:49:44] thanks [14:50:06] I have an example command in the task if you wanna copy / paste for the liberica restarts [14:51:02] thanks for handling the liberica bits :) [14:52:01] swfrench-wmf: yeah please feel free to share, will save me looking it up [14:53:01] `sudo cookbook sre.loadbalancer.upgrade -t T430909 --seamless --alias liberica-ulsfo --reason 'Clear control-plane connections to etcd' restart` <- is what I used yesterday in ulsfo [14:53:02] T430909: codfw: rack B3 maintenance - https://phabricator.wikimedia.org/T430909 [14:53:30] thanks, so I will do that for ulsfo and eqsin [14:53:38] * swfrench-wmf thumbs up [14:58:11] pybal's look good moving on to Liberica [15:17:11] swfrench-wmf: all done on my end as well [15:17:24] alright, looks like we're all done, then. thanks for your help, sukhe! [15:17:31] things look good, please let me know if not [15:17:34] <3 [15:17:48] * swfrench-wmf now turns to improving the docs on this [15:27:09] thanks for doing that! [19:09:18] JennH: okay to puppet-merge yours? [19:09:37] yes [19:09:42] doing, thanks! [19:09:48] tyt!\