[08:40:34] inflatador: I think generally thanos-swift doesn't have quotas set, but if you run swift stat with the relevant account's credentials, it'll tell you [09:27:58] hnowlan: ok to merge yours? [09:34:08] hnowlan: ping? [09:39:39] seems harmless enough that I've just merged to get my patches unstuck [09:43:43] taavi: my bad, thanks [10:06:41] on-callers: I'm about to start re-imaging thanos-swift (frontend) nodes to trixie (per T429630). I'll go slowly, but if anything goes bang LMK [10:06:42] T429630: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630 [10:15:10] merging a change to SMART monitoring - it's well tested so it Should Be Fine but just, yknow [10:15:27] (there is a chance we'll also suddenly learn about a bunch of failing BBU batteries) [11:16:33] I've re-imaged (to trixie) thanos-fe2004 and repooled it, and it's showing up in confctl as pooled, weight 100, but there's still no traffic going to it except the health-checks (which look to be OK). Do I need to be more patient, or has something gone wrong? [11:17:46] [if it makes a difference, the node was also re-numbered by the move-vlan cookbook] [11:34:42] Ah, https://wikitech.wikimedia.org/wiki/Conftool#Server_changing_IP_address [12:08:53] yeah :D [12:41:06] FYI,I'm disabling Puppet merges for about 10 mins shortly to reboot two Puppet servers [12:41:24] unless anyone tells me that currently is not a good time, ofc [12:42:51] disabling now [12:48:22] sorry didn't see it in time, I merged but I can wait [12:50:50] first server reboot, second starting now (took a little longer in BIOS etc) [12:51:54] I'll leave a note when puppet merges are back [12:55:43] all done [12:55:48] elukey: you can merge your patch now [12:56:04] dankeee [14:36:40] swfrench-wmf: hi. thanks for detailed instructions in https://phabricator.wikimedia.org/T430909#12091874 [14:36:56] outside of coordination with Traffic, do you need us to drive some part of this? [14:37:02] * swfrench-wmf stops typing in -traffic [14:37:05] hehe [14:37:13] haha. there was stuff happening there so I thought it's best here [14:40:51] thanks for checking! so: [14:40:51] * I'm happy to handle the pybal restarts, as long as it's at a time where I can grab Traffic's attention in case something goes wrong. [14:40:51] * I'm slightly unsure of how to proceed with the iberica-cp restarts after the SRV records are switched. it would be great if Traffic could confirm the procedure and / or drive that aspect. [14:41:00] *liberica-cp [14:41:35] swfrench-wmf: sre.loadbalancer.admin cookbook does it all [14:41:44] I think Traffic should handle both parts because you should focus on other stuff [14:41:56] and also because it's technically for us to do it and it's unfair to put that on you [14:42:10] we are planning this on Wednesday? [14:43:27] cdanis: thanks! yes, I vaguely recall this from when we were working on the PKI migration. [14:43:27] sukhe: the work needs to complete during our (Americas TZs) day today (Tuesday) in advance of the maintenance happening at 8:00 UTC Wednesday [14:43:41] ah fair [14:43:52] I was thinking we would do this on the day of the maintenance itself but yeah [14:44:14] swfrench-wmf: yeah today works. we have the Traffic meeting soon, so perhaps 09:30 PT? [14:45:02] Emperor thanks, the user is 'search:platform'. I'm guessing it's OK as we were able to snapshot the OpenSearch instance [14:47:20] sukhe: would starting about an hour later at 17:30 UTC (10:30 PT) work? I have a couple of conflicts between now and then. if someone from Traffic is available sooner than that, then review on the two patches I've posted to the task would be swell. [14:47:31] also fine yeah [14:47:37] * https://gerrit.wikimedia.org/r/c/operations/dns/+/1308114 [14:47:37] * https://gerrit.wikimedia.org/r/c/operations/puppet/+/1308115 [14:47:49] I reviewed the first one, looking at the other now [14:48:51] oh, awesome - thank you :) [15:00:22] I need to disappear for an appointment, but should be back (albeit in meetings) by 16:00 UTC [16:11:05] I'm closing out some old "Core Platform Team" and "API Platform" team tasks (both re-orged without direct successor in 2023), mainly non-software tasks about process, plans, etc not bugs or feature requests in currently deployed software. I stumbled on this one which I'm not sure is still relevant: "Generic strategy to deal with high volume / expensive traffic from cloud providers" by gehel [16:11:06] https://phabricator.wikimedia.org/T326782 [16:12:23] Can someone look at whether this has been superseded by other work, or if not, what team should perhaps be tagged? [16:34:23] I think gehel would be best to have the final say on closing it but I think that has been entirely superseded by the last year of work across SRE [16:34:49] rzl: jenkins is working again. actually an entirely new jenkins on a new machine and new Java [16:35:14] mutante: thanks! [16:38:12] sorry for the interruption of the puppet window. was on calendar but only added last night, admittedly. very happy it worked this time though, because it did not last time. [16:54:23] hnowlan: Krinkle: one +1 from at least, fwiw [17:34:21] sukhe: I'm back and ready to get started on the etcd prep work. let me know if Traffic still has availability in this time range, or whether we should target some time later. [17:34:50] swfrench-wmf: let's do it. [17:35:51] great, thank you! so, there are two parts to this, which should be safe to proceed in parallel: the LVS changes to switch pybals to eqiad and the SRV record changes for everyone else [17:36:21] let me take care of the former and you can do the latter? [17:36:23] I'll take the latter, as I'll need to coordinate it with client restarts (e.g., confd) [17:36:26] thanks, sounds good [17:36:27] awesome [17:37:11] sukhe: once the DNS changes are live, I'll check on the liberica-cp daemons to see whether we'll need restarts on that side as well. that can happen after the LVS bits are done. [17:38:08] cool, can also help with that if required [17:38:44] awesome, thank you! [17:44:11] 2014 looks good moving on [17:44:23] awesome [17:45:51] I'm going to let the 5m TTL expire on the SRV records while monitoring as (some) client traffic slowly shifts to eqiad. then I'll get started with client restarts. [17:49:16] that looks like what I'd expect: MediaWiki -> etcd traffic in codfw is slowly tapering off and picking up in eqiad (see drop-off in non-quorum GET traffic in https://grafana.wikimedia.org/goto/ffrev1htiejggc?orgId=default) [17:53:09] swfrench-wmf: all good and done on pybals in codfw [17:53:27] thank you <3 [17:53:54] I'm going to get the client restarts going now, then go check on libericas. I'll keep you posted on that latter. [17:54:28] thanks [18:08:04] sukhe: as suspected, liberica-cp daemons in ulsfo and eqsin are hanging onto cached connections toward codfw etcd nodes, and will thus need restarts. [18:09:05] swfrench-wmf: ok, I can get to it in ~15ish [18:09:09] or wait, brett is back [18:09:52] sounds good, thank you! I'm also happy to "learn to fish" as it were [18:10:30] please go ahead :) [18:27:47] alright, reading through the various cookbooks, I think what we actually want is `cookbook sre.loadbalancer.upgrade --seamless --reason 'Drop etcd connections' restart` rather than anything specifically offered by `sre.loadbalancer.admin` (unless we know for a fact that a `config_reload` action will recreate all watches) [18:28:46] this is based on the guidance originally posted in https://phabricator.wikimedia.org/T352245#10843210 for this kind of a scenario (modulo the typo in the name of the cookbook; it's clear that it's `upgrade` due to the `--seamless`) [18:29:09] ^ sukhe or brett: could I get a +1 on that before proceeding? [18:30:22] Without digging into the liberica code, I'd say that sounds good [18:30:45] swfrench-wmf: +1 [18:31:05] thanks, all :) [18:31:08] good digging :) [18:31:19] cool, I'll proceed with that, starting in ulsfo, then eqsin [18:31:31] I've encountered this code path before, although I think it was re-resolution of something else [18:31:33] (i.e., "usual order" in the event that things go south) [18:38:49] sorry folks [18:40:18] swfrench-wmf: things are looking good? [18:41:03] sukhe: yes, thanks! the above cookbook incantation seems to be doing the right thing :) [18:41:06] thankfully edge sites Liberica work is less stressful, not only because of Liberica but because worse case we depool :) [18:41:16] swfrench-wmf: nice! [18:41:23] and yeah, in Liberica we trust [18:41:28] +1 [18:42:10] * swfrench-wmf is going to hold for a few minutes before bounding eqsin [18:42:14] *bouncing [18:54:09] alright, I believe we are done for today. thank you all very much for your time, especially sukhe for wrangling the LVS restarts :) [18:54:43] swfrench-wmf: you did all the hard work. thanks! hope you enjoyed the Liberica experience [18:55:06] did do :) [18:55:13] and then once the maintenance is done, we will do this dance again [18:55:29] * swfrench-wmf nods [18:56:17] and probably a couple more times in the coming weeks, since we (service ops) need to do some conf* reimages. this was a _very_ useful saw-sharpening exercise ahead of that. [18:56:39] nice [21:00:22] hm, admin-ng has some undeployed changes in eqiad, codfw, and staging, related to cassandra-ml-cache-a -- does this ring a bell? https://www.irccloud.com/pastebin/tiEJ5OSw/ [21:01:08] not sure if it was something deleted from the charts repo but not deployed yet (in which case, is it safe to deploy?) or something never committed to the charts repo but deployed from someone's homedir [21:03:29] the ml-cache hosts were decommisioned a few days ago: https://phabricator.wikimedia.org/T430654 [21:03:42] so if these are all removals, they seem fine [21:06:01] ah thanks! [21:06:47] I see https://gerrit.wikimedia.org/r/c/operations/deployment-charts/+/1306681 unmerged too, hm [21:16:21] btullis: wouldn't ping you at this hour, but I see you helmfiling :) happen to have context on the above? [21:17:09] in particular I see admin_ng was deployed for the ml-* and dse-* clusters but not for wikikube or aux, pretty sure I could do so without stepping on ongoing work, but I'd like to confirm before doing it :)