[01:11:57] inflatador: Sorry for the delay. Glad things are looking good! [06:05:48] 10netops, 06Traffic, 06Data-Persistence, 06Infrastructure-Foundations, 06ServiceOps new: codfw: rack B3 maintenance - https://phabricator.wikimedia.org/T430909#12093069 (10MoritzMuehlenhoff) [07:23:41] 06Traffic, 07affects-translatewiki.net, 13Patch-For-Review: Translatewiki.net unable to fetch updates from gerrit over https - https://phabricator.wikimedia.org/T430613#12093214 (10Nikerabbit) [08:27:51] hi folks, FYI I'm about to move dumps-nfs service to the shared public IP with http/rsync, no impact expected https://gerrit.wikimedia.org/r/c/operations/puppet/+/1307338 [10:07:38] 06Traffic, 06Data-Engineering (Q4 FS25/26 April 1st - June 30st), 13Patch-For-Review: Add X-Provenance data to webrequest_sampled_live - https://phabricator.wikimedia.org/T427068#12093643 (10Fabfur) 05Open→03Resolved I think this can be closed [12:14:29] \o I have rolled out the non-pybal/lvs changes for updating our (ML) k8s cluster control planes to IPIP (https://gerrit.wikimedia.org/r/c/operations/puppet/+/1308080). I have also run the IPIP test cookbook (sre.loadbalancer.check-ipip) and dry-ran the final cookbook sre.loadbalancer.migrate-service-ipip for both sites. I'd like to now run the actual cookbook once [12:14:31] https://gerrit.wikimedia.org/r/c/operations/puppet/+/1308094 is reviewed and merged. So I am asking for your kind permission to restart pybal (twice). [12:30:58] (I have taken the liberty of adding g_odog as a reviewer) [12:39:46] 10netops, 06Traffic, 10Cloud-VPS, 06Data-Platform-SRE, and 3 others: Plan to make clouddumps more resilient and easier to operate - https://phabricator.wikimedia.org/T411248#12094337 (10fgiunchedi) >>! In T411248#12084087, @fgiunchedi wrote: >>>! In T411248#12080727, @ayounsi wrote: >>>>! In T411248#120760... [12:56:33] 10netops, 06Infrastructure-Foundations, 06SRE: cr2-esams rpd failure after enabling bgp 'graceful-shutdown' (June 2026) - https://phabricator.wikimedia.org/T429386#12094364 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=99c0df32-80dc-453a-920f-5d3957387031) set by cmooney@cumin1003 for 1... [12:57:13] 10netops, 06Infrastructure-Foundations, 06SRE: cr2-esams rpd failure after enabling bgp 'graceful-shutdown' (June 2026) - https://phabricator.wikimedia.org/T429386#12094370 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=8f5cefde-0060-47c2-8575-70bfa00d21f7) set by cmooney@cumin1003 for 1... [13:20:25] klausman: hi. how does 14:00 UTC sound? [13:21:14] In 40m or so? Sure, wfm [13:21:47] yes thanks, I can help review then [13:23:08] 10netops, 06Infrastructure-Foundations, 06SRE: cr2-esams rpd failure after enabling bgp 'graceful-shutdown' (June 2026) - https://phabricator.wikimedia.org/T429386#12094443 (10cmooney) Upgraded cr3-eqsin also as PIC needed bouncing. And created some instructions to do the traffic drain without 'graceful-shu... [13:28:28] hello [13:28:44] how to feature request: Wikimedia DNS over QUIC (DoQ) support? [13:32:54] 06Traffic: Upgrade Traffic hosts to trixie - https://phabricator.wikimedia.org/T401832#12094482 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by cdobbins@cumin1003 for host dns7002.wikimedia.org with OS trixie [13:58:12] oh they left [14:01:03] klausman: looking at the patches [14:01:15] ty! [14:05:01] klausman: +1ed. let us know if you need us to do something :) [14:06:14] ack, ty! Will run the testst+dry run again, then make one last note here before running the actual cb [14:12:35] tests and dry-runs ok, will now run the cb for codfw, restarting the pybal there in the process. [14:13:16] thanks, lvs2014 first as you probably know [14:13:21] the backup LVS [14:13:46] The cb does the restarts for me, so I trusts it :) [14:20:19] CB completed without error. Keeping an eye on things for a few mins (and poking the ctrl API), then I'll do the eqiad side. [14:20:44] thanks, looking [14:22:25] looks good [14:22:36] yep [14:22:51] thx! will proceed in a Zurich minute [14:22:53] I've noted to add some unit test for this too [14:23:01] ooops sorry, wrong channel [14:28:47] hey folks, i'm trying to reimage a k8s worker node and move the vlan it's on, but we're failing to update dns on dns7002 with "sudo: rec_control: command not found" - is this something that'll be retryable shortly? it looks like dns7002 was recently reimaged, so i wasn't sure how to proceed [14:29:30] proceeding with IPIP bits for eqiad [14:30:14] bjensen: yeah, that host is under reimaging [14:30:42] it is technically depooled but I see why it is still showing up there hmm [14:31:32] cjd91: where is the cookbook at currently? [14:31:52] bjensen: is there any option to skip the host temporarily? I am assuming not but I am curious what options are presented right now? [14:32:15] sukhe: i could skip, it seems the command was successful on all the other dns hosts, i just wasn't sure if that would leave us with a problematic state [14:32:26] bjensen: ah ok [14:32:36] 10netops, 06Traffic, 06Data-Persistence, 06Infrastructure-Foundations, and 2 others: codfw: rack B3 maintenance - https://phabricator.wikimedia.org/T430909#12094907 (10Scott_French) Thanks @ayounsi - I'll aim to have this complete during my day today and will follow up here once it's safe to proceed. While... [14:32:38] and yeah, it should not be a problem. the host is depooled from serving DNS queries and authdns-update [14:32:41] so we should be good yeah [14:32:44] grand, thanks very much :) [14:32:57] unfortunately there is no real fix per se. for long-term depools, we remove the host from Puppet [14:33:18] for shorter reimages like this, which are very rare, the DNS host is not serving any queries [14:33:26] but it still shows up because it's in the cumin alias [14:33:35] gotcha, that makes sense [14:33:41] for some cookbooks, we do a dynamic check that's based on the pooled state of the DNS host but not all [14:34:58] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-ulsfo, 06SRE: ULSFO: OOB IPV6 down - https://phabricator.wikimedia.org/T430599#12094934 (10Papaul) @cmooney see below email from DR thanks ` From our investigation, we confirmed that the source network (2607:fb58:9000:7::/64) is reachable within ou... [14:36:44] sukhe: cookbook complete for eqiad, no errors, ML clusters looking good [14:37:02] klausman: nice and thanks! [14:37:13] right back at ya :) [14:53:52] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-ulsfo, 06SRE: ULSFO: OOB IPV6 down - https://phabricator.wikimedia.org/T430599#12095011 (10cmooney) @papaul hmm it looks like it's working different now: ` cmooney@mr1-ulsfo> traceroute 2620:0:861:3:208:80:154:78 routing-instance oob-transit wait 2... [15:47:08] rzl: Anything I can do to support you with the ipoid removal? [15:48:13] brett: thanks, I was just about to give a heads up :) nothing specific, but I'll let you know if I have trouble [15:48:24] gl! [16:36:09] brett: fyi, waited for CI to be back online but I'm going ahead now [16:36:16] ack! [16:38:12] brett: looking ahead to step 3 at https://wikitech.wikimedia.org/wiki/LVS#Remove_the_service_from_the_load-balancers_and_the_backend_servers -- which backup LVS servers should I be looking at? [16:38:19] and then same for the active at step 5 [16:38:44] I'm assuming those cumin queries get me to the right place, but double-checking :) [16:39:12] they do. I think it's mostly for those that might be a little less familiar than you are of all this :) [16:41:16] haha uh oh, don't give me too much credit [16:42:40] uh hang on, what's the correct cumin query for the backup servers? step 3.1 says A:lvs-low-traffic-eqiad twice [16:43:30] oof yeah, that's really bad documenting [16:44:22] should that be lvs-secondary-eqiad and then lvs-low-traffic-eqiad? trying to piece it together from 3 and 5 [16:45:13] rzl: yes it should. I'll update it [16:45:28] to double-check: ipoid is of class low-traffic? [16:46:38] I was just checking that, where's the best place? [16:47:43] It *should* be hieradata/common/service.yaml but I see that it references k8s-ingress-wikikube stuff [16:47:56] sukhe: Are there special instructions for removal/addition of a k8s service in lvs? [16:48:07] I've not done this before and the lvs doc page is just terrible [16:48:36] (for clarity there are k8s steps after the lvs steps, I've got that handled, but I don't know what the traffic class is if none is listed in service.yaml) [16:50:03] I'm wondering if it somehow piggybacks off of the kubernetes service [16:51:13] brett: yeah, the documentation is ... [16:51:14] k8s-ingress-wikikube is low-traffic, so I guess it must be, in order to go through the same pybals -- and I mean, low-traffic has to be the right answer anyway [16:51:17] but I'd love to understand how that works [16:51:25] yeah, all k8s stuff is low-traffic [16:51:39] just confirmed, it's on lvs2013 for codfw [16:51:48] so 2014 first and then 2013 after [16:51:48] ye [16:51:53] thanks :) [16:52:06] we should mention this somewhere in the documentation, we don't do much of these removals [16:52:23] and also in general at least I don't do a good job at updating the docs, so there's that [16:52:34] well it's intimidating to add new docs to a tire fire like that [16:53:00] most of this will go away with Liberica, the steps are quite different [16:53:05] the basic stuff won't but the pybal steps [16:53:17] so the cleanup is scheduled™ for when that happens, even per our last discussion on this [16:55:10] oh yay and now wt is down [16:55:24] brett: wfm? [16:55:43] thankfully was minor and is back up. Just feels like everything today doesn't wanna work :) [17:04:47] looking ahead again, after the pybal restarts -- is there actually an ipvsadm step? the IP and port are shared, as written, I would --delete-service the cluster ingress IP and port, which doesn't sound like what I want [17:05:09] yeah that shouldn't happen in this case [17:05:18] (and we should add to the docs, in addition to the above :>) [17:06:46] 👍 [17:07:00] in that case, after the 4x pybal restarts, I'm hands-off? [17:08:34] rzl: from the operational side yeah but note the "final removal" step in https://wikitech.wikimedia.org/wiki/LVS#Remove_a_load_balanced_service [17:08:49] basically removing this entire stanza from service.yaml and the related conftool bits if any [17:09:08] yep perfect [17:09:10] thanks! [17:09:18] 301 to brett on that but yeah :) [17:09:24] (the thanks) [17:30:00] gone from service.yaml, I'll follow up with the k8s side of the work later on :) thanks again both! [17:43:36] thank you! [18:19:36] 06Traffic, 06Data-Persistence: Data Lake - mediarequests - adapt HQL query to include counting mediarequests from webrequest_text - https://phabricator.wikimedia.org/T431475 (10Ottomata) 03NEW [18:19:50] 06Traffic, 06Data-Engineering, 06Data-Persistence: Data Lake - mediarequests - adapt HQL query to include counting mediarequests from webrequest_text - https://phabricator.wikimedia.org/T431475#12096445 (10Ottomata) [19:32:47] 10netops, 06Traffic, 06Data-Persistence, 06Infrastructure-Foundations, 06ServiceOps new: codfw: rack B3 maintenance - https://phabricator.wikimedia.org/T430909#12096752 (10ssingh) >>! In T430909#12096739, @Scott_French wrote: > **Status** > > Alright, the (updated) procedure in T430909#12091874 is compl... [23:35:53] 06Traffic, 13Patch-For-Review, 07Sustainability (Incident Followup): Experiment with single backend CDN nodes - https://phabricator.wikimedia.org/T288106#12097579 (10BCornwall) It's been a week since running and the results can be viewed: [[ https://grafana-rw.wikimedia.org/goto/bfrfm0ixeeqkga?orgId=default...