[07:27:04] We'll need one more rolling PyBal restart at some point today on `lvs1019`/`lvs1020` so `wdqs1014` picks up its new address; it's locally healthy but safely out of pool until then [07:27:44] Meanwhile I think I finally understand the issue that has been producing this: setting `pooled=no` (including `sudo depool`) retains pybal's cached IP across a VLAN move; `pooled=inactive` removes the object, so re-pooling resolves DNS again [07:28:10] (A somewhat unfortunate result of the overloaded terminology with `depool`, which usually means the equivalent of `sudo depool` but in this case properly depooling means it must be set inactive) [09:35:43] I am about to do a restart of pybal on lvs1019 and lsv1020 to pick up the 1315900 patch that I mentioned above. [09:50:28] I'm getting errors from k8s-ingress-dse-postgresql_31432 in postgresql after the restart. Connection to the realservers on port 31432 failed. Investigating now. [10:00:59] btullis: Are you finding anything interesting? [10:04:30] I think I may have got a port number incorrect, but I don't know whether it's safe for me just to keep working on this and 'fail forward' - or if it would be better for me to revert. [10:05:47] The pybal service on lvs1019 and lvs1020 is healthy, but there is this alert. https://alerts.wikimedia.org/?q=alertname%3D~Pybal for k8s-ingress-dse-postgresql_31432 [10:06:31] I think the worst that can happen is that PyBal will complain that "wrong" port went away, but that's fixable by restarting PyBal [10:07:03] OK, thanks. [10:07:52] We can also rollback and do a new patch with the correct port, but I don't think that's required. [10:12:21] Oh it might be something else. I'm continuing to look at this, as long as you don't mind the alert being there. I can't find a way to silence just the postgres relayed pybal alert. [10:12:36] No worries, we know what it is :-) [10:18:41] Thx [10:20:56] I think I have found it. [10:32:41] Cool. What is it? [10:38:42] It's this: https://gerrit.wikimedia.org/r/c/operations/deployment-charts/+/1318075 - A missing networkpolicy allowing inbound connections to the new port on the ingressgateways. [10:41:40] aaah :-) [11:10:18] That has fixed it and pybal is now happy - or at least happier with me. There are a couple of alerts about kubstage hosts, but they're not mine. [11:21:47] hello! I'm going to depool ulsfo for a router upgrade and deploying a new feature [11:50:50] 10netops, 06Infrastructure-Foundations, 06SRE: cr4-ulsfo: upgrade to Junos 23.4R2-S8 - https://phabricator.wikimedia.org/T431752#12158344 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=fe3278b9-2acb-46d1-af9d-66d64acae08f) set by ayounsi@cumin1003 for 2:00:00 on 3 host(s) and their servi... [11:55:39] btullis: Ah, the "someone else problem field" strikes again [12:37:01] 10netops, 06Infrastructure-Foundations, 06SRE: cr4-ulsfo: upgrade to Junos 23.4R2-S8 - https://phabricator.wikimedia.org/T431752#12158517 (10ayounsi) 05Open→03Resolved a:03ayounsi upgraded! [12:56:08] 10netops, 06Infrastructure-Foundations, 13Patch-For-Review: Junos: investigate BGP rib sharding - https://phabricator.wikimedia.org/T320264#12158568 (10ayounsi) BGP rib sharding has been enabled on cr4-ulsfo Some non-scientific numbers. Before enabling RIB sharding, removing the "disabled" statement from t... [13:30:34] 10netops, 06Traffic, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: eqsin downtime scheduling for switch upgrade - https://phabricator.wikimedia.org/T433097#12158671 (10ssingh) Thanks @BCornwall. @RobH: Is there a preference for the day of the week from your end? I am assuming we can't do tomorrow Sin... [13:58:56] 10netops, 06Traffic, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: eqsin downtime scheduling for switch upgrade - https://phabricator.wikimedia.org/T433097#12158773 (10RobH) Right now it is 13:55 GMT = 21:55 Singapore, so if we schedule now it won't be confirmed until Tuesday, for possibly a Wednesda... [14:07:54] 10netops, 06Traffic, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: eqsin downtime scheduling for switch upgrade - https://phabricator.wikimedia.org/T433097#12158818 (10ssingh) >>! In T433097#12158773, @RobH wrote: > > Proposed window: 2026-07-30 @ 01:00 GMT (9AM Singapore). Depool to take place a fe... [14:09:33] 10netops, 06Infrastructure-Foundations: Junos: investigate BGP rib sharding - https://phabricator.wikimedia.org/T320264#12158831 (10ayounsi) [14:10:27] 06Traffic, 06Data-Platform-SRE (2026-07-03 - 2026-07-31), 13Patch-For-Review: Standardize on a single web proxy hostname - https://phabricator.wikimedia.org/T429338#12158840 (10brouberol) Based on the link @MoritzMuehlenhoff shared, I'm going to say that having Airflow DAGs use urldownloader to interact with... [14:10:44] 06Traffic, 06Data-Platform-SRE (2026-07-03 - 2026-07-31), 13Patch-For-Review: Standardize on a single web proxy hostname - https://phabricator.wikimedia.org/T429338#12158843 (10BTullis) >>! In T429338#12023924, @cmooney wrote: > There are also the hcaptcha-proxy VMs too. > > In terms of the original point I... [14:15:30] 06Traffic, 06Data-Platform-SRE (2026-07-03 - 2026-07-31), 13Patch-For-Review: Standardize on a single web proxy hostname - https://phabricator.wikimedia.org/T429338#12158932 (10brouberol) ` brouberol@an-worker1200:~$ telnet url-downloader.eqiad.wikimedia.org 8080 Trying 2620:0:861:2:208:80:154:136... Connect... [14:20:00] 06Traffic, 06Data-Platform-SRE (2026-07-03 - 2026-07-31), 13Patch-For-Review: Standardize on a single web proxy hostname - https://phabricator.wikimedia.org/T429338#12158966 (10BTullis) >>! In T429338#12158840, @brouberol wrote: > Based on the link @MoritzMuehlenhoff shared, I'm going to say that having Airf... [14:25:40] 10netops, 06Traffic, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: eqsin downtime scheduling for switch upgrade - https://phabricator.wikimedia.org/T433097#12158992 (10RobH) I'm going to estimate the work starting at 9AM Singapore will take up to 4 hours to complete, so we'd like the site depooled fr... [14:30:01] 10netops, 06Traffic, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: eqsin downtime scheduling for switch upgrade - https://phabricator.wikimedia.org/T433097#12159006 (10ssingh) >>! In T433097#12158992, @RobH wrote: > I'm going to estimate the work starting at 9AM Singapore will take up to 4 hours to c... [14:37:04] 10netops, 06Traffic, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: eqsin downtime scheduling for switch upgrade - https://phabricator.wikimedia.org/T433097#12159020 (10RobH) I've sent an email to Jin to confirm the date window of 2026-07-29 @ 23:00 GMT to 2026-07-30 @ 03:00 GMT, which is 2026-07-30 @... [14:43:50] 06Traffic, 06Data-Platform-SRE (2026-07-03 - 2026-07-31), 13Patch-For-Review: Standardize on a single web proxy hostname - https://phabricator.wikimedia.org/T429338#12159043 (10brouberol) 05Open→03In progress [16:08:25] 10netops, 06Traffic, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: eqsin downtime scheduling for switch upgrade - https://phabricator.wikimedia.org/T433097#12159656 (10BCornwall) Confirming I can do those times (2026-07-29 @ 16:00-20:00 PDT) [16:42:49] 10netops, 06Traffic, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: eqsin downtime scheduling for switch upgrade - https://phabricator.wikimedia.org/T433097#12159857 (10RobH) Jin confirmed via email just now for the window: 2026-07-29 @ 23:00 GMT to 2026-07-30 @ 03:00 GMT I'll update our directions... [16:43:40] 10netops, 06Traffic, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: eqsin downtime scheduling for switch upgrade - https://phabricator.wikimedia.org/T433097#12159861 (10RobH) [21:30:34] hello traffic friends - we would like to proceed with some maintenance on the codfw conf* hosts this week (reimage [0]), which will require depooling client traffic from them for at least 24h (i.e., similar to what we did for the rack B3 maintenance a couple of weeks ago). [21:30:34] * as usual, this will involve codfw pybal restarts for [2] and some number of liberica restarts after [3] is merged (along with rolling confd restarts, etc.) [21:30:34] * any concerns about that happening this week, potentially as early as tomorrow? we should be able to pause any operations that overlap with the eqsin switch maintenance [21:30:34] [0] https://phabricator.wikimedia.org/T428495 [21:30:34] [1] https://phabricator.wikimedia.org/T430909 [21:30:35] [2] https://gerrit.wikimedia.org/r/c/operations/puppet/+/1314076 [21:30:35] [3] https://gerrit.wikimedia.org/r/c/operations/dns/+/1314077 [22:24:03] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-eqiad, 06SRE: Standardize management routers interfaces - https://phabricator.wikimedia.org/T421674#12161199 (10Papaul)