[08:08:11] 10netops, 06Infrastructure-Foundations, 10netbox: Automatically run Capirca Netbox script regularly - https://phabricator.wikimedia.org/T361549#12135720 (10ayounsi) 05Open→03Resolved event-rule enabled in prod: https://netbox.wikimedia.org/extras/event-rules/1/ Added a few lines about it in the doc:... [09:29:14] 06Traffic, 10hCaptcha, 06Product Safety and Integrity, 06ServiceOps new, and 2 others: memcached errors seen in hCaptcha health checks - https://phabricator.wikimedia.org/T430340#12136060 (10MLechvien-WMF) [11:50:16] 10netops, 06Infrastructure-Foundations, 06SRE, 13Patch-For-Review: Create alerting for saturation on sub-rated interfaces - https://phabricator.wikimedia.org/T374614#12136425 (10ayounsi) `lang=diff Changes for 1 devices: ['cr2-drmrs.wikimedia.org'] [edit interfaces et-0/0/0] - description "Transport: Ar... [12:22:56] 06Traffic, 06cloud-services-team, 10Infrastructure Security, 06Infrastructure-Foundations, and 5 others: July 2026 Bookworm/Trixie reboots: Data Platform SRE - https://phabricator.wikimedia.org/T431826#12136526 (10klausman) [12:23:30] 06Traffic, 06cloud-services-team, 10Infrastructure Security, 06Infrastructure-Foundations, and 5 others: July 2026 Bookworm/Trixie reboots: Data Platform SRE - https://phabricator.wikimedia.org/T431826#12136528 (10klausman) >>! In T431826#12131542, @elukey wrote: > Hey folks! IIUC this task is to track bot... [12:24:36] 06Traffic, 10Infrastructure Security, 06Infrastructure-Foundations, 06SRE, and 4 others: July 2026 Bookworm/Trixie reboots: Data Platform SRE - https://phabricator.wikimedia.org/T431826#12136531 (10taavi) [13:55:15] 06Traffic, 10bot-traffic-requests, 07SEO: Link preview from iMessage on macOS blocked - https://phabricator.wikimedia.org/T432373#12136768 (10Krinkle) [13:55:32] 06Traffic, 10bot-traffic-requests, 06Future-Audiences, 07SEO: Link preview from iMessage on macOS blocked - https://phabricator.wikimedia.org/T432373#12136769 (10Krinkle) SRE has rolled out a fix. I'll verify later today. [13:55:45] 06Traffic, 10bot-traffic-requests, 06Future-Audiences, 06MediaWiki-Core-Platform-Team (Radar), 07SEO: Link preview from iMessage on macOS blocked - https://phabricator.wikimedia.org/T432373#12136771 (10Krinkle) [13:55:52] 06Traffic, 10bot-traffic-requests, 06Future-Audiences, 06MediaWiki-Core-Platform-Team (Radar), 07SEO: Link preview from iMessage on macOS blocked - https://phabricator.wikimedia.org/T432373#12136772 (10Krinkle) p:05Triage→03High [13:58:00] hey traffic, question on https://wikitech.wikimedia.org/wiki/LVS#Remove_a_load_balanced_service . Do I need 2 puppet changes for `change state: production to state: lvs_setup in hieradata/common/service.yaml` and `remove the service stanza from profile::lvs::realserver::pools` or can we do that in a single change? [14:00:56] inflatador: yeah, you need one change for state: production -> state: lvs_setup and then do the puppet run on A:dnsbox [14:01:12] then another change for lvs_setup to service_setup with ideally a separate patch for the final cleanup [14:01:30] it's not a big deal for the latter but it's best to avoid alerting issues and such [14:01:57] does that make sense? [14:07:00] sukhe thanks, I think I get it. I'll break the above change into 2 changes [14:07:49] haven't done the patch for the final cleanup either, will get that too [14:08:03] inflatador: let us know if we can help [14:11:00] ACK, will do [14:13:38] dear traffic, could be please coord with someone eg on Wed to roll out https://phabricator.wikimedia.org/T432445 ? we have prepped the patches [14:14:02] ideally on EU timezones :) [14:20:11] hmm, this is a good template for my work too ;) [14:22:01] effie: we just have one SRE right now in EU and that's fab(ulous) fabfur [14:22:18] if he can't do it, then I can do it at 14:00 UTC [14:22:29] (the poor fabfur) [14:22:53] effie: if it can't be postponed in the early afternoon (our afternoon) it's ok for me to follow this [14:26:19] fabfur: on Wed? [14:27:09] yes [14:27:52] we can be flexible ! or wait for sukhe a bit later, his timezone is also convenient [14:28:11] I will DM you with details and we coord. Thank you folks! [14:28:27] effie: works for me yeah [14:28:30] let us know [14:29:47] cheers [14:29:55] thanks [14:47:13] sukhe heads up that I merged the above patch and ran `sudo cumin 'A:dnsbox' run-puppet-agent` which initially failed on `dns6002`, but succeeded when I tried again [14:47:30] inflatador: what was the failure on dns6002? [14:47:35] puppet ran without errors the first time on all other hosts [14:48:14] sukhe the only thing cumin said was `6.2% (1/16) of nodes failed to execute command #1: 'run-puppet-agent': dns6002.wikimedia.org`. Maybe an intermittent connection error? [14:48:34] inflatador: could be yeah but thanks for flagging it [14:49:04] trying another run manually [14:50:28] ACK, note that it did succeed when I ran it the 2nd time [14:53:34] inflatador: yeah looks good, so we will attribute it to being intermittent [14:54:01] the thing on my mind is that we reimaged dns7002 to trixie and it's the only host on that, so just wanted to be sure it is not that (I know you said dns6002) and any connection issues in general [14:54:41] https://gerrit.wikimedia.org/r/c/operations/puppet/+/1312517 here's the 2nd patch as requested if you or someone else has time to look. No hurry on this one, we can merge on y'all's schedule if you prefer [14:58:00] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-codfw, 06SRE: codfw: setup MPC10E-10C and SCBE3 - https://phabricator.wikimedia.org/T393552#12137148 (10cmooney) a:05cmooney→03VRiley-WMF @VRiley-WMF hope it's ok to re-assign to you. Juniper are chasing for details in terms of the shipping... [14:59:12] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-codfw, 06SRE: codfw: setup MPC10E-10C and SCBE3 - https://phabricator.wikimedia.org/T393552#12137157 (10cmooney) a:05VRiley-WMF→03cmooney Sorry wrong ticket meant to re-assign the one we're working on in eqiad. [17:52:28] FIRING: KeyholderUnarmed: 1 unarmed Keyholder key(s) on acmechief1002:9100 - https://wikitech.wikimedia.org/wiki/Keyholder - TODO - https://alerts.wikimedia.org/?q=alertname%3DKeyholderUnarmed [17:52:55] yes, fixing [17:58:00] 06Traffic, 10Beta-Cluster-Infrastructure: Project deployment-prep instance deployment-cache-text08 is down - https://phabricator.wikimedia.org/T432498#12138191 (10bd808) 05Open→03Invalid `lang=shell-session bd808@mbp03:~/projects/wmf$ ssh deployment-cache-text08.deployment-prep.eqiad1.wikimedia.cloud L... [17:59:02] 06Traffic, 10GitLab (Project Migration): Migrate DNS repository from Gerrit to Gitlab - https://phabricator.wikimedia.org/T355906#12138216 (10Dzahn) Since T341991#12137429 was declined this might also be declined nowadays. [18:02:28] FIRING: [2x] KeyholderUnarmed: 1 unarmed Keyholder key(s) on acmechief1002:9100 - https://wikitech.wikimedia.org/wiki/Keyholder - TODO - https://alerts.wikimedia.org/?q=alertname%3DKeyholderUnarmed [18:02:42] this is done [18:05:25] FIRING: SystemdUnitFailed: acme-chief-certs-sync.service on acmechief2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [18:05:38] hm ok [18:05:41] let's see [18:06:18] that was because of the keyholder so let's see how it is now [18:07:28] RESOLVED: [2x] KeyholderUnarmed: 1 unarmed Keyholder key(s) on acmechief1002:9100 - https://wikitech.wikimedia.org/wiki/Keyholder - TODO - https://alerts.wikimedia.org/?q=alertname%3DKeyholderUnarmed [18:10:25] RESOLVED: SystemdUnitFailed: acme-chief-certs-sync.service on acmechief2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [19:37:54] 10netops, 06Infrastructure-Foundations: Apply egress Source Address Validation on the Wikimedia core routers - https://phabricator.wikimedia.org/T372158#12138458 (10Southparkfan) Sorry for the delay! Thank you for responding :-). It's an interesting exercise to review what we worked on two years ago. >>! In T3... [19:41:06] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-eqiad, 06SRE: eqiad row A/B switch refresh prep - https://phabricator.wikimedia.org/T418012#12138461 (10RobH) [20:03:00] 10netops, 06Infrastructure-Foundations: Publish, and maintain ASPA records for valid AS14907 upstreams - https://phabricator.wikimedia.org/T372161#12138539 (10Southparkfan) As of January 2026, we can provision ASPA records through's ARIN Hosted RPKI: https://www.arin.net/resources/manage/rpki/aspa/. According... [20:52:50] sukhe: around perchance? im fighting an ipvs diff check attributable to wdqs-internal-main. We expect 1 out of the 3 hosts to be down but all 3 are marked as down-but-pooled rn. my theory is after a reimage with move-vlan pybal is still tryhing to talk to their old IPs not new ones. could use help confirming, and figuring out what the fix is (naively I'd think bounce pybal, which ofc im not touching without u guys around :D) [20:54:27] I guess `enabled/down/not pooled` is more accurate than marked-as-down-but-pooled since I think only 2020 is reporting that specifically [20:56:17] ryankemper: yeah, that all points to the a pybal restart that is required [20:56:26] wdqs-internal-main? [20:56:33] and where did the IPs change, eqiad or codfw? [20:56:33] yup wdqs-internal-main codfw [20:57:15] checking 2014 first (backup) [20:58:42] ryankemper: only wdqs2007 is down. that adds up? [20:58:54] wdqs2007 is actively being reimaged [20:59:07] ok [20:59:32] So in https://config-master.wikimedia.org/pybal/codfw/wdqs-internal-main, `wdqs2020` is actually down, 201[8-9] are stale (servers itself healthy but pybal not talking to them properly) [21:00:06] can you check now? should be better [21:00:13] it was the depool threshold keeping them pooled [21:00:42] wdqs2020 seems unhappy as well [21:01:07] Yeah, dc ops is looking at that one ;( ref T430880 [21:01:08] T430880: Migrate WDQS hosts to Bookworm or later - https://phabricator.wikimedia.org/T430880 [21:01:13] sounds like it might have a power issue or something [21:01:13] oh -intermail vs -mail again, I always get this wrong [21:01:36] sukhe@lvs2013:~$ curl localhost:9090/pools/wdqs-internal-main_443 [21:01:36] wdqs2018.codfw.wmnet: enabled/up/pooled [21:01:36] wdqs2019.codfw.wmnet: enabled/up/pooled [21:01:40] that seems fine now? [21:02:32] sukhe: looks correct now [21:02:48] ok great. so yeah, the IP change probably needed a bump [21:03:04] and the hosts were pooled because of depool threshold [21:03:08] which is .3 [21:03:16] for wdqs-internal-main [21:03:39] yeah makes sense, 2 hosts invisible to pybal, 1 host visible but dead so kept pooled until i manually depooled it [21:04:26] yep [21:04:34] do we not do vlan changes at wmf enough to hit this consistently? I'm surprised this hasn't bit other teams before (maybe bad assumption) [21:04:59] then again i guess it requires either a small pool (like what we have here) or a large pool but vlan moves on almost every host [21:05:05] there is a tricky race condition here which I have been observing more often that I don't have a good explanation for yet [21:05:14] did you depool the host before you started the maintenance? [21:05:24] my suspicion is that if you don't depool manually, pybal does notpick up the changes [21:06:03] but take that with a grain or a fist of salt since I haven't actually confirmed this :) [21:06:09] (nor traced it out fully why that matters) [21:07:13] not pybal specifically but how it interacts with LVS (so most likely LVS needs to be reconfigured for the IP change and that only happens when Pybal restarts) [21:07:20] that seems to be the reasonable explanation [21:07:36] anyway, I think we can consider this resolved now [21:07:59] I always assumed reimage would depool a host automatically, but I guess that's a bad assumption [21:08:31] IIRC reimage requires a --conftool option otherwise it won't [21:08:38] in theory yeah, since the healthcheck fails. what I am not sure is if pybal reattempts that against the new IP [21:09:23] ryankemper is correct re: depool options https://github.com/wikimedia/operations-cookbooks/blob/master/cookbooks/sre/hosts/reimage.py#L66 [21:09:58] still, it would be nice if the cookbook warned us about this, since it encourages you to use `move-vlan`. I'll make a ticket [21:10:42] inflatador: by this you mean that the IP will change or? [21:11:18] I have to step away for now, please ping bret.t if we can help :> [21:11:31] \o [21:11:39] sukhe I mean the cookbook should tell you that if you don't depool the host it will cause problems with pybal [21:11:50] when you're running with `--move-vlan` that is [21:13:24] wait, it doesn't depool?! [21:15:14] -.- [21:18:05] (possibly) adjacently relevant: https://phabricator.wikimedia.org/T430226 [21:19:26] So maybe if `--move-vlan` is supplied, the cookbook should check if conftool flag was passed, and if not passed, if the host has conftool objects then it refuses to go ahead without explicit override [21:19:35] (bc we need to avoid breaking --move-vlan for hosts that aren't pybal backed) [21:20:58] jasmine_ ACK, thanks for sharing [21:21:16] I created T432657 to ask for a warning, feel free to add/edit if anyone has more details or opinions [21:21:17] T432657: Reimage cookbook sets off Pybal alerts when using `--move-vlan` on pooled hosts - https://phabricator.wikimedia.org/T432657