[00:32:07] 10netops, 06Traffic, 06Data-Persistence, 06Data-Platform-SRE, and 5 others: codfw: rack B2 maintenance 2026-07-01 11:00 am CT - https://phabricator.wikimedia.org/T429861#12068530 (10Papaul) @BCornwall about 40 minutes from adding the image to reboot [05:24:42] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-ulsfo, 06SRE: ULSFO: OOB IPV6 down - https://phabricator.wikimedia.org/T430599#12068734 (10ayounsi) FYI it works from my home internet: ` laptop:~$ ping -6 mr1-ulsfo.oob.wikimedia.org PING mr1-ulsfo.oob.wikimedia.org (2607:fb58:9000:7::2) 56 data by... [06:38:03] 06Traffic, 07affects-translatewiki.net: Translatewiki.net unable to fetch updates from gerrit over https - https://phabricator.wikimedia.org/T430613 (10Nikerabbit) 03NEW [06:40:09] Could someone have a look at https://phabricator.wikimedia.org/T430613? It completely breaks translation update workflow for translatewiki.net. [07:02:33] 06Traffic, 07affects-translatewiki.net: Translatewiki.net unable to fetch updates from gerrit over https - https://phabricator.wikimedia.org/T430613#12068876 (10Nikerabbit) There is a report that we are at least two weeks behind picking up new strings to translate: https://translatewiki.net/w/i.php?title=Suppo... [07:08:04] 06Traffic, 07affects-translatewiki.net: Translatewiki.net unable to fetch updates from gerrit over https - https://phabricator.wikimedia.org/T430613#12068880 (10Nikerabbit) [07:09:38] Nikerabbit: it was due to heavy scraping of gerrit, it should be ok now [07:26:10] arnaudb: yes but that is not the root cause... we wouldn't have delay of two weeks of fetching updates if it was just temporary scraping issue [07:29:16] Nikerabbit: looking at the ticket, thanks. It could be also helpful if possible to report the whole error message related to the 429 status code [07:29:22] if you have access to [07:40:06]

Too many requests (1059162)

[07:40:08] this part? [07:41:34] yes, thanks! [07:50:28] 06Traffic, 10Liberica, 10Prod-Kubernetes, 06ServiceOps new, 07Kubernetes: Add missing wikikube workers to conftool-data - https://phabricator.wikimedia.org/T420729#12068997 (10MLechvien-WMF) [07:50:51] 06Traffic, 06MW-Interfaces-Team, 06ServiceOps new, 10ServiceOps-SharedInfra: map the /api/ prefix to /w/rest.php - https://phabricator.wikimedia.org/T364400#12069002 (10MLechvien-WMF) [07:50:57] 06Traffic, 10Liberica, 10Prod-Kubernetes, 06ServiceOps new, 07Kubernetes: Migrate Wikikube k8s apiserver and services to IPIP - https://phabricator.wikimedia.org/T420436#12069007 (10MLechvien-WMF) [07:56:44] 06Traffic, 10Liberica, 10Prod-Kubernetes, 06ServiceOps new, 07Kubernetes: Migrate Wikikube k8s apiserver and services to IPIP - https://phabricator.wikimedia.org/T420436#12069038 (10MLechvien-WMF) a:03JMeybohm [07:57:56] 06Traffic, 10Liberica, 10Prod-Kubernetes, 06ServiceOps new, 07Kubernetes: Add missing wikikube workers to conftool-data - https://phabricator.wikimedia.org/T420729#12069046 (10MLechvien-WMF) a:03JMeybohm [08:59:20] 06Traffic, 07affects-translatewiki.net: Translatewiki.net unable to fetch updates from gerrit over https - https://phabricator.wikimedia.org/T430613#12069339 (10neriah) Seems related to T418087. [09:12:14] 10netops, 06Traffic, 10Cloud-VPS, 06Data-Platform-SRE, and 4 others: Plan to make clouddumps more resilient and easier to operate - https://phabricator.wikimedia.org/T411248#12069417 (10fgiunchedi) [09:22:52] hi folks, I'm adding a new service (dumps-nfs) and per https://wikitech.wikimedia.org/wiki/LVS#Configure_the_load_balancers I got https://gerrit.wikimedia.org/r/c/operations/puppet/+/1306631 out, ok to proceed ? [09:25:37] slyngs: https://wikimedia.slack.com/archives/C01H5HCH5DZ/p1782811501379969 [09:27:26] Nice [09:41:44] godog: ack [09:43:18] 06Traffic, 07affects-translatewiki.net: Translatewiki.net unable to fetch updates from gerrit over https - https://phabricator.wikimedia.org/T430613#12069635 (10neriah) [09:44:24] fabfur: thank you, proceeding [11:08:43] FIRING: HaproxyKafkaSocketDroppedMessages: Sustained high rate of dropped messages from HaproxyKafka - https://wikitech.wikimedia.org/wiki/HAProxyKafka#HaproxyKafkaSocketDroppedMessages - https://grafana.wikimedia.org/d/d3e4e37c-c1d9-47af-9aad-a08dae2b3fd5/haproxykafka?orgId=1&var-site=esams&var-instance=cp3067&viewPanel=panel-19 - https://alerts.wikimedia.org/?q=alertname%3DHaproxyKafkaSocketDroppedMessages [11:13:43] FIRING: [6x] HaproxyKafkaSocketDroppedMessages: Sustained high rate of dropped messages from HaproxyKafka - https://wikitech.wikimedia.org/wiki/HAProxyKafka#HaproxyKafkaSocketDroppedMessages - https://alerts.wikimedia.org/?q=alertname%3DHaproxyKafkaSocketDroppedMessages [11:18:43] FIRING: [6x] HaproxyKafkaSocketDroppedMessages: Sustained high rate of dropped messages from HaproxyKafka - https://wikitech.wikimedia.org/wiki/HAProxyKafka#HaproxyKafkaSocketDroppedMessages - https://alerts.wikimedia.org/?q=alertname%3DHaproxyKafkaSocketDroppedMessages [11:23:43] RESOLVED: [6x] HaproxyKafkaSocketDroppedMessages: Sustained high rate of dropped messages from HaproxyKafka - https://wikitech.wikimedia.org/wiki/HAProxyKafka#HaproxyKafkaSocketDroppedMessages - https://alerts.wikimedia.org/?q=alertname%3DHaproxyKafkaSocketDroppedMessages [11:25:25] 06Traffic, 10Lift-Wing, 06Machine-Learning-Team (Q4 FY2025-26), 13Patch-For-Review: Host Qwen 3.6-27B as an inference service - https://phabricator.wikimedia.org/T425680#12070060 (10gkyziridis) **Attempt: deploy qwen36-27b (FP8) on ml-serve1012 ** Changes (vs `ml-serve1013/1014`: - `affinity` → `ml-serve1... [11:35:13] \o I'll be submitting https://gerrit.wikimedia.org/r/c/operations/puppet/+/1305623 (IPIP update for ML k8s staging) soon and then running the sre.loadbalancer.migrate-service-ipip --- is that okay (it implies a pybal restart) [11:56:11] 10netops, 06Traffic, 06Infrastructure-Foundations, 06cloud-services-team (FY2025/2026-Q3-Q4): LOAD_BALANCER_HEALTH_CHECKS firewall set and asymmetry lvs1018 / lvs1020 - https://phabricator.wikimedia.org/T430651 (10fgiunchedi) 03NEW [11:57:55] 10netops, 06Traffic, 06Infrastructure-Foundations, 06cloud-services-team (FY2025/2026-Q3-Q4), 13Patch-For-Review: LOAD_BALANCER_HEALTH_CHECKS firewall set and asymmetry lvs1018 / lvs1020 - https://phabricator.wikimedia.org/T430651#12070204 (10fgiunchedi) https://gerrit.wikimedia.org/r/c/operations/puppet... [11:58:43] the bandaid in ^ is to get unblocked, unless folks have a better suggestion ! [12:07:57] 10netops, 06Infrastructure-Foundations, 06SRE: Blackbox probe for TLS cert expriy failing on multiple eqiad SR-Linux nodes - https://phabricator.wikimedia.org/T429242#12070253 (10cmooney) lsw1-d2-eqiad has been left in a "bad" state and I opened 05827811 against it with Nokia. [12:25:49] 10netops, 06Traffic, 06Infrastructure-Foundations, 06cloud-services-team (FY2025/2026-Q3-Q4), 13Patch-For-Review: LOAD_BALANCER_HEALTH_CHECKS firewall set and asymmetry lvs1018 / lvs1020 - https://phabricator.wikimedia.org/T430651#12070301 (10MoritzMuehlenhoff) All the IPs allowed by LOAD_BALANCER_HEALTH... [12:41:18] 10netops, 10Cloud-VPS, 06Infrastructure-Foundations, 06SRE, 06tools-infrastructure-team: Upgrade cloudsw1-e4-eqiad - https://phabricator.wikimedia.org/T429013#12070397 (10ayounsi) Unfortunately there is no easy way to perform the maintenance without impacting the devices connected to that switch. [12:44:16] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-eqiad, 06SRE: Install new MPC10E-10C line cards on cr1-eqiad and cr2-eqiad slot 0. - https://phabricator.wikimedia.org/T426343#12070412 (10cmooney) I spoke to @VRiley-WMF about this and we are going to schedule the change for Wed July 1st starting a... [12:45:32] godog: about https://phabricator.wikimedia.org/T430651 are you using IPIP for the new service? [12:47:03] XioNoX: seems like it [12:47:13] I had to check, genuinely had no idea [12:47:27] I copied what we're doing for dumps-rsync and dumps-http tbh [12:47:57] ok, I don't know the service well enough, but iirc with IPIP the healthchecks follow the data path [12:48:24] makes sense yeah [12:54:00] Friendly ping about merging https://gerrit.wikimedia.org/r/c/operations/puppet/+/1305623 and running sre.loadbalancer.migrate-service-ipip on ML k8s staging (implies pybal restart, hence me asking here) [12:55:28] fabfur ^ (I see that you +1ed the change) [12:57:53] XioNoX: they sure do for liberica (katran), IIRC also for LVS too [12:59:27] 10netops, 06Traffic, 10Cloud-VPS, 06Data-Platform-SRE, and 3 others: Plan to make clouddumps more resilient and easier to operate - https://phabricator.wikimedia.org/T411248#12070478 (10fgiunchedi) [13:08:10] 06Traffic, 06ServiceOps new, 06Machine-Learning-Team (Q4 FY2025-26): k8s changes needed to allow article topic (and other future isvcs) to use the kserve v2 inference protocol (and gRPC) - https://phabricator.wikimedia.org/T424049#12070508 (10isarantopoulos) 05Open→03Resolved All changes have been re... [13:09:38] 10netops, 06Traffic, 10Cloud-VPS, 06Data-Platform-SRE, and 3 others: Plan to make clouddumps more resilient and easier to operate - https://phabricator.wikimedia.org/T411248#12070530 (10fgiunchedi) I'll be starting the trials on `toolsbeta` and specifically `toolsbeta-test-k8s-worker-nfs-8.toolsbeta` by ov... [13:33:03] 10netops, 06Traffic, 06Infrastructure-Foundations, 06cloud-services-team (FY2025/2026-Q3-Q4), 13Patch-For-Review: LOAD_BALANCER_HEALTH_CHECKS firewall set and asymmetry lvs1018 / lvs1020 - https://phabricator.wikimedia.org/T430651#12070626 (10cmooney) @fgiunchedi thanks for the task. I think @MoritzMueh... [13:57:57] fabfur: I think Arzhel was asking about my pybal restart :) [14:02:35] klausman: shouldn't be an issue in restarting pybal, assuming the cookbook knows how to do it :) [14:02:39] so ack for me [14:04:57] ty! [14:23:25] dry-run passed without errors, doing real run now. [14:28:16] cookbook completed, no errors [14:33:02] 👍 [14:37:03] I'll have to re-run it since I forgot to puppet-merge %-) [14:47:23] second run also a-ok [16:14:38] 06Traffic: GeoDNS: consider sending CN to eqsin - https://phabricator.wikimedia.org/T378744#12071878 (10BCornwall) Based on @CDobbins 's findings and the other comments, it seems ulsfo should be restated. Thoughts, @CDobbins / @ssingh ? [17:48:42] 06Traffic, 06Infrastructure-Foundations, 06ServiceOps new, 06SRE: Scaling urldownloaders by adding redundancy and load balancing - https://phabricator.wikimedia.org/T429175#12072629 (10ssingh) After some discussion in Traffic and confirming with @BBlack, we are going with Option 3, putting this behind the... [18:58:28] 06Traffic, 13Patch-For-Review, 07Sustainability (Incident Followup): Experiment with single backend CDN nodes - https://phabricator.wikimedia.org/T288106#12072995 (10BCornwall) As of this comment cp3081 is now running on single-backend and with one drive [20:10:22] 10netops, 06Infrastructure-Foundations, 06SRE: cr2-esams rpd failure after enabling bgp 'graceful-shutdown' (June 2026) - https://phabricator.wikimedia.org/T429386#12073435 (10cmooney) 05Open→03Resolved [20:27:14] 10netops, 06Infrastructure-Foundations, 06SRE: Create alerting for saturation on sub-rated interfaces - https://phabricator.wikimedia.org/T374614#12073531 (10cmooney) We now have the //CoreRouterInterfaceDropPercent// alerts, which to an extent covers this scenario. It's not exactly the same, but we are in...