[07:22:03] haproxy fails to restart on various cache nodes: https://puppetboard.wikimedia.org/report/cp4038.ulsfo.wmnet/a191aca50464cf3d0f9117b29e73972d3279af33 [07:22:19] let me check! [07:22:33] journalctl isn't very verbose, only mentions "Some tasks resisted to hard-stop, exiting now" [07:24:41] yeah, checking syntax shows an error on a lua function but that shouldn't be touched since long time [07:25:49] and I don't see relevant changes lately [07:26:30] FIRING: HAProxyRestarted: HAProxy server restarted on cp4038:9100 - https://wikitech.wikimedia.org/wiki/HAProxy#HAProxy_for_edge_caching - https://grafana.wikimedia.org/d/gQblbjtnk/haproxy-drilldown?orgId=1&var-site=ulsfo%20prometheus/ops&var-instance=cp4038&viewPanel=10 - https://alerts.wikimedia.org/?q=alertname%3DHAProxyRestarted [07:26:44] ^^ me [07:28:10] [ALERT] (2606534) : config : parsing [/etc/haproxy/haproxy.cfg:21] : Lua runtime error: The MaxMind DB file contains invalid metadata [07:28:17] this causes the other issues [07:30:25] FIRING: SystemdUnitFailed: haproxy.service on cp4038:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [07:31:29] RESOLVED: HAProxyRestarted: HAProxy server restarted on cp4038:9100 - https://wikitech.wikimedia.org/wiki/HAProxy#HAProxy_for_edge_caching - https://grafana.wikimedia.org/d/gQblbjtnk/haproxy-drilldown?orgId=1&var-site=ulsfo%20prometheus/ops&var-instance=cp4038&viewPanel=10 - https://alerts.wikimedia.org/?q=alertname%3DHAProxyRestarted [07:33:11] given the latest touched file is the proxy.mmdb I suspect there's some issue with that [07:51:33] 06Traffic, 06Machine-Learning-Team, 06ServiceOps new: k8s changes needed to allow article topic (and other future isvcs) to use the kserve v2 inference protocol (and gRPC) - https://phabricator.wikimedia.org/T424049#12088920 (10isarantopoulos) [07:51:43] ok I think I "fixed" it, re-running the service that generates proxy.mmdb on puppetserver and re-running puppet [07:51:51] now I'll try on cp4038 and repool it [07:56:15] mmm that's funny, puppet brings the new proxy.mmdb file on cp7001 but not on cp4038 [07:56:29] no it does, it's just slower [08:00:25] RESOLVED: SystemdUnitFailed: haproxy.service on cp4038:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [08:03:10] 06Traffic, 06collaboration-services, 06Data-Persistence, 06Infrastructure-Foundations, and 2 others: codfw: rack B8 maintenance - https://phabricator.wikimedia.org/T430929#12088990 (10MatthewVernon) apus-be2004 is an apus backend, which will need some care (otherwise the cluster will notice the host as fai... [08:28:38] 06Traffic, 10Lift-Wing, 10Semantic Search, 07Essential-Work, 06Machine-Learning-Team (Q4 FY2025-26): Transparent DNS Routing for LiftWing Services (eqiad vs Multi-DC) - https://phabricator.wikimedia.org/T422253#12089070 (10DPogorzelski-WMF) @isarantopoulos I think I agree with @elukey here and currently... [08:41:42] 10netops, 06Infrastructure-Foundations: Upgrade netflow hosts to Trixie - https://phabricator.wikimedia.org/T424478#12089148 (10ayounsi) Please keep in mind that if a site has more than one instance with the gnmi class, it will split the targets between them, so it's not possible to have a full "test" host. T... [08:50:20] 06Traffic, 10Lift-Wing, 10Semantic Search, 07Essential-Work, 06Machine-Learning-Team (Q4 FY2025-26): Transparent DNS Routing for LiftWing Services (eqiad vs Multi-DC) - https://phabricator.wikimedia.org/T422253#12089169 (10isarantopoulos) 05Open→03Resolved Sounds good. I'll be resolving this task... [08:52:34] 10netops, 06Infrastructure-Foundations, 06SRE: Blackbox probe for TLS cert expriy failing on multiple eqiad SR-Linux nodes - https://phabricator.wikimedia.org/T429242#12089176 (10ayounsi) I had a quick look at the doc and indeed I couldn't find a way to increase the check interval. [09:30:04] 06Traffic, 06Data-Engineering (Q4 FS25/26 April 1st - June 30st): Add X-Provene - https://phabricator.wikimedia.org/T431256 (10JAllemandou) 03NEW [09:32:33] 06Traffic, 06Data-Engineering (Q1 FS26/27 July 1st - September 30th): Add X-Provenance to `wmf.webrequest` - https://phabricator.wikimedia.org/T431256#12089320 (10JAllemandou) [09:32:46] 06Traffic, 06Data-Engineering (Q1 FS26/27 July 1st - September 30th): Add X-Provenance to `wmf.webrequest` - https://phabricator.wikimedia.org/T431256#12089324 (10JAllemandou) a:05Ahoelzl→03JAllemandou [09:49:37] 06Traffic, 10Liberica, 10Prod-Kubernetes, 06Data-Platform-SRE (2026-07-03 - 2026-07-31), and 3 others: Migrate ML k8s apiserver and services to IPIP - https://phabricator.wikimedia.org/T420438#12089435 (10Gehel) 05Open→03In progress [09:49:38] 06Traffic, 10Liberica, 10Prod-Kubernetes, 06Data-Platform-SRE (2026-07-03 - 2026-07-31), and 3 others: Migrate ML k8s apiserver and services to IPIP - https://phabricator.wikimedia.org/T420438#12089434 (10Gehel) [09:52:45] !! [09:54:42] 06Traffic, 06MediaWiki-Media-Platform-Team: Change the webp threshold based on access distribution - https://phabricator.wikimedia.org/T431150#12089448 (10Ladsgroup) a:03Ladsgroup [09:56:24] 06Traffic, 10Lift-Wing, 10Semantic Search, 07Essential-Work, 06Machine-Learning-Team (Q4 FY2025-26): Transparent DNS Routing for LiftWing Services (eqiad vs Multi-DC) - https://phabricator.wikimedia.org/T422253#12089463 (10achou) +1. I suggest documenting this on the [[ https://wikitech.wikimedia.org... [10:22:15] 06Traffic, 07affects-translatewiki.net, 13Patch-For-Review: Translatewiki.net unable to fetch updates from gerrit over https - https://phabricator.wikimedia.org/T430613#12089606 (10Nikerabbit) It isn't resolved. [11:59:17] 06Traffic, 10Lift-Wing, 06Machine-Learning-Team, 10Semantic Search, 07Essential-Work: Transparent DNS Routing for LiftWing Services (eqiad vs Multi-DC) - https://phabricator.wikimedia.org/T422253#12090009 (10isarantopoulos) [12:09:17] 10netops, 06Traffic, 06Data-Persistence, 06Infrastructure-Foundations, 06ServiceOps new: codfw: rack B3 maintenance - https://phabricator.wikimedia.org/T430909#12090057 (10ops-monitoring-bot) Draining ganeti2048.codfw.wmnet of running VMs [12:10:23] 06Traffic, 06Data-Platform-SRE, 10Liberica, 10Prod-Kubernetes, and 3 others: Migrate ML k8s apiserver and services to IPIP - https://phabricator.wikimedia.org/T420438#12090069 (10isarantopoulos) [13:35:17] 06Traffic, 06Data-Platform-SRE, 10Liberica, 10Prod-Kubernetes, and 3 others: Migrate ML k8s apiserver and services to IPIP - https://phabricator.wikimedia.org/T420438#12090498 (10ops-monitoring-bot) Roll-reboot of nodes in ml-serve-codfw cluster started by klausman: * ml-serve-ctrl[2001-2002].codfw.wmnet [13:41:27] hey traffic folks! fyi I intend to merge https://gerrit.wikimedia.org/r/c/operations/puppet/+/1290731 tomorrow morning to start testing Gitlab's config behind CDN before moving the DNS record, let me know if I should abstain [13:45:24] arnaudb: I guess the various settings are on my mind, the ATS proxy.config.http ones [13:45:44] yes! [13:46:06] can we add some reference or understanding of how we arrived at those? [13:51:45] these are guesstimates based on https://phabricator.wikimedia.org/T428903#12014964 and extrapolated frm typical git traffic pattern, I'll amend the comments to reflect that [13:54:40] thanks [13:54:50] happy to look at it then [14:13:33] thanks sukhe fwiw I also posted a recap (https://phabricator.wikimedia.org/T430901#12081230) of what's been done regarding network reset we have in CI on Gerrit since it's been moved behind CDN. Mentioning this because I might try these timers on Gerrit once they are confirmed valid for Gitlab (https://gerrit.wikimedia.org/r/c/operations/puppet/+/1304808) [14:19:33] thanks arnaudb [14:19:51] give us some time to discuss it (tomorrow's Traffic meeting) and we will take the review. does that seem fine? [14:20:00] np! [14:20:02] thanks :-) [14:20:19] (postponing the tests to Wednesday) [14:28:57] 06Traffic, 06collaboration-services, 06Data-Persistence, 06Infrastructure-Foundations, and 2 others: codfw: rack B8 maintenance - https://phabricator.wikimedia.org/T430929#12090702 (10cmooney) p:05Triage→03Medium [15:28:53] 10netops, 06Traffic, 06Data-Persistence, 06Infrastructure-Foundations, 06ServiceOps new: codfw: rack B3 maintenance - https://phabricator.wikimedia.org/T430909#12090952 (10MoritzMuehlenhoff) [15:37:36] 10netops, 06Infrastructure-Foundations, 06SRE, 13Patch-For-Review: Do we need prometheus-ethtool-exporter? - https://phabricator.wikimedia.org/T371375#12090984 (10LSobanski) Is there a plan to also remove prometheus-ethtool-exporter from Trixie hosts provisioned before the change was merged? Collab gets lo... [15:44:45] 10netops, 06Infrastructure-Foundations, 06SRE, 13Patch-For-Review: Do we need prometheus-ethtool-exporter? - https://phabricator.wikimedia.org/T371375#12091022 (10MoritzMuehlenhoff) >>! In T371375#12090984, @LSobanski wrote: > Is there a plan to also remove prometheus-ethtool-exporter from Trixie hosts pro... [15:47:39] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-ulsfo, 06SRE: ULSFO: OOB IPV6 down - https://phabricator.wikimedia.org/T430599#12091030 (10Papaul) ` Hello Papaul, Apologies for the late response. We are currently coordinating with our internal team for further investigation. We will keep yo... [15:48:53] 06Traffic, 06Data-Persistence, 13Patch-For-Review: Move thumbnail caching from upload cluster to text - https://phabricator.wikimedia.org/T427465#12091033 (10Ladsgroup) [15:56:04] 10netops, 06Infrastructure-Foundations, 06SRE, 13Patch-For-Review: Do we need prometheus-ethtool-exporter? - https://phabricator.wikimedia.org/T371375#12091066 (10LSobanski) You are right, my assumption was invalid and this was in fact a Bookworm host. [16:52:21] 10netops, 06Traffic, 10Cloud-Services, 06collaboration-services, and 7 others: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918#12091218 (10Dzahn) [16:53:56] 10netops, 06Traffic, 10Cloud-Services, 06collaboration-services, and 7 others: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918#12091221 (10Dzahn) re: contint2002 - This is not the main CI server, that is currently contint1002, but it is still used to a certain extent. That being sai... [17:03:33] 06Traffic: Add DNS recrods for VMC/BIMI (Verified Mark Certificate) - https://phabricator.wikimedia.org/T431062#12091245 (10nisrael) Thank you all for working on this so quickly! [17:08:29] 06Traffic, 10Lift-Wing, 06Machine-Learning-Team, 10Semantic Search, 07Essential-Work: Transparent DNS Routing for LiftWing Services (eqiad vs Multi-DC) - https://phabricator.wikimedia.org/T422253#12091255 (10dcausse) Thanks for this looking, so I understand that for internal usecases (e.g. qwen3-embe... [17:20:18] 06Traffic: Create PTR records for LVS interfaces - https://phabricator.wikimedia.org/T431324 (10BCornwall) 03NEW [17:20:31] 06Traffic: Create PTR records for LVS interfaces - https://phabricator.wikimedia.org/T431324#12091373 (10BCornwall) p:05Triage→03Low [17:33:12] 06Traffic: Create PTR records for codfw LVS interfaces - https://phabricator.wikimedia.org/T431324#12091414 (10BCornwall) [17:58:59] 06Traffic, 10Lift-Wing, 06Machine-Learning-Team, 10Semantic Search, 07Essential-Work: Transparent DNS Routing for LiftWing Services (eqiad vs Multi-DC) - https://phabricator.wikimedia.org/T422253#12091490 (10isarantopoulos) Hey @dcausse! yes that is correct. [19:01:46] 06Traffic: Create PTR records for codfw LVS interfaces - https://phabricator.wikimedia.org/T431324#12091657 (10BCornwall) 05Open→03Declined Asked netops about whether this can be safely ignored with the upcoming move to liberica. The answer was yes, so I'm going to just decline this. [19:44:38] 10netops, 06Traffic, 06Data-Persistence, 06Infrastructure-Foundations, 06ServiceOps new: codfw: rack B3 maintenance - https://phabricator.wikimedia.org/T430909#12091876 (10Scott_French) Thanks for your patience while I was out last week, @ayounsi. In short, conf2004 (or any conf* host) will require manu... [20:13:05] 06Traffic, 10MediaWiki-File-management, 10MediaWiki-General, 06MW-Interfaces-Team, and 5 others: Allow MediaWiki to make authenticated requests to other instances of MediaWiki - https://phabricator.wikimedia.org/T424768#12092035 (10HCoplin-WMF) p:05High→03Medium Hey there -- thanks for flagging this an... [22:27:09] FIRING: LVSHighRX: Excessive RX traffic on lvs2013:9100 (eno12399np0) - https://bit.ly/wmf-lvsrx - https://grafana.wikimedia.org/d/000000377/host-overview?orgId=1&var-server=lvs2013 - https://alerts.wikimedia.org/?q=alertname%3DLVSHighRX [22:52:49] brett are you still around? We started an OpenSearch snapshot to Thanos-swift and I have a feeling that is what is causing that HighRX [22:55:37] I'm gonna kill the snapshot and try again at 1/10 rate [23:02:09] RESOLVED: LVSHighRX: Excessive RX traffic on lvs2013:9100 (eno12399np0) - https://bit.ly/wmf-lvsrx - https://grafana.wikimedia.org/d/000000377/host-overview?orgId=1&var-server=lvs2013 - https://alerts.wikimedia.org/?q=alertname%3DLVSHighRX [23:05:59] OK,I restarted the snapshot at 1/4 rate, keeping an eye on the graphs [23:21:57] everything looks OK, we were at 5 Gbps before I rate-limited, now it's sticking around 2Gbps [23:39:15] 06Traffic, 06collaboration-services, 06Data-Persistence, 06Infrastructure-Foundations, and 2 others: codfw: rack B8 maintenance - https://phabricator.wikimedia.org/T430929#12092630 (10Eevans) [23:42:26] 10netops, 06Traffic, 10Cloud-Services, 06collaboration-services, and 7 others: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918#12092634 (10Eevans) [23:42:50] 10netops, 06Traffic, 06Data-Persistence, 06Infrastructure-Foundations, 06ServiceOps new: codfw: rack B3 maintenance - https://phabricator.wikimedia.org/T430909#12092635 (10Eevans)