[01:37:20] FIRING: DnsboxServiceMismatch: Service ntp-a state mismatch on dns5003:9100 - https://wikitech.wikimedia.org/wiki/DNS#DnsboxServiceMismatch - https://grafana.wikimedia.org/d/96fb573c-0f3c-456a-886c-e50c29f3ed48/dns-box-service-state?var-site=eqsin&var-instance=dns5003:9100 - https://alerts.wikimedia.org/?q=alertname%3DDnsboxServiceMismatch [01:47:20] RESOLVED: DnsboxServiceMismatch: Service ntp-a state mismatch on dns5003:9100 - https://wikitech.wikimedia.org/wiki/DNS#DnsboxServiceMismatch - https://grafana.wikimedia.org/d/96fb573c-0f3c-456a-886c-e50c29f3ed48/dns-box-service-state?var-site=eqsin&var-instance=dns5003:9100 - https://alerts.wikimedia.org/?q=alertname%3DDnsboxServiceMismatch [01:51:45] 06Traffic, 06MediaWiki-Media-Platform-Team, 07Wikimedia-Performance-recommendation: Change the webp threshold based on access distribution - https://phabricator.wikimedia.org/T431150#12113186 (10Samwilson) This sounds good to me. [07:00:15] 06Traffic, 06cloud-services-team, 10Infrastructure Security, 06Infrastructure-Foundations, and 5 others: July 2026 Bookworm/Trixie reboots: Data Platform SRE - https://phabricator.wikimedia.org/T431826#12113401 (10Jelto) [07:56:29] 06Traffic, 06SRE, 10SRE-SLO: Page on ATS backend errors relative to traffic - https://phabricator.wikimedia.org/T400675#12113544 (10fgiunchedi) This is still an issue: SRE got paged for 10 5xx req/s on swift which was doing 2k req/s at the time (esams). I very much doubt it is worth paging engineers on absol... [08:02:20] 06Traffic, 06Data-Persistence, 13Patch-For-Review: Move thumbnail caching from upload cluster to text - https://phabricator.wikimedia.org/T427465#12113554 (10JAllemandou) Hi folks, this change will impact some of our data jobs. Can you share here a bit more on the planned schedule? [10:17:21] 06Traffic, 06Data-Persistence, 13Patch-For-Review: Move thumbnail caching from upload cluster to text - https://phabricator.wikimedia.org/T427465#12113948 (10Ladsgroup) Hi, The idea is to move 0.1% of URLs in the next week or after (which could roughly mean 0.1% of requests but if we get unlucky, it could me... [11:13:19] 10netops, 06Traffic, 10Cloud-Services, 06collaboration-services, and 7 others: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918#12114149 (10hnowlan) graphite2004 doesn't need depooling. centrallog will have some gaps in logs on-disk but that's unavoidable and low-impact. [13:53:55] 10netops, 06Traffic, 06Infrastructure-Foundations, 06cloud-services-team (FY2025/2026-Q3-Q4): LOAD_BALANCER_HEALTH_CHECKS firewall set and asymmetry lvs1018 / lvs1020 - https://phabricator.wikimedia.org/T430651#12114971 (10ssingh) >>! In T430651#12109242, @cmooney wrote: > @fabfur have you any insight into... [13:57:23] hello traffic friends - any concerns if I roll out (via rolling run-puppet-agent) an ATS backend mapping config change shortly? we need to revert a change that recently removed a Lua plugin from one of the map entries [13:58:00] swfrench-wmf: no concerns at all, thanks for checking [13:58:29] awesome, thanks sukhe! I'll do the usual disable puppet -> rolling run agent thing [14:05:23] 06Traffic, 13Patch-For-Review: Upgrade Traffic hosts to trixie - https://phabricator.wikimedia.org/T401832#12114990 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by cdobbins@cumin2002 for host dns7002.wikimedia.org with OS trixie [14:05:49] swfrench-wmf: <3 [14:37:11] Hi! I'd like to update PyBal config on high-traffic2 servers, https://gerrit.wikimedia.org/r/1310117. Is it a good time? [14:40:26] 10netops, 06Traffic, 10Cloud-Services, 06collaboration-services, and 7 others: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918#12115161 (10cmooney) [14:41:55] atsukoito: looking [14:43:10] atsukoito: as long as you have verified the url, looks good [14:43:29] you can merge whenever you want, or let us know and we can do it for you [14:43:53] (this requires a restart only in eqiad, so lvs1020 first and then lvs1018, high-traffic2) [14:44:11] for posterity, the mapping is in modules/profile/manifests/lvs/configuration.pp [14:44:22] 'high-traffic2' => $::realm ? { [14:44:22] 'production' => $::site ? { [14:44:22] 'eqiad' => [ 'lvs1018', 'lvs1020' ], [14:51:08] thanks sukhe [14:55:39] 10netops, 06Infrastructure-Foundations, 07Sustainability (Incident Followup), 07Wikimedia-Incident: Review thresholds for paging on DDoS alerts - https://phabricator.wikimedia.org/T431683#12115229 (10ayounsi) p:05Triage→03Medium [15:20:21] 10netops, 06Traffic, 06Infrastructure-Foundations, 06tools-infrastructure-team: LOAD_BALANCER_HEALTH_CHECKS firewall set and asymmetry lvs1018 / lvs1020 - https://phabricator.wikimedia.org/T430651#12115415 (10fnegri) [15:20:51] sukhe: I was about to start but then noticed this one in the log [15:21:00] On lvs1020 [15:21:31] 10netops, 10Cloud-VPS, 06Infrastructure-Foundations, 06tools-infrastructure-team: Establish a blackbox network probe vantage point into cloud realm - https://phabricator.wikimedia.org/T429451#12115433 (10fnegri) [15:21:45] https://www.irccloud.com/pastebin/X9qLuO3b/ [15:23:58] 10netops, 10Cloud-VPS, 06Infrastructure-Foundations, 06tools-infrastructure-team: cloud: edge network suffers downtime if one cloudsw is down - https://phabricator.wikimedia.org/T375259#12115460 (10fnegri) [15:24:16] 10netops, 06Traffic, 10Cloud-VPS, 06Data-Platform-SRE, and 4 others: Plan to make clouddumps more resilient and easier to operate - https://phabricator.wikimedia.org/T411248#12115461 (10fnegri) [15:26:12] atsukoito: yeah it is interesting [15:26:48] but what hosts are these though? [15:26:54] for the tcpdump [15:28:15] so cirrussearch1122:9200 should be checked by lvs1020, but the only traffic I see is prometheus [15:29:07] specifically I guess, why is only that failing and no other host in search [15:29:13] cirrussearch1122.eqiad.wmnet: enabled/down/not pooled [15:29:20] everything else looks healthy hmm [15:29:26] is there something specific about 1122? [15:29:30] Maybe we never enabled that host in confctl? [15:29:45] inflatador: it shows pooled there though [15:29:45] sukhe@puppetserver1001:~$ sudo confctl select 'name=cirrussearch1122.eqiad.wmnet' get [15:29:50] {"cirrussearch1122.eqiad.wmnet": {"weight": 10, "pooled": "yes"}, "tags": "dc=eqiad,cluster=elasticsearch,service=elasticsearch"} [15:29:53] {"cirrussearch1122.eqiad.wmnet": {"weight": 10, "pooled": "yes"}, "tags": "dc=eqiad,cluster=elasticsearch,service=elasticsearch-psi-ssl"} [15:29:56] {"cirrussearch1122.eqiad.wmnet": {"weight": 10, "pooled": "yes"}, "tags": "dc=eqiad,cluster=elasticsearch,service=elasticsearch-ssl"} [15:32:32] it was one of the hosts that was migrated to a new vlan, but there are no problems wit other hosts [15:32:58] no OS-level difference as well? or in the hiera override config? [15:33:24] Shouldn't be, and the host is responding from lvs1020 with `atsuko@lvs1020:~$ curl -4 http://cirrussearch1122.eqiad.wmnet:9200` [15:33:27] sukhe oops, it's not that then. Is the problem is that it's not failing out of rotation when you deliberately block the ports with firewall rules? [15:35:19] inflatador: atsukoito: did the host IP change recently? [15:36:01] yes, we did update the IP on 8th of July, https://phabricator.wikimedia.org/T431311#12101743 [15:36:09] ok [15:36:19] restarting pybal [15:36:37] yeah it's happy again [15:37:02] thanks! [15:38:11] so that happens when the host IP changes and I am assuming another step goes missing, perhaps an explicit depool and then pool [15:38:16] anyway, you should be good to go ahead [15:43:14] thanks, merging the healthcheck [15:46:03] 10netops, 06Traffic, 10Cloud-Services, 06collaboration-services, and 7 others: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918#12115582 (10CWilliams-WMF) [15:50:34] salrestarting pybal on lvs1020 [15:52:44] `ip addr` and `ipvsadm -L -n` doesn't have any changes :) [15:53:39] atsukoito: `curl localhost:9090/pools/search_9200` looks good [15:55:30] it is only this one this time `curl localhost:9090/pools/cloudelasticlb_9643` [15:55:44] I'll make a diff for the rest after this one [15:55:46] ah right, I was still stuck at cirrussearch [15:56:11] yeah looks good too. +1 for moving to lvs1018 [15:59:58] restarting pybal on lvs1018 [16:03:01] it was much faster, and `ipvsadm -L -n` output has changed, but `ip addr` is the same [16:03:18] which one was primary, again? [16:03:28] yeah that's expected on lvs1018, since it only hands ht2 [16:03:42] lvs1020 (backup) handles ht1, ht2, low-traffic so has more pools to take care about [16:03:45] atsukoito: lvs1018 [16:03:55] i see, thanks! i was a little bit worries [16:03:58] *worried [16:29:47] sukhe: I'll push the change for other services tomorrow, https://gerrit.wikimedia.org/r/1310129 [16:29:56] thank you for the help [16:30:46] atsukoito: sounds good, let us know if you need a review for that [16:38:25] FIRING: [2x] SystemdUnitFailed: anycast-healthchecker.service on dns7002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [16:39:48] FIRING: PuppetFailure: Puppet has failed on dns7002:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [16:39:53] sukhe: what is other lvs for `search_9200`(`low-traffic` class in eqiad)? it seems that it has the same healthcheck problem (based on prom: https://w.wiki/_bjCc) [16:40:00] FIRING: AnycastHealthcheckerRestarted: anycast-healthchecker service restarted on dns7002:9100 - https://wikitech.wikimedia.org/wiki/Anycast#Anycast_healthchecker_not_running - https://grafana.wikimedia.org/d/dxbfeGDZk/anycast?orgId=1&var-protocol=BGP&var-site=magru&var-cluster=All&var-ip_version=All - https://alerts.wikimedia.org/?q=alertname%3DAnycastHealthcheckerRestarted [16:40:07] cjd91: ^ let's downtime dns7002 [16:40:32] atsukoito: that would be lvs1019 [16:40:43] 'eqiad' => [ 'lvs1019', 'lvs1020' ], [16:40:47] 'lvs1019' => 'low-traffic', [16:41:11] or, cumin alias, A:lvs-low-traffic-eqiad [16:41:16] and similary, A:lvs-low-traffic-codfw [16:41:23] these only exist in eqiad/codfw for now [16:41:25] Thanks! shall I restart pybal myself? [16:41:38] atsukoito: yeah or you can ask us, whatever works [16:41:45] but same process, first on the backup, lvs1020, see if everything works [16:41:51] then do the primary low-traffic, lvs1019 [16:42:30] ah, I didn't do any config change, it is that old problem with `cirrussearch1122.eqiad.wmnet` (old ip is stuck in pybal) [16:42:51] ah yeah, so same thing there [16:42:59] just lvs1018 in that case [16:43:07] (no need for lvs1020 since we did that alreadY) [16:43:47] `Jul 13 16:40:59 lvs1019 pybal[2107970]: [search-https_9243 ProxyFetch] WARN: cirrussearch1122.eqiad.wmnet (enabled/down/not pooled): Fetch failed (https://localhost/), 3.057 s` [16:44:15] `lvs1019` tho, sorry to be confusing with two classes and two separate issues [16:44:30] yes that, you are right :) [16:44:32] 12:41:51 < sukhe> then do the primary low-traffic, lvs1019 [16:44:41] 12:42:59 < sukhe> just lvs1018 in that case <-- wrong [16:45:09] I was a bit surprised on that note to see cirrussearch on ht2. that surprised me, but I am sure there is a reason [16:47:18] restarted, `cirrussearch1122.eqiad.wmnet: enabled/up/pooled` [16:47:29] no changes in `ip addr` [16:47:42] cool, thanks :) [16:50:11] thanks for help, signing off for today [16:50:38] enjoy! [17:53:25] FIRING: [2x] SystemdUnitFailed: anycast-healthchecker.service on dns7002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [17:58:25] FIRING: [2x] SystemdUnitFailed: anycast-healthchecker.service on dns7002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [18:01:18] 06Traffic: several HTTP 503 and 504 errors - https://phabricator.wikimedia.org/T413622#12116329 (10Aklapper) 05Stalled→03Invalid Unfortunately closing this Phabricator task as no further information has been provided. @doctaxon: After you have provided the information asked for and if this still happens... [18:03:25] RESOLVED: [2x] SystemdUnitFailed: anycast-healthchecker.service on dns7002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [18:05:00] RESOLVED: AnycastHealthcheckerRestarted: anycast-healthchecker service restarted on dns7002:9100 - https://wikitech.wikimedia.org/wiki/Anycast#Anycast_healthchecker_not_running - https://grafana.wikimedia.org/d/dxbfeGDZk/anycast?orgId=1&var-protocol=BGP&var-site=magru&var-cluster=All&var-ip_version=All - https://alerts.wikimedia.org/?q=alertname%3DAnycastHealthcheckerRestarted [18:20:07] 06Traffic: Upgrade Traffic hosts to trixie - https://phabricator.wikimedia.org/T401832#12116422 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by cdobbins@cumin2002 for host dns7002.wikimedia.org with OS trixie completed: - dns7002 (**WARN**) - Downtimed on Icinga/Alertmanager - Disabled... [18:24:48] RESOLVED: PuppetFailure: Puppet has failed on dns7002:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [18:28:36] 06Traffic, 06SRE, 10SRE-SLO: Page on ATS backend errors relative to traffic - https://phabricator.wikimedia.org/T400675#12116461 (10ssingh) Thanks for the update @fgiunchedi. We should certainly move forward on this, one way or the other. @hnowlan: Would appreciate some input from you, or someone on olly on... [18:46:41] 06Traffic, 06SRE, 10SRE-SLO: Page on ATS backend errors relative to traffic - https://phabricator.wikimedia.org/T400675#12116524 (10ssingh) We will be discussing this tomorrow in the Traffic meeting as well and will follow up. [18:58:24] 10netops, 06Traffic, 10Cloud-Services, 06collaboration-services, and 7 others: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918#12116562 (10BTullis) I have cordoned `dse-k8s-worker2003` ` root@deploy2003:~# kubectl cordon dse-k8s-worker2003.codfw.wmnet node/dse-k8s-worker2003.codfw.wm...