[06:14:31] Could someone do a rolling pybal restart on `lvs1019`/`lvs1020`? After `wdqs1013`’s VLAN move, DNS changed from `10.64.32.105` to `10.64.171.13`, but pybal is still probing the old IP and therefore keeps the host down/not pooled. Readiness checks are passing on the new IP (T430880) [06:14:32] T430880: Migrate WDQS hosts to Bookworm or later - https://phabricator.wikimedia.org/T430880 [06:14:38] No immediate service impact; wdqs1013 is the only affected backend, and the other 10 backends in eqiad/wdqs-main remain healthy [07:28:37] hi everyone! I have a traffic question. I launched two services yesterday in wikikube - https://query-next.wikidata.org and https://query-scholarly-next.wikidata.org. Both are behind a User-Agent filter (reach out to me on slack at Arthur Taylor WMDE for the header). [07:28:42] What I'm finding is that the https://query-scholarly-next.wikidata.org host only works when I route traffic via drmrs. If I route via esams I see 'Unconfigured Domain' [07:28:57] can someone here take a look? Should I file a ticket? [07:54:52] (in fact I checked all the other datacentres - it is only broken in esams) [08:04:22] file as https://phabricator.wikimedia.org/T432932 [08:04:22] 06Traffic, 10Wikidata, 06Wikidata-Omega: Service at query-scholarly-next.wikidata.org not available in esams - https://phabricator.wikimedia.org/T432932 (10ArthurTaylor) 03NEW [08:05:43] 06Traffic, 10Wikidata, 06Wikidata-Omega (The Board): Service at query-scholarly-next.wikidata.org not available in esams - https://phabricator.wikimedia.org/T432932#12148687 (10ArthurTaylor) [09:59:11] 06Traffic, 06MediaWiki-API-Platform-Team, 06ServiceOps new, 10ServiceOps-SharedInfra, 07Epic: Remove api-gateway service traffic - https://phabricator.wikimedia.org/T432445#12149098 (10jijiki) a:03Blake [09:59:52] 06Traffic, 06MediaWiki-API-Platform-Team, 06ServiceOps new, 10ServiceOps-SharedInfra, 07Epic: Remove api-gateway service traffic - https://phabricator.wikimedia.org/T432445#12149106 (10jijiki) 05In progress→03Resolved @Blake removed all the things along with @ssingh. Thank you folks! [10:00:25] 10Domains: Redirect fil.wikipedia.org to tl.wikipedia.org - https://phabricator.wikimedia.org/T432946#12149114 (10Aklapper) 05Open→03Stalled > The Language committee [has been informed of this proposal](https://meta.wikimedia.org/wiki/Talk:Language_committee#Redirecting_fil.wikipedia.org_to_tl.wikipedia.org)... [11:17:45] 06Traffic, 06collaboration-services, 06Data-Persistence, 06Infrastructure-Foundations, and 2 others: codfw: rack B8 maintenance - Thursday July 23rd 14:00 UTC - https://phabricator.wikimedia.org/T430929#12149488 (10cmooney) [11:20:39] 06Traffic, 06collaboration-services, 06Data-Persistence, 06Infrastructure-Foundations, and 2 others: codfw: rack B8 maintenance - Thursday July 23rd 14:00 UTC - https://phabricator.wikimedia.org/T430929#12149498 (10cmooney) [11:52:25] 06Traffic, 06Data-Engineering, 06Data-Persistence: Data Lake - mediarequests - adapt HQL query to include counting mediarequests from webrequest_text - https://phabricator.wikimedia.org/T431475#12149626 (10JAllemandou) I have spoken with @Ladsgroup . For our `mediarequest` and `mediacount` jobs to continue t... [13:05:14] ryankemper: I can take care of the pybal restart now [13:05:50] codders: that is a very odd one, thanks for filing a task [13:05:53] I will look shortly [13:14:48] ryankemper: all done on the restarts [13:31:41] 06Traffic, 10Wikidata, 06Wikidata-Omega (The Board): Service at query-scholarly-next.wikidata.org not available in esams - https://phabricator.wikimedia.org/T432932#12149945 (10ssingh) Hi @ArthurTaylor: thanks for reporting. We have tried to reproduce this in a variety of ways but have failed to do so. Are y... [13:31:55] codders: responded on the task; fwiw, it works for so it's a bit odd you are seeing this [13:32:05] can you let us know (on the task) if there are other users as well? [14:07:48] 10netops, 06cloud-services-team, 06Data-Persistence, 06Infrastructure-Foundations, and 6 others: codfw: rack B7 maintenance - Tuesday July 21st 14:00 UTC - https://phabricator.wikimedia.org/T430928#12150101 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=7f1c1ea0-38e4-46b3-89a9-b606... [14:08:49] ryankemper sukhe FYI I definitely depooled that one prior to reimage `RI=wdqs1013; sudo cumin --force ${RI}* 'depool'; sudo cookbook sre.hosts.reimage --os bookworm ${RI} --move-vlan` [14:08:55] 06Traffic, 06collaboration-services, 06Data-Persistence, 06Infrastructure-Foundations, and 2 others: codfw: rack B8 maintenance - Thursday July 23rd 14:00 UTC - https://phabricator.wikimedia.org/T430929#12150113 (10Eevans) apus-be2004 has been placed in maintenance mode; good to go [14:10:22] 10netops, 06cloud-services-team, 06Data-Persistence, 06Infrastructure-Foundations, and 6 others: codfw: rack B7 maintenance - Tuesday July 21st 14:00 UTC - https://phabricator.wikimedia.org/T430928#12150115 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=e840132a-4d0f-4713-b838-f931... [14:20:03] 06Traffic, 06collaboration-services, 06Data-Persistence, 06Infrastructure-Foundations, and 2 others: codfw: rack B8 maintenance - Thursday July 23rd 14:00 UTC - https://phabricator.wikimedia.org/T430929#12150152 (10ops-monitoring-bot) Completed depooling of db2163 by cmooney@cumin2003: codfw rack B8 depool... [14:20:31] 06Traffic, 06collaboration-services, 06Data-Persistence, 06Infrastructure-Foundations, and 2 others: codfw: rack B8 maintenance - Thursday July 23rd 14:00 UTC - https://phabricator.wikimedia.org/T430929#12150156 (10ops-monitoring-bot) Completed depooling of db2164 by cmooney@cumin2003: codfw rack B8 depool... [14:21:34] 06Traffic, 06collaboration-services, 06Data-Persistence, 06Infrastructure-Foundations, and 2 others: codfw: rack B8 maintenance - Thursday July 23rd 14:00 UTC - https://phabricator.wikimedia.org/T430929#12150158 (10ops-monitoring-bot) Completed depooling of db2189 by cmooney@cumin2003: codfw rack B8 depool... [14:21:49] 06Traffic, 06collaboration-services, 06Data-Persistence, 06Infrastructure-Foundations, and 2 others: codfw: rack B8 maintenance - Thursday July 23rd 14:00 UTC - https://phabricator.wikimedia.org/T430929#12150160 (10ops-monitoring-bot) Completed depooling of db2249 by cmooney@cumin2003: codfw rack B8 depool... [14:22:08] 06Traffic, 06collaboration-services, 06Data-Persistence, 06Infrastructure-Foundations, and 2 others: codfw: rack B8 maintenance - Thursday July 23rd 14:00 UTC - https://phabricator.wikimedia.org/T430929#12150162 (10ops-monitoring-bot) Completed depooling of es2052 by cmooney@cumin2003: codfw rack B8 depool... [14:59:18] 06Traffic, 06collaboration-services, 06Data-Persistence, 06Infrastructure-Foundations, and 2 others: codfw: rack B8 maintenance - Thursday July 23rd 14:00 UTC - https://phabricator.wikimedia.org/T430929#12150336 (10cmooney) Switch is back online on the new JunOS. Everything looks healthy at first glance. [15:02:56] 06Traffic, 06collaboration-services, 06Data-Persistence, 06Infrastructure-Foundations, and 2 others: codfw: rack B8 maintenance - Thursday July 23rd 14:00 UTC - https://phabricator.wikimedia.org/T430929#12150368 (10ops-monitoring-bot) Starting pool of db2163 by cmooney@cumin2003: codfw rack B8 re-pool afte... [15:08:46] 06Traffic, 06collaboration-services, 06Data-Persistence, 06Infrastructure-Foundations, and 2 others: codfw: rack B8 maintenance - Thursday July 23rd 14:00 UTC - https://phabricator.wikimedia.org/T430929#12150394 (10Eevans) apus-be2004 has been removed from maintenance. `lang=sh-session root@apus-be2005:/#... [15:15:51] 06Traffic, 06collaboration-services, 06Data-Persistence, 06Infrastructure-Foundations, and 2 others: codfw: rack B8 maintenance - Thursday July 23rd 14:00 UTC - https://phabricator.wikimedia.org/T430929#12150424 (10ssingh) [15:30:19] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-eqiad, and 2 others: eqiad row A/B switch refresh prep - https://phabricator.wikimedia.org/T418012#12150480 (10cmooney) I've added the main links for the leaf<->spine single-mode fibre connections in Netbox now, with status "planned": https://netbox.... [15:34:44] inflatador: sorry, there was other stuff happening [15:35:03] so we did depool before the reimaging and then repooled after the reimaging? [15:48:05] 06Traffic, 06collaboration-services, 06Data-Persistence, 06Infrastructure-Foundations, and 2 others: codfw: rack B8 maintenance - Thursday July 23rd 14:00 UTC - https://phabricator.wikimedia.org/T430929#12150605 (10ops-monitoring-bot) Completed pooling of db2163 by cmooney@cumin2003: codfw rack B8 re-pool... [15:48:16] 06Traffic, 06collaboration-services, 06Data-Persistence, 06Infrastructure-Foundations, and 2 others: codfw: rack B8 maintenance - Thursday July 23rd 14:00 UTC - https://phabricator.wikimedia.org/T430929#12150607 (10ops-monitoring-bot) Starting pool of db2164 by cmooney@cumin2003: codfw rack B8 re-pool afte... [16:33:24] 06Traffic, 06collaboration-services, 06Data-Persistence, 06Infrastructure-Foundations, and 2 others: codfw: rack B8 maintenance - Thursday July 23rd 14:00 UTC - https://phabricator.wikimedia.org/T430929#12150836 (10ops-monitoring-bot) Completed pooling of db2164 by cmooney@cumin2003: codfw rack B8 re-pool... [16:33:37] 06Traffic, 06collaboration-services, 06Data-Persistence, 06Infrastructure-Foundations, and 2 others: codfw: rack B8 maintenance - Thursday July 23rd 14:00 UTC - https://phabricator.wikimedia.org/T430929#12150837 (10ops-monitoring-bot) Starting pool of db2189 by cmooney@cumin2003: codfw rack B8 re-pool afte... [17:18:47] 06Traffic, 06collaboration-services, 06Data-Persistence, 06Infrastructure-Foundations, and 2 others: codfw: rack B8 maintenance - Thursday July 23rd 14:00 UTC - https://phabricator.wikimedia.org/T430929#12150970 (10ops-monitoring-bot) Completed pooling of db2189 by cmooney@cumin2003: codfw rack B8 re-pool... [17:18:58] 06Traffic, 06collaboration-services, 06Data-Persistence, 06Infrastructure-Foundations, and 2 others: codfw: rack B8 maintenance - Thursday July 23rd 14:00 UTC - https://phabricator.wikimedia.org/T430929#12150971 (10ops-monitoring-bot) Starting pool of es2052 by cmooney@cumin2003: codfw rack B8 re-pool afte... [17:31:50] 06Traffic, 06DC-Ops, 10ops-drmrs: hw troubleshooting: Memory failure for cp6008.drmrs.wmnet - https://phabricator.wikimedia.org/T431651#12151013 (10RobH) I'll check with the folks in both eqiad and codfw to see if they have any spare '32GB RDIMM, 3200MT/s, dual Rank 8Gb BASE x4' from decoms. [18:04:13] 06Traffic, 06collaboration-services, 06Data-Persistence, 06Infrastructure-Foundations, and 2 others: codfw: rack B8 maintenance - Thursday July 23rd 14:00 UTC - https://phabricator.wikimedia.org/T430929#12151096 (10ops-monitoring-bot) Completed pooling of es2052 by cmooney@cumin2003: codfw rack B8 re-pool... [18:28:38] 06Traffic, 10DNS, 06Fundraising-Backlog, 06Infrastructure-Foundations, and 2 others: Add support for Brand Indicators for Message Identification (BIMI) for wiki mail - https://phabricator.wikimedia.org/T311685#12151174 (10ssingh) @nisrael and I discussed this today, since Noah has received the PEM and SVG... [18:42:44] 06Traffic: Ephemeral port range clashes with bound daemon ports on cache proxy hosts - https://phabricator.wikimedia.org/T433005 (10BCornwall) 03NEW [18:42:56] 06Traffic: Ephemeral port range clashes with bound daemon ports on cache proxy hosts - https://phabricator.wikimedia.org/T433005#12151230 (10BCornwall) [18:43:20] 06Traffic: Ephemeral port range clashes with bound daemon ports on cache proxy hosts - https://phabricator.wikimedia.org/T433005#12151233 (10BCornwall) p:05Triage→03Low [20:26:40] FIRING: [10x] VarnishHighThreadCount: cp1110:0's Varnish thread count nearly maxed out - https://wikitech.wikimedia.org/wiki/Varnish - https://alerts.wikimedia.org/?q=alertname%3DVarnishHighThreadCount [20:31:40] FIRING: [16x] VarnishHighThreadCount: cp1110:0's Varnish thread count nearly maxed out - https://wikitech.wikimedia.org/wiki/Varnish - https://alerts.wikimedia.org/?q=alertname%3DVarnishHighThreadCount [20:36:40] FIRING: [16x] VarnishHighThreadCount: cp1110:0's Varnish thread count nearly maxed out - https://wikitech.wikimedia.org/wiki/Varnish - https://alerts.wikimedia.org/?q=alertname%3DVarnishHighThreadCount [20:41:40] RESOLVED: [16x] VarnishHighThreadCount: cp1110:0's Varnish thread count nearly maxed out - https://wikitech.wikimedia.org/wiki/Varnish - https://alerts.wikimedia.org/?q=alertname%3DVarnishHighThreadCount [23:27:23] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: EQSIN:Switch refresh diagram and wiring - https://phabricator.wikimedia.org/T423724#12152035 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=af46b3e8-aed1-4d85-ad29-e512cf16f00c) set by pt1979@cumin1003 for 2:00:00...