[00:15:31] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host cirrussearch1110.eqiad.wmnet with OS trixie [00:26:58] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host cirrussearch1099.eqiad.wmnet with OS trixie [00:32:50] !log bking@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage [00:37:39] (03CR) 10Ladsgroup: [C:03+1] "Hi @fnegri@wikimedia.org. If you have time, could you take care of it? Thanks. If not, I'll try to deploy it on Thursday." [puppet] - 10https://gerrit.wikimedia.org/r/1307100 (https://phabricator.wikimedia.org/T386456) (owner: 10Dreamy Jazz) [00:38:06] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage [00:41:53] !log bking@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage [00:44:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.04% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [00:44:35] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage [00:49:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.04% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [00:53:43] (03CR) 10Dzahn: [V:03+1 C:03+1] "thank you :)" [puppet] - 10https://gerrit.wikimedia.org/r/1215329 (owner: 10Dzahn) [00:57:59] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1110.eqiad.wmnet with OS trixie [01:03:35] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1099.eqiad.wmnet with OS trixie [01:12:45] (03PS1) 10TrainBranchBot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1308244 [01:12:45] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1308244 (owner: 10TrainBranchBot) [01:14:00] (03PS1) 10Dzahn: ci: add contint2003 to list of hosts allowed to git/9148 [puppet] - 10https://gerrit.wikimedia.org/r/1308245 (https://phabricator.wikimedia.org/T418521) [01:21:04] (03CR) 10CI reject: [V:04-1] Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1308244 (owner: 10TrainBranchBot) [01:41:25] (03PS1) 10Dzahn: ci: add parameter to toggle jenkins package install, remove it from role::ci [puppet] - 10https://gerrit.wikimedia.org/r/1308249 (https://phabricator.wikimedia.org/T418521) [01:45:02] (03PS1) 10Ladsgroup: Init minwikiquote [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308250 (https://phabricator.wikimedia.org/T429922) [01:48:01] (03CR) 10TrainBranchBot: [C:03+2] "Approved by ladsgroup@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308250 (https://phabricator.wikimedia.org/T429922) (owner: 10Ladsgroup) [01:48:58] (03Merged) 10jenkins-bot: Init minwikiquote [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308250 (https://phabricator.wikimedia.org/T429922) (owner: 10Ladsgroup) [01:49:36] !log ladsgroup@deploy1003 Started scap sync-world: Backport for [[gerrit:1308250|Init minwikiquote (T429922)]] [01:49:39] T429922: Create Wikiquote Minangkabau - https://phabricator.wikimedia.org/T429922 [01:51:38] !log ladsgroup@deploy1003 ladsgroup: Backport for [[gerrit:1308250|Init minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [01:55:02] !log ladsgroup@deploy1003 ladsgroup: Continuing with deployment [01:59:22] !log ladsgroup@deploy1003 Finished scap sync-world: Backport for [[gerrit:1308250|Init minwikiquote (T429922)]] (duration: 09m 46s) [01:59:26] T429922: Create Wikiquote Minangkabau - https://phabricator.wikimedia.org/T429922 [02:07:20] (03PS1) 10Ladsgroup: Activate minwikiquote [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308251 (https://phabricator.wikimedia.org/T429922) [02:09:27] FIRING: [2x] JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:14:22] RESOLVED: [2x] JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:17:37] (03CR) 10TrainBranchBot: [C:03+2] "Approved by ladsgroup@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308251 (https://phabricator.wikimedia.org/T429922) (owner: 10Ladsgroup) [02:18:39] (03Merged) 10jenkins-bot: Activate minwikiquote [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308251 (https://phabricator.wikimedia.org/T429922) (owner: 10Ladsgroup) [02:19:01] !log ladsgroup@deploy1003 Started scap sync-world: Backport for [[gerrit:1308251|Activate minwikiquote (T429922)]] [02:19:05] T429922: Create Wikiquote Minangkabau - https://phabricator.wikimedia.org/T429922 [02:21:03] !log ladsgroup@deploy1003 ladsgroup: Backport for [[gerrit:1308251|Activate minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [02:22:55] !log ladsgroup@deploy1003 ladsgroup: Continuing with deployment [02:27:15] !log ladsgroup@deploy1003 Finished scap sync-world: Backport for [[gerrit:1308251|Activate minwikiquote (T429922)]] (duration: 08m 14s) [02:27:19] T429922: Create Wikiquote Minangkabau - https://phabricator.wikimedia.org/T429922 [02:55:25] FIRING: CoreRouterInterfaceDown: Core router interface down - cr2-eqiad:et-1/1/0 () - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [02:57:20] FIRING: [2x] CirrusSearchBackendMemoryIssue: CirrusSearch backend failed 1434.4383966666664 times in the last 10 minutes due to memory usage (mw@codfw to dnsdisc) - https://wikitech.wikimedia.org/wiki/Search#Health/Activity_Monitoring - https://grafana.wikimedia.org/d/dc04b9f2-b8d5-4ab6-9482-5d9a75728951/elasticsearch-percentiles?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchBackendMemoryIssue [03:07:21] RESOLVED: [2x] CirrusSearchBackendMemoryIssue: CirrusSearch backend failed 338.88888888888886 times in the last 10 minutes due to memory usage (mw@codfw to dnsdisc) - https://wikitech.wikimedia.org/wiki/Search#Health/Activity_Monitoring - https://grafana.wikimedia.org/d/dc04b9f2-b8d5-4ab6-9482-5d9a75728951/elasticsearch-percentiles?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchBackendMemoryIssue [03:07:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [03:18:21] FIRING: CirrusSearchBackendMemoryIssue: CirrusSearch backend failed 1513.333333333333 times in the last 10 minutes due to memory usage (mw@eqiad to dnsdisc) - https://wikitech.wikimedia.org/wiki/Search#Health/Activity_Monitoring - https://grafana.wikimedia.org/d/dc04b9f2-b8d5-4ab6-9482-5d9a75728951/elasticsearch-percentiles?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchBackendMemoryIssue [03:23:20] FIRING: [2x] CirrusSearchBackendMemoryIssue: CirrusSearch backend failed 533.3333333333334 times in the last 10 minutes due to memory usage (mw@codfw to dnsdisc) - https://wikitech.wikimedia.org/wiki/Search#Health/Activity_Monitoring - https://grafana.wikimedia.org/d/dc04b9f2-b8d5-4ab6-9482-5d9a75728951/elasticsearch-percentiles?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchBackendMemoryIssue [03:28:20] RESOLVED: [2x] CirrusSearchBackendMemoryIssue: CirrusSearch backend failed 533.3333333333334 times in the last 10 minutes due to memory usage (mw@codfw to dnsdisc) - https://wikitech.wikimedia.org/wiki/Search#Health/Activity_Monitoring - https://grafana.wikimedia.org/d/dc04b9f2-b8d5-4ab6-9482-5d9a75728951/elasticsearch-percentiles?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchBackendMemoryIssue [03:56:02] (03CR) 10Giuseppe Lavagetto: [C:03+1] cache::text: raise global_auth_ratelimit limits [puppet] - 10https://gerrit.wikimedia.org/r/1307220 (owner: 10BCornwall) [04:47:50] (03CR) 10Samwilson: [C:03+1] "Looks ok to me, but I don't know much about how all this works yet." [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1308156 (https://phabricator.wikimedia.org/T333445) (owner: 10Ladsgroup) [04:50:33] (03CR) 10Arnaudb: [C:03+2] gerrit: filter paths on httpd [puppet] - 10https://gerrit.wikimedia.org/r/1308148 (https://phabricator.wikimedia.org/T431398) (owner: 10Arnaudb) [04:55:49] !log ryankemper@cumin2002 START - Cookbook sre.hosts.reimage for host cirrussearch1101.eqiad.wmnet with OS trixie [04:56:28] !log ryankemper@cumin2002 START - Cookbook sre.hosts.reimage for host cirrussearch1107.eqiad.wmnet with OS trixie [04:58:28] !log ryankemper@cumin2002 START - Cookbook sre.hosts.reimage for host cirrussearch1124.eqiad.wmnet with OS trixie [05:10:54] !log ryankemper@cumin2002 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage [05:11:12] !log ryankemper@cumin2002 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage [05:13:35] !log ryankemper@cumin2002 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage [05:13:41] !log ryankemper@cumin2002 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage [05:17:37] !log ryankemper@cumin2002 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage [05:21:40] !log ryankemper@cumin2002 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage [05:31:24] !log ryankemper@cumin2002 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1101.eqiad.wmnet with OS trixie [05:32:13] (03PS4) 10Ladsgroup: swift: Load the file async [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1308156 (https://phabricator.wikimedia.org/T333445) [05:33:30] (03CR) 10Ladsgroup: "Oh I don't know how it works either :P" [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1308156 (https://phabricator.wikimedia.org/T333445) (owner: 10Ladsgroup) [05:35:51] !log ryankemper@cumin2002 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1124.eqiad.wmnet with OS trixie [05:42:14] !log ryankemper@cumin2002 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1107.eqiad.wmnet with OS trixie [06:00:05] Deploy window MediaWiki infrastructure (UTC early) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260708T0600) [06:03:27] (03CR) 10Samwilson: [C:03+1] swift: Load the file async (031 comment) [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1308156 (https://phabricator.wikimedia.org/T333445) (owner: 10Ladsgroup) [06:06:25] FIRING: SystemdUnitFailed: dump_proxy_ranges.service on puppetserver2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [06:06:29] (03CR) 10Kevin Bazira: [C:03+1] ml-services: Remove experimental qwen3-14b deployment. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308085 (https://phabricator.wikimedia.org/T431268) (owner: 10Bartosz Wójtowicz) [06:31:32] (03PS21) 10Elukey: Add sre.hosts.bmc-user-mgmt.py [cookbooks] - 10https://gerrit.wikimedia.org/r/1302859 (https://phabricator.wikimedia.org/T426180) [06:33:48] PROBLEM - MariaDB disk space on db1208 is CRITICAL: DISK CRITICAL - /srv is not accessible: Input/output error https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting [06:33:54] PROBLEM - MariaDB read only matomo on db1208 is CRITICAL: Could not connect to localhost:3351 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Master_comes_back_in_read_only [06:33:56] PROBLEM - MariaDB read only analytics_meta on db1208 is CRITICAL: Could not connect to localhost:3352 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Master_comes_back_in_read_only [06:36:46] PROBLEM - mysqld processes on db1208 is CRITICAL: PROCS CRITICAL: 0 processes with command name mysqld https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting [06:36:46] PROBLEM - MariaDB Replica IO: analytics_meta on db1208 is CRITICAL: CRITICAL slave_io_state could not connect https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [06:36:46] PROBLEM - MariaDB Replica IO: matomo on db1208 is CRITICAL: CRITICAL slave_io_state could not connect https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [06:36:46] PROBLEM - MariaDB Replica SQL: matomo on db1208 is CRITICAL: CRITICAL slave_sql_state could not connect https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [06:36:48] PROBLEM - MariaDB Replica SQL: analytics_meta on db1208 is CRITICAL: CRITICAL slave_sql_state could not connect https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [06:38:01] !log ryankemper@cumin2002 START - Cookbook sre.hosts.reimage for host cirrussearch1125.eqiad.wmnet with OS trixie [06:40:34] !log elukey@cumin1003 START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet [06:40:41] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet [06:43:48] PROBLEM - MariaDB Replica Lag: matomo on db1208 is CRITICAL: CRITICAL slave_sql_lag could not connect https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [06:43:48] PROBLEM - MariaDB Replica Lag: analytics_meta on db1208 is CRITICAL: CRITICAL slave_sql_lag could not connect https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [06:49:11] (03CR) 10Elukey: [C:03+2] Add sre.hosts.bmc-user-mgmt.py (032 comments) [cookbooks] - 10https://gerrit.wikimedia.org/r/1302859 (https://phabricator.wikimedia.org/T426180) (owner: 10Elukey) [06:50:17] !log ryankemper@cumin2002 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage [06:52:49] !log upgrade all trixie hosts to pywmflib 3.1 - T430552 [06:52:51] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [06:52:52] T430552: Deploy wmflib 3.0.0 to production - https://phabricator.wikimedia.org/T430552 [06:53:10] 06SRE, 10SRE-tools, 06Infrastructure-Foundations: Deploy wmflib 3.0.0 to production - https://phabricator.wikimedia.org/T430552#12097992 (10elukey) 05Open→03Resolved a:03elukey [06:54:21] !log ryankemper@cumin2002 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage [06:55:25] FIRING: CoreRouterInterfaceDown: Core router interface down - cr2-eqiad:et-1/1/0 () - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [06:56:51] !log elukey@cumin1003 START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm [06:59:27] FIRING: [3x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-1/0/2 (Transport: cr1-eqiad:et-1/1/2 (Arelion, IC-374549) {#12267}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [07:00:05] Amir1, urbanecm, and awight: How many deployers does it take to do UTC morning backport window deploy? (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260708T0700). [07:00:05] VadymTS1: A patch you scheduled for UTC morning backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [07:07:32] 06SRE, 10SRE-Access-Requests, 06Data-Engineering: Personal Hive Database Username - https://phabricator.wikimedia.org/T431412#12098007 (10catherine.kelsey.wmde) Thank you very much (feel quite stupid I missed this). Thanks for taking a look and clarifying, have done this now! [07:07:37] 06SRE, 10SRE-Access-Requests, 06Data-Engineering: Personal Hive Database Username - https://phabricator.wikimedia.org/T431412#12098008 (10catherine.kelsey.wmde) 05Open→03Resolved a:03catherine.kelsey.wmde [07:07:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [07:07:53] !log elukey@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage [07:13:31] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage [07:13:43] !log ryankemper@cumin2002 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1125.eqiad.wmnet with OS trixie [07:14:22] FIRING: [5x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-1/0/2 (Transport: cr1-eqiad:et-1/1/2 (Arelion, IC-374549) {#12267}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [07:17:53] (03PS2) 10Arnaudb: gerrit: drop 132.219.138.0/24 on Gerrit [puppet] - 10https://gerrit.wikimedia.org/r/1308351 (https://phabricator.wikimedia.org/T431398) [07:21:33] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host thanos-fe1005.eqiad.wmnet with OS trixie [07:21:49] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12098045 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host thanos-fe1005.eqiad.wmnet with OS trixie [07:28:28] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.5 point update - https://phabricator.wikimedia.org/T427072#12098054 (10MoritzMuehlenhoff) [07:29:09] !log installing gnutls28 security updates [07:29:10] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [07:36:51] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm [07:38:20] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage [07:39:28] (03PS1) 10Brouberol: dse-k8s: upgrade cloudnative-pg from 1.23 to 1.27 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308354 (https://phabricator.wikimedia.org/T424493) [07:40:06] (03PS1) 10Ryan Kemper: pontoon: add rkemper-kafka stack [puppet] - 10https://gerrit.wikimedia.org/r/1308355 (https://phabricator.wikimedia.org/T276088) [07:42:22] (03PS2) 10Brouberol: dse-k8s: upgrade cloudnative-pg from 1.23 to 1.27 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308354 (https://phabricator.wikimedia.org/T424493) [07:44:34] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage [07:48:51] jouncebot: nowandnext [07:48:52] For the next 0 hour(s) and 11 minute(s): UTC morning backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260708T0700) [07:48:52] In 2 hour(s) and 11 minute(s): MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260708T1000) [07:50:08] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1121 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [07:50:08] PROBLEM - OpenSearch health check for shards on 9200 on cirrussearch1121 is CRITICAL: CRITICAL - elasticsearch http://localhost:9200/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9200): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [07:50:11] andre: so somehow the Deployment calendar had put the train later tonight even though Jaime, you and me are flagged as running it during emea morning [07:50:14] https://wikitech.wikimedia.org/wiki/Deployments#Wednesday,_July_8 [07:50:38] possibly because the entry has task TNONE [07:51:00] so the script was unable to find the metadata and fallback to the default evening window [07:51:02] fun twist [07:54:37] 06SRE, 06Data-Engineering, 06Data-Platform-SRE, 10Kafka-Infrastructure, and 3 others: Configuration Management for Kafka settings - https://phabricator.wikimedia.org/T276088#12098197 (10RKemper) Uploaded pontoon patch and spun up the cluster, seems to be working thus far. Note we're at 2 brokers rather tha... [07:54:56] (03PS3) 10Arnaudb: gerrit: remove unused bad_browser reference [puppet] - 10https://gerrit.wikimedia.org/r/1308351 (https://phabricator.wikimedia.org/T431398) [07:55:23] 06SRE, 06Data-Engineering, 10Kafka-Infrastructure, 06serviceops-radar, and 3 others: Configuration Management for Kafka settings - https://phabricator.wikimedia.org/T276088#12098200 (10RKemper) [07:56:32] !log ayounsi@cumin1003 START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B3 [07:57:24] hashar: uuuuuurgh [07:58:19] (03CR) 10Hashar: "Looks good!" [puppet] - 10https://gerrit.wikimedia.org/r/1308351 (https://phabricator.wikimedia.org/T431398) (owner: 10Arnaudb) [07:59:32] ayounsi@cumin1003 depool-rack (PID 2401892) is awaiting input [08:02:36] andre: I am fixing it [08:02:46] thanks [08:03:02] !log ayounsi@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 21 hosts with reason: codfw rack B3 depool for maintenance [08:03:41] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1005.eqiad.wmnet with OS trixie [08:03:55] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12098239 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host thanos-fe1005.eqiad.wmnet with OS trixie compl... [08:05:08] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host thanos-fe1004.eqiad.wmnet with OS trixie [08:05:26] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12098256 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host thanos-fe1004.eqiad.wmnet with OS trixie [08:05:40] !log ayounsi@cumin1003 START - Cookbook sre.mysql.depool depool es2051: codfw rack B3 depool for maintenance [08:06:00] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: codfw rack B3 depool for maintenance [08:06:36] (03PS1) 10STran: Deploy IRS to enwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308553 (https://phabricator.wikimedia.org/T431316) [08:06:54] (03PS2) 10STran: Deploy IRS to enwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308553 (https://phabricator.wikimedia.org/T431316) [08:07:16] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, July 08 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployc" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308553 (https://phabricator.wikimedia.org/T431316) (owner: 10STran) [08:07:22] !log ayounsi@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet [08:08:55] jouncebot: refresh [08:08:55] I refreshed my knowledge about deployments. [08:08:58] jouncebot: now [08:08:59] For the next 1 hour(s) and 51 minute(s): MediaWiki train - Utc morning Version (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260708T0800) [08:09:07] andre: fixed! [08:09:44] I might have missed some bits but at least the time and blocker task are rights for the whole week [08:10:01] lets proceed it, this time with SpiderPig! [08:10:29] (03PS1) 10TrainBranchBot: group0 to 1.47.0-wmf.10 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308555 (https://phabricator.wikimedia.org/T430829) [08:10:32] (03CR) 10TrainBranchBot: [C:03+2] "Initiated by hashar@deploy1003" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308555 (https://phabricator.wikimedia.org/T430829) (owner: 10TrainBranchBot) [08:10:51] scary :P [08:11:25] (03Merged) 10jenkins-bot: group0 to 1.47.0-wmf.10 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308555 (https://phabricator.wikimedia.org/T430829) (owner: 10TrainBranchBot) [08:12:41] (03CR) 10Mszwarc: [C:03+1] Deploy IRS to enwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308553 (https://phabricator.wikimedia.org/T431316) (owner: 10STran) [08:13:37] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet [08:13:54] andre: I had been running it out of a terminal/screen for a decade, I guess I had some habits :b [08:14:16] I still haven't tried the UI, sigh [08:15:14] !log ayounsi@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet [08:15:14] !log ayounsi@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b3-codfw,lsw1-b3-codfw IPv6,lsw1-b3-codfw.mgmt with reason: Switch maintenance [08:15:51] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet [08:15:51] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B3 [08:16:11] it is quite good [08:16:36] if you want to use the UI but would rather drive it from the CLI, you can always use webdriver.io / cypress Oo, [08:16:52] !log cwilliams@cumin1003 START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 [08:17:26] what I'd like ideally is to be able to prompt it "promote group0" [08:17:42] get the command as a result + confirmation and then I can approve the execution [08:17:44] !log hashar@deploy1003 rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.10 refs T430829 [08:17:47] T430829: 1.47.0-wmf.10 deployment blockers - https://phabricator.wikimedia.org/T430829 [08:18:00] group0 done [08:18:04] (03PS1) 10Ayounsi: configcluster: profile::server_depool fix typo [puppet] - 10https://gerrit.wikimedia.org/r/1308556 [08:18:22] (03CR) 10Bartosz Wójtowicz: [C:03+2] ml-services: Remove experimental qwen3-14b deployment. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308085 (https://phabricator.wikimedia.org/T431268) (owner: 10Bartosz Wójtowicz) [08:18:36] (03CR) 10Muehlenhoff: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1308556 (owner: 10Ayounsi) [08:19:17] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 [08:19:19] !log lsw1-b3-codfw> request system reboot - T430909 [08:19:21] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [08:19:22] T430909: codfw: rack B3 maintenance - https://phabricator.wikimedia.org/T430909 [08:20:49] (03Merged) 10jenkins-bot: ml-services: Remove experimental qwen3-14b deployment. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308085 (https://phabricator.wikimedia.org/T431268) (owner: 10Bartosz Wójtowicz) [08:21:54] (03CR) 10Ayounsi: [C:03+2] configcluster: profile::server_depool fix typo [puppet] - 10https://gerrit.wikimedia.org/r/1308556 (owner: 10Ayounsi) [08:22:09] PROBLEM - Host dse-k8s-etcd2002 is DOWN: PING CRITICAL - Packet loss = 100% [08:22:30] expected (non DRDB VM) [08:22:39] FIRING: CirrusSearchNodeIndexingNotIncreasing: Elasticsearch instance cirrussearch1121-production-search-eqiad is not indexing - https://wikitech.wikimedia.org/wiki/Search/Elasticsearch_Administration#Indexing_hung_and_not_making_progress - https://grafana.wikimedia.org/d/JLK3I_siz/elasticsearch-indexing?orgId=1&from=now-3d&to=now&viewPanel=57 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchNodeIndexingNotIncreasing [08:22:41] !log fceratto@cumin1003 START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s5 [08:22:55] PROBLEM - BFD status on ssw1-a8-codfw.mgmt is CRITICAL: Down: 1 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [08:22:55] PROBLEM - BFD status on ssw1-a1-codfw.mgmt is CRITICAL: Down: 1 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [08:23:03] (03CR) 10Bartosz Wójtowicz: [C:03+2] "@tklausmann@wikimedia.org Could you sync the merged REST Gateway changes in a free second?" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308085 (https://phabricator.wikimedia.org/T431268) (owner: 10Bartosz Wójtowicz) [08:23:09] !log bwojtowicz@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . [08:23:12] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage [08:23:25] FIRING: [4x] ProbeDown: Service restbase2028-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [08:23:29] I'll let it settle a bit then process with group1 [08:23:33] (03Abandoned) 10Nikerabbit: Localisation updates from https://translatewiki.net. [phabricator/translations] (wmf/stable) - 10https://gerrit.wikimedia.org/r/1307779 (owner: 10L10n-bot) [08:23:39] FIRING: CoreBGPDown: Core BGP session down between ssw1-a8-codfw and lsw1-b3-codfw (10.192.252.12) - group EVPN_IBGP - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=ssw1-a8-codfw:9804&var-bgp_group=EVPN_IBGP&var-bgp_neighbor=lsw1-b3-codfw - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [08:23:43] (plus there are a bunch of things failing possibly due to the network operation :) ) [08:24:13] (03CR) 10Atsuko: [C:03+1] dse-k8s: upgrade cloudnative-pg from 1.23 to 1.27 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308354 (https://phabricator.wikimedia.org/T424493) (owner: 10Brouberol) [08:24:25] (03Abandoned) 10Nikerabbit: Localisation updates from https://translatewiki.net. [phabricator/translations] (wmf/stable) - 10https://gerrit.wikimedia.org/r/1304791 (owner: 10L10n-bot) [08:24:25] (03Abandoned) 10Nikerabbit: Localisation updates from https://translatewiki.net. [phabricator/translations] (wmf/stable) - 10https://gerrit.wikimedia.org/r/1307129 (owner: 10L10n-bot) [08:24:26] (03Abandoned) 10Nikerabbit: Localisation updates from https://translatewiki.net. [phabricator/translations] (wmf/stable) - 10https://gerrit.wikimedia.org/r/1305648 (owner: 10L10n-bot) [08:24:26] (03Abandoned) 10Nikerabbit: Localisation updates from https://translatewiki.net. [phabricator/translations] (wmf/stable) - 10https://gerrit.wikimedia.org/r/1302135 (owner: 10L10n-bot) [08:24:41] (03Abandoned) 10Nikerabbit: Localisation updates from https://translatewiki.net. [phabricator/translations] (wmf/stable) - 10https://gerrit.wikimedia.org/r/1306246 (owner: 10L10n-bot) [08:24:51] FIRING: [2x] SwitchCoreInterfaceDown: Switch core interface down - ssw1-a1-codfw:et-0/0/9 (Core: lsw1-b3-codfw:et-0/0/55 {#230403800006}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Switch_interface_down - https://alerts.wikimedia.org/?q=alertname%3DSwitchCoreInterfaceDown [08:27:51] FIRING: SwaggerProbeHasFailures: Not all openapi/swagger endpoints returned healthy - https://wikitech.wikimedia.org/wiki/Runbook#https://wikifeeds.svc.codfw.wmnet:4101 - https://grafana.wikimedia.org/d/_77ik484k/openapi-swagger-endpoint-state?var-site=codfw - https://alerts.wikimedia.org/?q=alertname%3DSwaggerProbeHasFailures [08:27:53] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage [08:28:25] FIRING: [6x] ProbeDown: Service restbase2028-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [08:28:39] FIRING: [2x] CoreBGPDown: Core BGP session down between ssw1-a1-codfw and lsw1-b3-codfw (10.192.252.12) - group EVPN_IBGP - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [08:29:22] FIRING: [5x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-1/0/2 (Transport: cr1-eqiad:et-1/1/2 (Arelion, IC-374549) {#12267}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [08:30:03] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis minwikiquote in section s5 [08:30:52] !log fceratto@cumin1003 START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5 [08:31:25] FIRING: [2x] SystemdUnitFailed: fetch-rings-thanos.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [08:31:34] switch back up, interfaces still initializing [08:32:01] (03CR) 10Brouberol: [C:03+2] dse-k8s: upgrade cloudnative-pg from 1.23 to 1.27 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308354 (https://phabricator.wikimedia.org/T424493) (owner: 10Brouberol) [08:32:33] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5 [08:32:43] RECOVERY - Host dse-k8s-etcd2002 is UP: PING OK - Packet loss = 0%, RTA = 37.52 ms [08:32:51] RESOLVED: SwaggerProbeHasFailures: Not all openapi/swagger endpoints returned healthy - https://wikitech.wikimedia.org/wiki/Runbook#https://wikifeeds.svc.codfw.wmnet:4101 - https://grafana.wikimedia.org/d/_77ik484k/openapi-swagger-endpoint-state?var-site=codfw - https://alerts.wikimedia.org/?q=alertname%3DSwaggerProbeHasFailures [08:32:55] RECOVERY - BFD status on ssw1-a1-codfw.mgmt is OK: UP: 17 AdminDown: 0 Down: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [08:32:55] RECOVERY - BFD status on ssw1-a8-codfw.mgmt is OK: UP: 17 AdminDown: 0 Down: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [08:33:35] switch is full back [08:33:39] FIRING: [2x] CoreBGPDown: Core BGP session down between ssw1-a1-codfw and lsw1-b3-codfw (10.192.252.12) - group EVPN_IBGP - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [08:33:53] !log fceratto@cumin1003 START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s3 [08:34:51] RESOLVED: [2x] SwitchCoreInterfaceDown: Switch core interface down - ssw1-a1-codfw:et-0/0/9 (Core: lsw1-b3-codfw:et-0/0/55 {#230403800006}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Switch_interface_down - https://alerts.wikimedia.org/?q=alertname%3DSwitchCoreInterfaceDown [08:35:17] !log fceratto@cumin1003 END (ERROR) - Cookbook sre.mysql.sanitize-wiki (exit_code=97) Managing sanitization for wikis minwikiquote in section s3 [08:38:25] RESOLVED: [6x] ProbeDown: Service restbase2028-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [08:38:30] !log javiermonton@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply [08:38:36] !log javiermonton@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply [08:38:39] RESOLVED: [2x] CoreBGPDown: Core BGP session down between ssw1-a1-codfw and lsw1-b3-codfw (10.192.252.12) - group EVPN_IBGP - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [08:40:55] (03Merged) 10jenkins-bot: dse-k8s: upgrade cloudnative-pg from 1.23 to 1.27 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308354 (https://phabricator.wikimedia.org/T424493) (owner: 10Brouberol) [08:40:57] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet [08:41:53] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet [08:42:27] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet [08:42:58] andre: logs look good in the OpenSearch dashboard and on https://grafana.wikimedia.org/d/000000102/mediawiki-production-logging [08:43:12] yeah I also took a quick look :) [08:43:17] !log ayounsi@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet [08:43:19] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet [08:43:19] XioNoX: thanks :] I am going to process with the rest of the mw train [08:44:12] though hmm there are a few PHP Warning: API call failed to get remote JsonConfig: status=
    [08:44:16] !log ayounsi@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet [08:44:24] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet [08:45:35] !log ayounsi@cumin1003 START - Cookbook sre.mysql.pool pool es2051: codfw rack B3 pool after maintenance [08:46:36] they are preexistent [08:47:00] (03PS1) 10TrainBranchBot: group1 to 1.47.0-wmf.10 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308560 (https://phabricator.wikimedia.org/T430829) [08:47:03] (03CR) 10TrainBranchBot: [C:03+2] "Initiated by hashar@deploy1003" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308560 (https://phabricator.wikimedia.org/T430829) (owner: 10TrainBranchBot) [08:48:17] (03Merged) 10jenkins-bot: group1 to 1.47.0-wmf.10 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308560 (https://phabricator.wikimedia.org/T430829) (owner: 10TrainBranchBot) [08:49:22] FIRING: [5x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-1/0/2 (Transport: cr1-eqiad:et-1/1/2 (Arelion, IC-374549) {#12267}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [08:50:07] !log ladsgroup@cumin1003 START - Cookbook sre.mysql.sanitarium_restart [08:50:07] !log ladsgroup@cumin1003 END (FAIL) - Cookbook sre.mysql.sanitarium_restart (exit_code=99) [08:50:23] !log ladsgroup@cumin1003 START - Cookbook sre.mysql.sanitarium_restart [08:50:28] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1004.eqiad.wmnet with OS trixie [08:50:39] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12098491 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host thanos-fe1004.eqiad.wmnet with OS trixie compl... [08:54:10] 06SRE: Problems with SRE Team Vacations Calendar sync - https://phabricator.wikimedia.org/T431303#12098505 (10fgiunchedi) 05Open→03Resolved Tentatively resolving, feel free to reopen if sth is amiss [08:54:37] !log hashar@deploy1003 rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.10 refs T430829 [08:54:41] T430829: 1.47.0-wmf.10 deployment blockers - https://phabricator.wikimedia.org/T430829 [08:56:30] (03CR) 10Cathal Mooney: [C:03+2] HE Transport esams<->eqiad, lower OSPF cost to make live [homer/public] - 10https://gerrit.wikimedia.org/r/1308152 (https://phabricator.wikimedia.org/T424839) (owner: 10Cathal Mooney) [08:57:26] !log merge patch to shift eqiad <-> esams traffic onto new 40G circuit [08:57:27] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [08:58:02] (03Merged) 10jenkins-bot: HE Transport esams<->eqiad, lower OSPF cost to make live [homer/public] - 10https://gerrit.wikimedia.org/r/1308152 (https://phabricator.wikimedia.org/T424839) (owner: 10Cathal Mooney) [09:02:34] !log ladsgroup@cumin1003 END (PASS) - Cookbook sre.mysql.sanitarium_restart (exit_code=0) [09:04:44] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12098548 (10MatthewVernon) [09:07:21] !log fceratto@cumin1003 START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet [09:07:22] !log fceratto@cumin1003 END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet [09:08:43] (03PS1) 10Bartosz Wójtowicz: ml-services: declare ephemeral-storage for all MI300 LLM predictors [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308563 (https://phabricator.wikimedia.org/T431089) [09:19:01] RECOVERY - OpenSearch health check for shards on 9200 on cirrussearch1121 is OK: OK - elasticsearch status production-search-eqiad: cluster_name: production-search-eqiad, status: yellow, timed_out: False, number_of_nodes: 55, number_of_data_nodes: 55, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1287, active_shards: 3877, relocating_shards: 0, initializing_shards: 2, unassigned_shards: 0, delayed_unass [09:19:01] ards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 99.94844031967001 https://wikitech.wikimedia.org/wiki/Search%23Administration [09:19:01] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch1121 is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: green, timed_out: False, number_of_nodes: 30, number_of_data_nodes: 30, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 5196, relocating_shards: 2, initializing_shards: 0, unassigned_shards: 0, delaye [09:19:01] gned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [09:21:35] (03CR) 10Hnowlan: [C:03+1] "One nit, otherwise lgtm!" [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1308156 (https://phabricator.wikimedia.org/T333445) (owner: 10Ladsgroup) [09:22:39] RESOLVED: CirrusSearchNodeIndexingNotIncreasing: Elasticsearch instance cirrussearch1121-production-search-eqiad is not indexing - https://wikitech.wikimedia.org/wiki/Search/Elasticsearch_Administration#Indexing_hung_and_not_making_progress - https://grafana.wikimedia.org/d/JLK3I_siz/elasticsearch-indexing?orgId=1&from=now-3d&to=now&viewPanel=57 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchNodeIndexingNotIncreasing [09:22:50] (03CR) 10Hnowlan: [C:03+2] cassandra: remove disabled CQL check [puppet] - 10https://gerrit.wikimedia.org/r/1307167 (https://phabricator.wikimedia.org/T407120) (owner: 10Hnowlan) [09:25:10] (03CR) 10Hnowlan: [C:03+1] logstash: add parameters needed for security plugin [puppet] - 10https://gerrit.wikimedia.org/r/1305769 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [09:26:03] (03CR) 10Hnowlan: [C:03+1] "lgtm!" [puppet] - 10https://gerrit.wikimedia.org/r/1305718 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [09:26:41] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. [09:27:08] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. [09:27:42] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.5 point update - https://phabricator.wikimedia.org/T427072#12098723 (10MoritzMuehlenhoff) [09:30:16] (03PS1) 10Klausman: hiera/ml-k8s: Add IPIP bits for production workers [puppet] - 10https://gerrit.wikimedia.org/r/1308575 (https://phabricator.wikimedia.org/T420438) [09:31:03] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: codfw rack B3 pool after maintenance [09:31:55] (03CR) 10SD0001: tables-catalog: set betafeatures_user_counts to public visibility (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1298329 (https://phabricator.wikimedia.org/T402145) (owner: 10SD0001) [09:32:22] (03CR) 10Clément Goubert: [C:03+2] redioscope: fix metrics prefix [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308151 (owner: 10Clément Goubert) [09:32:49] !log fceratto@cumin1003 START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet [09:32:49] !log fceratto@cumin1003 END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet [09:33:50] !log fceratto@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Test [09:34:36] !log fceratto@cumin1003 START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet [09:34:36] !log fceratto@cumin1003 END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet [09:34:48] (03Merged) 10jenkins-bot: redioscope: fix metrics prefix [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308151 (owner: 10Clément Goubert) [09:35:14] !log cgoubert@deploy1003 helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply [09:35:40] 06SRE, 10SRE-Access-Requests: Grant shell access (fr-tech-devs) for laurabarluzzi - https://phabricator.wikimedia.org/T431338#12098764 (10cmooney) @greg can I get a thumbs up from you to add Laura to the fr-tech-devs group? Thanks. [09:37:14] 06SRE, 10SRE-Access-Requests: Requesting access to cloudcontrol servers for tlepage - https://phabricator.wikimedia.org/T431010#12098768 (10cmooney) [09:37:33] (03PS1) 10Bartosz Wójtowicz: ml-services: run qwen36-27b on a single GPU on ml-serve1012 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308582 (https://phabricator.wikimedia.org/T431268) [09:43:17] !log cgoubert@deploy1003 helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply [09:43:22] !log cgoubert@deploy1003 helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply [09:43:35] !log cgoubert@deploy1003 helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply [09:47:51] (03CR) 10Ladsgroup: swift: Load the file async (031 comment) [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1308156 (https://phabricator.wikimedia.org/T333445) (owner: 10Ladsgroup) [09:49:16] (03PS5) 10Ladsgroup: swift: Load the file async [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1308156 (https://phabricator.wikimedia.org/T333445) [09:50:01] (03PS1) 10Cathal Mooney: Add shell access to group 'wmcs-roots' for tlepage [puppet] - 10https://gerrit.wikimedia.org/r/1308585 (https://phabricator.wikimedia.org/T431010) [09:50:10] (03CR) 10Btullis: [C:03+1] "Looks good." [puppet] - 10https://gerrit.wikimedia.org/r/1308575 (https://phabricator.wikimedia.org/T420438) (owner: 10Klausman) [09:50:41] (03CR) 10Klausman: [C:03+2] hiera/ml-k8s: Add IPIP bits for production workers [puppet] - 10https://gerrit.wikimedia.org/r/1308575 (https://phabricator.wikimedia.org/T420438) (owner: 10Klausman) [09:53:32] (03CR) 10Hnowlan: "lgtm! btw I have been following SRE conventions on +1 vs +2 on this repo given the need for the image deploy to get changes out - up to ye" [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1308156 (https://phabricator.wikimedia.org/T333445) (owner: 10Ladsgroup) [10:00:05] Deploy window MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260708T1000) [10:00:57] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.5 point update - https://phabricator.wikimedia.org/T427072#12098922 (10MoritzMuehlenhoff) [10:01:18] !log cwilliams@cumin1003 START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet [10:01:36] !log atsuko@cumin1003 START - Cookbook sre.hosts.reimage for host cirrussearch1074.eqiad.wmnet with OS trixie [10:01:40] !log cwilliams@cumin1003 END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet [10:04:45] !log cwilliams@cumin1003 START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet [10:12:08] !log cwilliams@cumin1003 END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet [10:12:12] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host thanos-be1005.eqiad.wmnet with OS trixie [10:12:24] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12098980 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host thanos-be1005.eqiad.wmnet with OS trixie [10:13:33] (03CR) 10Ladsgroup: "I'll ask the team to be clear. My current, very unscientific stance is to have +2 here and +1 for the deployment charts repo since until t" [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1308156 (https://phabricator.wikimedia.org/T333445) (owner: 10Ladsgroup) [10:17:19] !log atsuko@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage [10:20:50] (03CR) 10Hnowlan: opensearch_dashboards: add settings required to enable the security plugin (036 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1305774 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [10:24:21] !log atsuko@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage [10:25:55] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage [10:29:21] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage [10:32:30] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet [10:32:42] PROBLEM - Blazegraph Port for wdqs-blazegraph on wdqs2013 is CRITICAL: connect to address 127.0.0.1 and port 9999: Connection refused https://wikitech.wikimedia.org/wiki/Wikidata_query_service/Runbook [10:33:42] RECOVERY - Blazegraph Port for wdqs-blazegraph on wdqs2013 is OK: TCP OK - 0.000 second response time on 127.0.0.1 port 9999 https://wikitech.wikimedia.org/wiki/Wikidata_query_service/Runbook [10:40:53] !log atsuko@cumin1003 START - Cookbook sre.hosts.reimage for host cirrussearch1075.eqiad.wmnet with OS trixie [10:45:00] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1005.eqiad.wmnet with OS trixie [10:45:14] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12099142 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host thanos-be1005.eqiad.wmnet with OS trixie compl... [10:46:24] !log atsuko@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1074.eqiad.wmnet with OS trixie [10:47:43] !log atsuko@cumin1003 START - Cookbook sre.hosts.reimage for host cirrussearch1076.eqiad.wmnet with OS trixie [10:50:37] (03PS1) 10Brouberol: Revert "dse-k8s: upgrade cloudnative-pg from 1.23 to 1.27" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308598 (https://phabricator.wikimedia.org/T424493) [10:50:40] (03PS1) 10Brouberol: cloudnative-pg-crds: upgrade CRDs to the cloudnative-pg-0.27.1 release [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308599 (https://phabricator.wikimedia.org/T424493) [10:50:45] FIRING: SwiftLowContainerAvailability: Swift eqiad container availability low - https://wikitech.wikimedia.org/wiki/Swift/How_To - https://grafana.wikimedia.org/d/OPgmB1Eiz/swift?panelId=8&fullscreen&orgId=1&var-DC=eqiad - https://alerts.wikimedia.org/?q=alertname%3DSwiftLowContainerAvailability [10:54:07] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.5 point update - https://phabricator.wikimedia.org/T427072#12099217 (10MoritzMuehlenhoff) [10:56:18] !log atsuko@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage [11:00:05] mvolz: OwO what's this, a deployment window?? Services – Citoid / Zotero. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260708T1100). nyaa~ [11:00:19] !log atsuko@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage [11:00:35] (03PS3) 10Blake: renumber-node: Add an option to run homer inline. [cookbooks] - 10https://gerrit.wikimedia.org/r/1308595 (https://phabricator.wikimedia.org/T431443) [11:01:43] (03PS1) 10Clément Goubert: linked-artifacts: Add service catalog and mesh entries [puppet] - 10https://gerrit.wikimedia.org/r/1308604 (https://phabricator.wikimedia.org/T392833) [11:02:26] (03CR) 10Kevin Bazira: [C:03+1] ml-services: declare ephemeral-storage for all MI300 LLM predictors [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308563 (https://phabricator.wikimedia.org/T431089) (owner: 10Bartosz Wójtowicz) [11:02:48] (03PS1) 10Clément Goubert: linked-artifacts: service::catalog to production [puppet] - 10https://gerrit.wikimedia.org/r/1308605 (https://phabricator.wikimedia.org/T392833) [11:03:25] !log atsuko@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage [11:05:19] (03CR) 10Kevin Bazira: [C:03+1] ml-services: run qwen36-27b on a single GPU on ml-serve1012 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308582 (https://phabricator.wikimedia.org/T431268) (owner: 10Bartosz Wójtowicz) [11:05:45] RESOLVED: SwiftLowContainerAvailability: Swift eqiad container availability low - https://wikitech.wikimedia.org/wiki/Swift/How_To - https://grafana.wikimedia.org/d/OPgmB1Eiz/swift?panelId=8&fullscreen&orgId=1&var-DC=eqiad - https://alerts.wikimedia.org/?q=alertname%3DSwiftLowContainerAvailability [11:07:08] (03PS2) 10Clément Goubert: linked-artifacts: service::catalog to production [puppet] - 10https://gerrit.wikimedia.org/r/1308605 (https://phabricator.wikimedia.org/T431108) [11:07:12] (03PS1) 10Brouberol: cloudnative-pg: reflect changes made up the 1.28.1 release [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308606 (https://phabricator.wikimedia.org/T424493) [11:07:15] (03PS1) 10Brouberol: helmfile/cloudnative-pg: grant permissions on new CRDs [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308607 (https://phabricator.wikimedia.org/T424493) [11:07:18] (03PS2) 10Clément Goubert: linked-artifacts: Add service catalog and mesh entries [puppet] - 10https://gerrit.wikimedia.org/r/1308604 (https://phabricator.wikimedia.org/T431108) [11:07:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [11:08:23] !log installing Linux 6.1.176 on Bookworm servers [11:08:24] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [11:08:32] (03PS1) 10Clément Goubert: mobileapps: Add linked-artifacts mesh listener [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308609 (https://phabricator.wikimedia.org/T431108) [11:08:40] !log atsuko@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage [11:08:43] (03PS2) 10Brouberol: cloudnative-pg: reflect changes made up the 1.28.1 release [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308606 (https://phabricator.wikimedia.org/T424493) [11:08:43] (03PS2) 10Brouberol: helmfile/cloudnative-pg: grant permissions on new CRDs [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308607 (https://phabricator.wikimedia.org/T424493) [11:11:15] (03CR) 10CI reject: [V:04-1] cloudnative-pg: reflect changes made up the 1.28.1 release [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308606 (https://phabricator.wikimedia.org/T424493) (owner: 10Brouberol) [11:11:17] (03CR) 10CI reject: [V:04-1] helmfile/cloudnative-pg: grant permissions on new CRDs [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308607 (https://phabricator.wikimedia.org/T424493) (owner: 10Brouberol) [11:11:19] (03PS3) 10Brouberol: cloudnative-pg: reflect changes made up the 1.28.1 release [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308606 (https://phabricator.wikimedia.org/T424493) [11:11:19] (03PS3) 10Brouberol: helmfile/cloudnative-pg: grant permissions on new CRDs [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308607 (https://phabricator.wikimedia.org/T424493) [11:11:34] (03CR) 10CI reject: [V:04-1] mobileapps: Add linked-artifacts mesh listener [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308609 (https://phabricator.wikimedia.org/T431108) (owner: 10Clément Goubert) [11:12:25] (03CR) 10Clément Goubert: "CI fails due to the puppet change dependency not being merged." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308609 (https://phabricator.wikimedia.org/T431108) (owner: 10Clément Goubert) [11:12:40] (03CR) 10Btullis: [C:03+1] Revert "dse-k8s: upgrade cloudnative-pg from 1.23 to 1.27" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308598 (https://phabricator.wikimedia.org/T424493) (owner: 10Brouberol) [11:13:13] (03CR) 10CI reject: [V:04-1] cloudnative-pg: reflect changes made up the 1.28.1 release [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308606 (https://phabricator.wikimedia.org/T424493) (owner: 10Brouberol) [11:13:18] (03CR) 10CI reject: [V:04-1] helmfile/cloudnative-pg: grant permissions on new CRDs [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308607 (https://phabricator.wikimedia.org/T424493) (owner: 10Brouberol) [11:14:07] (03CR) 10Btullis: [C:03+1] cloudnative-pg-crds: upgrade CRDs to the cloudnative-pg-0.27.1 release [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308599 (https://phabricator.wikimedia.org/T424493) (owner: 10Brouberol) [11:14:42] (03PS1) 10Majavah: haproxy: Fix default logging configuration [puppet] - 10https://gerrit.wikimedia.org/r/1308610 [11:16:51] (03CR) 10Btullis: cloudnative-pg: reflect changes made up the 1.28.1 release (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308606 (https://phabricator.wikimedia.org/T424493) (owner: 10Brouberol) [11:17:07] (03PS4) 10Brouberol: cloudnative-pg: reflect changes made up the 1.28.1 release [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308606 (https://phabricator.wikimedia.org/T424493) [11:17:07] (03PS4) 10Brouberol: helmfile/cloudnative-pg: grant permissions on new CRDs [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308607 (https://phabricator.wikimedia.org/T424493) [11:18:17] (03CR) 10Brouberol: cloudnative-pg: reflect changes made up the 1.28.1 release (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308606 (https://phabricator.wikimedia.org/T424493) (owner: 10Brouberol) [11:19:04] !log temporarily remove ganeti2031 from codfw cluster T430910 [11:19:06] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [11:19:07] T430910: codfw: rack B4 maintenance - https://phabricator.wikimedia.org/T430910 [11:19:10] (03CR) 10Clément Goubert: [C:03+1] "LGTM, just to note that if there is a core router in the `switches_to_update`, this will be long and block execution of the rest of the co" [cookbooks] - 10https://gerrit.wikimedia.org/r/1308595 (https://phabricator.wikimedia.org/T431443) (owner: 10Blake) [11:19:11] (03CR) 10CI reject: [V:04-1] cloudnative-pg: reflect changes made up the 1.28.1 release [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308606 (https://phabricator.wikimedia.org/T424493) (owner: 10Brouberol) [11:19:12] (03CR) 10CI reject: [V:04-1] helmfile/cloudnative-pg: grant permissions on new CRDs [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308607 (https://phabricator.wikimedia.org/T424493) (owner: 10Brouberol) [11:19:53] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2047.codfw.wmnet [11:19:55] (03PS2) 10Majavah: haproxy: Fix default logging configuration [puppet] - 10https://gerrit.wikimedia.org/r/1308610 [11:19:55] (03PS1) 10Majavah: P:wmcs::novaproxy: Fix ordering warning [puppet] - 10https://gerrit.wikimedia.org/r/1308611 (https://phabricator.wikimedia.org/T429930) [11:19:57] (03PS1) 10Majavah: P:wmcs::novaproxy: Include requested host in HAProxy access logs [puppet] - 10https://gerrit.wikimedia.org/r/1308612 (https://phabricator.wikimedia.org/T429930) [11:20:53] PROBLEM - ganeti-confd running on ganeti2031 is CRITICAL: PROCS CRITICAL: 0 processes with UID = 109 (gnt-confd), command name ganeti-confd https://wikitech.wikimedia.org/wiki/Ganeti [11:21:32] 10SRE-swift-storage, 06Commons, 07Wikimedia-production-error: Bell_telephone_magazine_(1922)_(14753300151).jpg and other files are missing - https://phabricator.wikimedia.org/T431160#12099335 (10AlexisJazz) [11:21:37] PROBLEM - ganeti-noded running on ganeti2031 is CRITICAL: PROCS CRITICAL: 0 processes with UID = 0 (root), command name ganeti-noded https://wikitech.wikimedia.org/wiki/Ganeti [11:21:42] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2047.codfw.wmnet [11:21:50] FIRING: ProbeDown: Service ganeti2031:1811 has failed probes (tcp_ganeti_noded_ip4) - https://wikitech.wikimedia.org/wiki/Ganeti - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [11:22:40] (03PS5) 10Brouberol: cloudnative-pg: reflect changes made up the 1.28.1 release [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308606 (https://phabricator.wikimedia.org/T424493) [11:22:41] (03PS5) 10Brouberol: helmfile/cloudnative-pg: grant permissions on new CRDs [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308607 (https://phabricator.wikimedia.org/T424493) [11:22:47] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.5 point update - https://phabricator.wikimedia.org/T427072#12099340 (10MoritzMuehlenhoff) [11:23:48] (03PS2) 10Majavah: P:wmcs::novaproxy: Include requested host in HAProxy access logs [puppet] - 10https://gerrit.wikimedia.org/r/1308612 (https://phabricator.wikimedia.org/T429930) [11:24:57] (03PS6) 10Brouberol: cloudnative-pg: reflect changes made up the 1.28.1 release [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308606 (https://phabricator.wikimedia.org/T424493) [11:24:57] (03PS6) 10Brouberol: helmfile/cloudnative-pg: grant permissions on new CRDs [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308607 (https://phabricator.wikimedia.org/T424493) [11:26:10] !log atsuko@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1075.eqiad.wmnet with OS trixie [11:26:13] (03PS3) 10Majavah: P:wmcs::novaproxy: Include requested host in HAProxy access logs [puppet] - 10https://gerrit.wikimedia.org/r/1308612 (https://phabricator.wikimedia.org/T429930) [11:27:52] (03PS4) 10Majavah: P:wmcs::novaproxy: Include requested host in HAProxy access logs [puppet] - 10https://gerrit.wikimedia.org/r/1308612 (https://phabricator.wikimedia.org/T429930) [11:32:00] (03CR) 10Blake: [C:03+2] "I still have some hosts left, so I'll give it a test and comment here with the results before proceeding to merge. Thanks for taking a loo" [cookbooks] - 10https://gerrit.wikimedia.org/r/1308595 (https://phabricator.wikimedia.org/T431443) (owner: 10Blake) [11:32:41] !log atsuko@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1076.eqiad.wmnet with OS trixie [11:33:28] 10SRE-swift-storage, 06Commons, 07Wikimedia-production-error: Bell_telephone_magazine_(1922)_(14753300151).jpg and other files are missing - https://phabricator.wikimedia.org/T431160#12099384 (10AlexisJazz) More missing files: * https://commons.wikimedia.org/wiki/File:Bell_telephone_magazine_(1922)_(1456940... [11:33:46] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.5 point update - https://phabricator.wikimedia.org/T427072#12099385 (10MoritzMuehlenhoff) [11:33:51] (03CR) 10Blake: renumber-node: Add an option to run homer inline. [cookbooks] - 10https://gerrit.wikimedia.org/r/1308595 (https://phabricator.wikimedia.org/T431443) (owner: 10Blake) [11:34:36] (03CR) 10Bartosz Wójtowicz: [C:03+2] ml-services: run qwen36-27b on a single GPU on ml-serve1012 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308582 (https://phabricator.wikimedia.org/T431268) (owner: 10Bartosz Wójtowicz) [11:35:25] (03PS3) 10Majavah: P:monitoring: Remove absented resources [puppet] - 10https://gerrit.wikimedia.org/r/1307745 [11:36:55] (03Merged) 10jenkins-bot: ml-services: run qwen36-27b on a single GPU on ml-serve1012 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308582 (https://phabricator.wikimedia.org/T431268) (owner: 10Bartosz Wójtowicz) [11:38:50] !log bwojtowicz@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . [11:43:09] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie [11:43:26] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12099440 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host thanos-be1006.eqiad.wmnet with OS trixie [11:44:08] 10SRE-swift-storage, 06Commons, 07Wikimedia-production-error: Bell_telephone_magazine_(1922)_(14753300151).jpg and other files are missing - https://phabricator.wikimedia.org/T431160#12099449 (10AlexisJazz) →14Duplicate dup:03T124101 [11:44:15] 06SRE, 10SRE-swift-storage, 06Commons, 10MediaWiki-File-management: Specific revisions of multiple files missing from Swift - 404 Not Found returned - https://phabricator.wikimedia.org/T124101#12099451 (10AlexisJazz) [11:46:22] 10ops-drmrs, 06SRE, 06Infrastructure-Foundations: FY2627 drmrs: upgrade infra links from 40G to 100G - https://phabricator.wikimedia.org/T431562 (10cmooney) 03NEW p:05Triage→03Medium [11:46:50] RESOLVED: ProbeDown: Service ganeti2031:1811 has failed probes (tcp_ganeti_noded_ip4) - https://wikitech.wikimedia.org/wiki/Ganeti - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [11:47:03] 06SRE, 10SRE-swift-storage, 06Commons, 10MediaWiki-File-management: Specific revisions of multiple files missing from Swift - 404 Not Found returned - https://phabricator.wikimedia.org/T124101#12099472 (10AlexisJazz) >>! In T124101#2842864, @MarkTraceur wrote: > I don't see this as "high" priority, but I'm... [11:47:20] FIRING: ProbeDown: Service ganeti2031:1811 has failed probes (tcp_ganeti_noded_ip4) - https://wikitech.wikimedia.org/wiki/Ganeti - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [11:47:48] (03CR) 10Btullis: [C:03+1] cloudnative-pg: reflect changes made up the 1.28.1 release (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308606 (https://phabricator.wikimedia.org/T424493) (owner: 10Brouberol) [11:48:26] (03CR) 10Filippo Giunchedi: [C:03+1] haproxy: Fix default logging configuration [puppet] - 10https://gerrit.wikimedia.org/r/1308610 (owner: 10Majavah) [11:48:38] (03CR) 10Filippo Giunchedi: [C:03+1] P:wmcs::novaproxy: Fix ordering warning [puppet] - 10https://gerrit.wikimedia.org/r/1308611 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [11:48:54] (03CR) 10Filippo Giunchedi: [C:03+1] P:wmcs::novaproxy: Include requested host in HAProxy access logs [puppet] - 10https://gerrit.wikimedia.org/r/1308612 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [11:49:27] 06SRE, 06MediaWiki-Media-Platform-Team, 06ServiceOps new, 10ServiceOps-Services-Oids, and 2 others: Thumbor-k8s performance improvements - https://phabricator.wikimedia.org/T333445#12099488 (10Krinkle) [11:50:20] 10SRE-swift-storage, 06Data-Persistence, 06MediaWiki-Media-Platform-Team, 10Prod-Kubernetes, and 6 others: Fix thumbor discovery records and make swift use them - https://phabricator.wikimedia.org/T397618#12099493 (10Krinkle) [11:50:43] (03CR) 10Majavah: [C:03+2] haproxy: Fix default logging configuration [puppet] - 10https://gerrit.wikimedia.org/r/1308610 (owner: 10Majavah) [11:50:59] (03CR) 10Majavah: [C:03+2] P:wmcs::novaproxy: Fix ordering warning [puppet] - 10https://gerrit.wikimedia.org/r/1308611 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [11:51:08] (03CR) 10Majavah: [C:03+2] P:wmcs::novaproxy: Include requested host in HAProxy access logs [puppet] - 10https://gerrit.wikimedia.org/r/1308612 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [11:52:20] RESOLVED: ProbeDown: Service ganeti2031:1811 has failed probes (tcp_ganeti_noded_ip4) - https://wikitech.wikimedia.org/wiki/Ganeti - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [11:54:42] mvernon@cumin2003 reimage (PID 4107300) is awaiting input [11:56:21] (03CR) 10Brouberol: [C:03+2] Revert "dse-k8s: upgrade cloudnative-pg from 1.23 to 1.27" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308598 (https://phabricator.wikimedia.org/T424493) (owner: 10Brouberol) [11:56:24] (03CR) 10Brouberol: [C:03+2] cloudnative-pg-crds: upgrade CRDs to the cloudnative-pg-0.27.1 release [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308599 (https://phabricator.wikimedia.org/T424493) (owner: 10Brouberol) [11:56:29] (03CR) 10Brouberol: [C:03+2] cloudnative-pg: reflect changes made up the 1.28.1 release [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308606 (https://phabricator.wikimedia.org/T424493) (owner: 10Brouberol) [11:56:33] (03CR) 10Brouberol: [C:03+2] helmfile/cloudnative-pg: grant permissions on new CRDs [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308607 (https://phabricator.wikimedia.org/T424493) (owner: 10Brouberol) [11:58:02] (03Abandoned) 10Brouberol: Revert "dse-k8s: upgrade cloudnative-pg from 1.23 to 1.27" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308598 (https://phabricator.wikimedia.org/T424493) (owner: 10Brouberol) [11:58:08] (03CR) 10CI reject: [V:04-1] cloudnative-pg-crds: upgrade CRDs to the cloudnative-pg-0.27.1 release [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308599 (https://phabricator.wikimedia.org/T424493) (owner: 10Brouberol) [11:58:09] (03CR) 10CI reject: [V:04-1] cloudnative-pg: reflect changes made up the 1.28.1 release [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308606 (https://phabricator.wikimedia.org/T424493) (owner: 10Brouberol) [11:58:10] (03CR) 10CI reject: [V:04-1] helmfile/cloudnative-pg: grant permissions on new CRDs [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308607 (https://phabricator.wikimedia.org/T424493) (owner: 10Brouberol) [11:58:40] (03PS2) 10Brouberol: cloudnative-pg-crds: upgrade CRDs to the cloudnative-pg-0.27.1 release [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308599 (https://phabricator.wikimedia.org/T424493) [11:58:55] (03PS7) 10Brouberol: cloudnative-pg: reflect changes made up the 1.28.1 release [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308606 (https://phabricator.wikimedia.org/T424493) [11:59:01] (03PS7) 10Brouberol: helmfile/cloudnative-pg: grant permissions on new CRDs [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308607 (https://phabricator.wikimedia.org/T424493) [12:00:29] (03CR) 10Filippo Giunchedi: "recheck" [puppet] - 10https://gerrit.wikimedia.org/r/1308126 (https://phabricator.wikimedia.org/T411248) (owner: 10Filippo Giunchedi) [12:00:39] (03CR) 10Brouberol: [V:03+2 C:03+2] cloudnative-pg-crds: upgrade CRDs to the cloudnative-pg-0.27.1 release [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308599 (https://phabricator.wikimedia.org/T424493) (owner: 10Brouberol) [12:00:45] (03CR) 10Brouberol: [V:03+2 C:03+2] cloudnative-pg: reflect changes made up the 1.28.1 release [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308606 (https://phabricator.wikimedia.org/T424493) (owner: 10Brouberol) [12:00:49] (03CR) 10Brouberol: [V:03+2 C:03+2] helmfile/cloudnative-pg: grant permissions on new CRDs [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308607 (https://phabricator.wikimedia.org/T424493) (owner: 10Brouberol) [12:01:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 21.8% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [12:01:55] !log mvernon@cumin2003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie [12:02:12] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12099580 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host thanos-be1006.eqiad.wmnet with OS trixie execu... [12:02:34] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie [12:02:45] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12099583 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host thanos-be1006.eqiad.wmnet with OS trixie [12:04:50] (03PS2) 10Filippo Giunchedi: openstack: deprecate icinga check-neutron-conntrack [puppet] - 10https://gerrit.wikimedia.org/r/1302764 (https://phabricator.wikimedia.org/T328502) [12:05:14] (03PS1) 10Bartosz Wójtowicz: ml-services: declare ephemeral-storage for all MI300 LLM predictors [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308618 (https://phabricator.wikimedia.org/T431089) [12:05:28] (03CR) 10Filippo Giunchedi: "Ack, fixed in next PS. Will followup with actual removal too" [puppet] - 10https://gerrit.wikimedia.org/r/1302764 (https://phabricator.wikimedia.org/T328502) (owner: 10Filippo Giunchedi) [12:06:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.56% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [12:06:25] FIRING: [2x] SystemdUnitFailed: fetch-rings-thanos.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [12:06:53] (03Abandoned) 10Bartosz Wójtowicz: ml-services: declare ephemeral-storage for all MI300 LLM predictors [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308618 (https://phabricator.wikimedia.org/T431089) (owner: 10Bartosz Wójtowicz) [12:07:33] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. [12:07:53] (03PS2) 10Bartosz Wójtowicz: ml-services: declare ephemeral-storage for all MI300 LLM predictors [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308563 (https://phabricator.wikimedia.org/T431089) [12:16:45] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.5 point update - https://phabricator.wikimedia.org/T427072#12099648 (10MoritzMuehlenhoff) [12:17:47] (03PS1) 10Raymond Ndibe: webservice-runner: Add pre-flight validity checks [docker-images/toollabs-images] - 10https://gerrit.wikimedia.org/r/1308620 (https://phabricator.wikimedia.org/T431563) [12:18:21] (03PS1) 10Brouberol: cloudnative-pg: restore cnpg-mutating-webhook-configuration resource [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308621 (https://phabricator.wikimedia.org/T424493) [12:18:35] (03PS1) 10Sbisson: Enable Article Guidance extension on itwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308622 (https://phabricator.wikimedia.org/T431540) [12:18:40] (03PS1) 10Majavah: P:wmcs::novaproxy: Add HAProxy in front of all requests [puppet] - 10https://gerrit.wikimedia.org/r/1308623 (https://phabricator.wikimedia.org/T429930) [12:20:26] (03PS2) 10Majavah: P:wmcs::novaproxy: Add HAProxy in front of all requests [puppet] - 10https://gerrit.wikimedia.org/r/1308623 (https://phabricator.wikimedia.org/T429930) [12:20:26] (03CR) 10CI reject: [V:04-1] cloudnative-pg: restore cnpg-mutating-webhook-configuration resource [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308621 (https://phabricator.wikimedia.org/T424493) (owner: 10Brouberol) [12:20:44] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, July 08 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployc" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308622 (https://phabricator.wikimedia.org/T431540) (owner: 10Sbisson) [12:20:45] (03PS2) 10Brouberol: cloudnative-pg: restore cnpg-mutating-webhook-configuration resource [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308621 (https://phabricator.wikimedia.org/T424493) [12:20:47] (03PS2) 10Aklapper: Phabricator: Remove unused fixed_settings.yaml stuff [puppet] - 10https://gerrit.wikimedia.org/r/1130323 (https://phabricator.wikimedia.org/T239355) [12:21:11] (03PS3) 10Majavah: P:wmcs::novaproxy: Add HAProxy in front of all requests [puppet] - 10https://gerrit.wikimedia.org/r/1308623 (https://phabricator.wikimedia.org/T429930) [12:21:17] (03PS3) 10Filippo Giunchedi: dumps: add nfs_client for dumps-nfs.w.o [puppet] - 10https://gerrit.wikimedia.org/r/1308126 (https://phabricator.wikimedia.org/T411248) [12:21:17] (03PS3) 10Filippo Giunchedi: statistics: optional support for dumps-nfs.w.o [puppet] - 10https://gerrit.wikimedia.org/r/1308127 (https://phabricator.wikimedia.org/T411248) [12:21:17] (03PS3) 10Filippo Giunchedi: wmcs: add new dumps-nfs.w.o support [puppet] - 10https://gerrit.wikimedia.org/r/1308128 (https://phabricator.wikimedia.org/T411248) [12:21:20] (03PS1) 10Filippo Giunchedi: fixup! dumps: add nfs_client for dumps-nfs.w.o [puppet] - 10https://gerrit.wikimedia.org/r/1308624 [12:21:59] (03PS3) 10Brouberol: cloudnative-pg: restore cnpg-mutating-webhook-configuration resource [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308621 (https://phabricator.wikimedia.org/T424493) [12:22:13] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, July 08 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployc" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1307836 (https://phabricator.wikimedia.org/T431334) (owner: 10Danielyepezgarces) [12:22:16] (03PS4) 10Filippo Giunchedi: dumps: add nfs_client for dumps-nfs.w.o [puppet] - 10https://gerrit.wikimedia.org/r/1308126 (https://phabricator.wikimedia.org/T411248) [12:22:16] (03PS4) 10Filippo Giunchedi: statistics: optional support for dumps-nfs.w.o [puppet] - 10https://gerrit.wikimedia.org/r/1308127 (https://phabricator.wikimedia.org/T411248) [12:22:17] (03PS4) 10Filippo Giunchedi: wmcs: add new dumps-nfs.w.o support [puppet] - 10https://gerrit.wikimedia.org/r/1308128 (https://phabricator.wikimedia.org/T411248) [12:22:28] (03Abandoned) 10Filippo Giunchedi: fixup! dumps: add nfs_client for dumps-nfs.w.o [puppet] - 10https://gerrit.wikimedia.org/r/1308624 (owner: 10Filippo Giunchedi) [12:23:55] (03CR) 10CI reject: [V:04-1] P:wmcs::novaproxy: Add HAProxy in front of all requests [puppet] - 10https://gerrit.wikimedia.org/r/1308623 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [12:24:18] (03PS4) 10Majavah: P:wmcs::novaproxy: Add HAProxy in front of all requests [puppet] - 10https://gerrit.wikimedia.org/r/1308623 (https://phabricator.wikimedia.org/T429930) [12:25:07] (03CR) 10CI reject: [V:04-1] dumps: add nfs_client for dumps-nfs.w.o [puppet] - 10https://gerrit.wikimedia.org/r/1308126 (https://phabricator.wikimedia.org/T411248) (owner: 10Filippo Giunchedi) [12:26:36] (03PS1) 10Filippo Giunchedi: openstack: remove openstack::monitor::neutron::l3_agent_conntrack [puppet] - 10https://gerrit.wikimedia.org/r/1308626 (https://phabricator.wikimedia.org/T328502) [12:27:05] (03CR) 10CI reject: [V:04-1] P:wmcs::novaproxy: Add HAProxy in front of all requests [puppet] - 10https://gerrit.wikimedia.org/r/1308623 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [12:27:16] !log mvernon@cumin2003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1006.eqiad.wmnet with OS trixie [12:27:31] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12099698 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host thanos-be1006.eqiad.wmnet with OS trixie execu... [12:27:43] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie [12:28:00] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12099700 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host thanos-be1006.eqiad.wmnet with OS trixie [12:28:29] (03PS5) 10Majavah: P:wmcs::novaproxy: Add HAProxy in front of all requests [puppet] - 10https://gerrit.wikimedia.org/r/1308623 (https://phabricator.wikimedia.org/T429930) [12:29:40] (03PS6) 10Majavah: P:wmcs::novaproxy: Add HAProxy in front of all requests [puppet] - 10https://gerrit.wikimedia.org/r/1308623 (https://phabricator.wikimedia.org/T429930) [12:30:28] (03PS5) 10Filippo Giunchedi: dumps: add nfs_client for dumps-nfs.w.o [puppet] - 10https://gerrit.wikimedia.org/r/1308126 (https://phabricator.wikimedia.org/T411248) [12:30:28] (03PS5) 10Filippo Giunchedi: statistics: optional support for dumps-nfs.w.o [puppet] - 10https://gerrit.wikimedia.org/r/1308127 (https://phabricator.wikimedia.org/T411248) [12:30:28] (03PS5) 10Filippo Giunchedi: wmcs: add new dumps-nfs.w.o support [puppet] - 10https://gerrit.wikimedia.org/r/1308128 (https://phabricator.wikimedia.org/T411248) [12:30:36] (03CR) 10CI reject: [V:04-1] P:wmcs::novaproxy: Add HAProxy in front of all requests [puppet] - 10https://gerrit.wikimedia.org/r/1308623 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [12:32:29] (03CR) 10Brouberol: [C:03+2] cloudnative-pg: restore cnpg-mutating-webhook-configuration resource [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308621 (https://phabricator.wikimedia.org/T424493) (owner: 10Brouberol) [12:32:36] (03CR) 10Filippo Giunchedi: "To be merged after I64df5db6b3f4c rollout" [puppet] - 10https://gerrit.wikimedia.org/r/1308626 (https://phabricator.wikimedia.org/T328502) (owner: 10Filippo Giunchedi) [12:33:37] (03CR) 10Hashar: [C:04-1] "The user addition to the group is guarded by `$::realm == 'labs'` because in production groups membership is managed via the data module. " [puppet] - 10https://gerrit.wikimedia.org/r/1308194 (https://phabricator.wikimedia.org/T418521) (owner: 10Dzahn) [12:33:54] mvernon@cumin2003 reimage (PID 4121941) is awaiting input [12:34:12] !log mvernon@cumin2003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie [12:34:32] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12099731 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host thanos-be1006.eqiad.wmnet with OS trixie execu... [12:34:44] !log mvernon@cumin2003 START - Cookbook sre.hosts.provision for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [12:36:46] !log blake@cumin1003 START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1051.eqiad.wmnet [12:36:50] !log blake@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1051.eqiad.wmnet [12:37:25] !log blake@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1051.eqiad.wmnet [12:38:16] !log blake@cumin1003 START - Cookbook sre.hosts.reimage for host wikikube-worker1051.eqiad.wmnet with OS trixie [12:38:43] !log blake@cumin1003 START - Cookbook sre.hosts.move-vlan for host wikikube-worker1051 [12:38:53] !log blake@cumin1003 START - Cookbook sre.dns.netbox [12:39:05] (03CR) 10Hnowlan: "lgtm, someone else on the team will give the +1 soon" [puppet] - 10https://gerrit.wikimedia.org/r/1305989 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [12:39:46] mvernon@cumin2003 provision (PID 4122673) is awaiting input [12:39:58] (03PS7) 10Majavah: P:wmcs::novaproxy: Add HAProxy in front of all requests [puppet] - 10https://gerrit.wikimedia.org/r/1308623 (https://phabricator.wikimedia.org/T429930) [12:40:51] (03CR) 10CI reject: [V:04-1] P:wmcs::novaproxy: Add HAProxy in front of all requests [puppet] - 10https://gerrit.wikimedia.org/r/1308623 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [12:43:05] !log blake@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" [12:43:10] !log blake@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003" [12:43:10] !log blake@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [12:43:10] !log blake@cumin1003 START - Cookbook sre.dns.wipe-cache wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [12:43:14] !log blake@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [12:43:15] !log blake@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1051 [12:43:26] !log installing Python 3.11 security updates [12:43:27] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [12:43:40] !log blake@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1051 [12:43:40] !log blake@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1051 [12:45:42] (03PS3) 10Hashar: ci::docker: ensure jenkins agent user is in docker group also in prod [puppet] - 10https://gerrit.wikimedia.org/r/1308194 (https://phabricator.wikimedia.org/T418521) (owner: 10Dzahn) [12:47:48] (03CR) 10Hashar: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308194 (https://phabricator.wikimedia.org/T418521) (owner: 10Dzahn) [12:49:27] (03CR) 10Muehlenhoff: [C:03+2] Update redis-misc-canary alias [puppet] - 10https://gerrit.wikimedia.org/r/1305344 (owner: 10Muehlenhoff) [12:50:01] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.5 point update - https://phabricator.wikimedia.org/T427072#12099880 (10MoritzMuehlenhoff) [12:50:23] (03PS8) 10Majavah: P:wmcs::novaproxy: Add HAProxy in front of all requests [puppet] - 10https://gerrit.wikimedia.org/r/1308623 (https://phabricator.wikimedia.org/T429930) [12:50:25] FIRING: [5x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-1/0/2 (Transport: cr1-eqiad:et-1/1/2 (Arelion, IC-374549) {#12267}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [12:50:43] !log mvernon@cumin2003 END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [12:52:28] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie [12:52:40] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12099895 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host thanos-be1006.eqiad.wmnet with OS trixie [12:55:32] (03CR) 10Majavah: [V:03+1] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/8950/co" [puppet] - 10https://gerrit.wikimedia.org/r/1308623 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [12:56:11] (03CR) 10Majavah: [C:03+1] openstack: deprecate icinga check-neutron-conntrack [puppet] - 10https://gerrit.wikimedia.org/r/1302764 (https://phabricator.wikimedia.org/T328502) (owner: 10Filippo Giunchedi) [12:56:31] (03PS1) 10Klausman: hiera/services: Switch ML k8s workers/ingress to IPIP [puppet] - 10https://gerrit.wikimedia.org/r/1308586 (https://phabricator.wikimedia.org/T420438) [12:56:48] (03CR) 10Fabfur: [C:03+1] "lgtm!" [puppet] - 10https://gerrit.wikimedia.org/r/1308586 (https://phabricator.wikimedia.org/T420438) (owner: 10Klausman) [12:58:19] !log klausman@cumin2002 START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@codfw [12:58:24] (03CR) 10Klausman: [V:03+1 C:03+2] hiera/services: Switch ML k8s workers/ingress to IPIP [puppet] - 10https://gerrit.wikimedia.org/r/1308586 (https://phabricator.wikimedia.org/T420438) (owner: 10Klausman) [13:00:05] Lucas_WMDE, urbanecm, and TheresNoTime: OwO what's this, a deployment window?? UTC afternoon backport window. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260708T1300). nyaa~ [13:00:05] Tran, stephanebisson, and Neriah: A patch you scheduled for UTC afternoon backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [13:00:09] o/ [13:00:11] hey [13:00:13] o/ [13:00:14] o/ [13:00:43] I’m in a meeting for a bit but should be able to deploy later [13:01:03] Tran are you able to deploy your change? [13:01:09] Tran, stephanebisson: you can probably both go ahead with your changes? [13:01:09] Yes I can self-service [13:01:14] (03CR) 10Kevin Bazira: [C:03+1] ml-services: declare ephemeral-storage for all MI300 LLM predictors [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308563 (https://phabricator.wikimedia.org/T431089) (owner: 10Bartosz Wójtowicz) [13:01:15] and/or deploy for Neriah :) but I can also do that later otherwise [13:01:20] !log blake@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage [13:01:28] I'm doing this in a meeting so can't deploy anyone else's sorry but I can get started if I'm first [13:01:29] (03CR) 10CDobbins: [V:03+1] "The reason for quoting all IPs as opposed to just IPv4 subnets and IPv6 ones is for simplicity and minimizing future problems. Checking fo" [puppet] - 10https://gerrit.wikimedia.org/r/1308150 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [13:01:37] Train, go ahead, I'll go after [13:01:43] I can do mine [13:02:25] RESOLVED: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:02:26] k I'll start then [13:02:42] (03CR) 10TrainBranchBot: [C:03+2] "Approved by stran@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308553 (https://phabricator.wikimedia.org/T431316) (owner: 10STran) [13:04:45] !log klausman@cumin2002 START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs [13:04:51] !log installing jq security updates [13:04:52] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [13:05:10] !log blake@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage [13:05:18] (03CR) 10Majavah: hieradata: issue fixes (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1308150 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [13:05:22] !log atsuko@cumin1003 START - Cookbook sre.hosts.reimage for host cirrussearch1081.eqiad.wmnet with OS trixie [13:05:46] !log klausman@cumin2002 END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs [13:05:47] !log klausman@cumin2002 END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@codfw [13:09:16] (03Merged) 10jenkins-bot: Deploy IRS to enwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308553 (https://phabricator.wikimedia.org/T431316) (owner: 10STran) [13:09:55] FIRING: [2x] SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:10:00] !log stran@deploy1003 Started scap sync-world: Backport for [[gerrit:1308553|Deploy IRS to enwiki (T431316)]] [13:10:03] T431316: Deploy IRS MVP on enwiki - https://phabricator.wikimedia.org/T431316 [13:10:10] RESOLVED: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:12:12] !log atsuko@cumin1003 START - Cookbook sre.hosts.reimage for host cirrussearch1118.eqiad.wmnet with OS trixie [13:12:16] !log stran@deploy1003 stran: Backport for [[gerrit:1308553|Deploy IRS to enwiki (T431316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [13:12:33] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage [13:12:37] (03CR) 10CDobbins: [V:03+1] hieradata: issue fixes (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1308150 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [13:12:40] !log atsuko@cumin1003 START - Cookbook sre.hosts.move-vlan for host cirrussearch1118 [13:13:13] (03CR) 10Lucas Werkmeister (WMDE): extwiki: Rename wgSitename to Güiquipedia (031 comment) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1307836 (https://phabricator.wikimedia.org/T431334) (owner: 10Danielyepezgarces) [13:13:19] Neriah: ^ question for you up there [13:13:56] !log atsuko@cumin1003 START - Cookbook sre.dns.netbox [13:14:11] * Lucas_WMDE afk for a moment to pick up pakige [13:14:12] (03PS3) 10CDobbins: hieradata: issue fixes [puppet] - 10https://gerrit.wikimedia.org/r/1308150 (https://phabricator.wikimedia.org/T401832) [13:14:55] FIRING: [2x] SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:15:03] i missed it [13:15:10] (03CR) 10CDobbins: [V:03+1] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/8951/co" [puppet] - 10https://gerrit.wikimedia.org/r/1308150 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [13:15:12] what did you ask? [13:15:34] Lucas_WMDE [13:15:50] (03CR) 10CDobbins: hieradata: issue fixes (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1308150 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [13:15:59] !log sukhe@cumin1003 START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox) [13:16:11] !log sukhe@cumin1003 START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough [13:17:24] Oh shoot I'm so sorry, I' mtesting now [13:17:32] 06SRE, 06Infrastructure-Foundations: Integrate Bookworm 12.14 point update - https://phabricator.wikimedia.org/T426759#12100035 (10MoritzMuehlenhoff) [13:17:48] !log stran@deploy1003 stran: Continuing with deployment [13:17:57] (03CR) 10David Caro: [C:03+1] "LGTM (in case it helps xd)" [puppet] - 10https://gerrit.wikimedia.org/r/1308585 (https://phabricator.wikimedia.org/T431010) (owner: 10Cathal Mooney) [13:18:48] !log atsuko@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" [13:18:53] !log atsuko@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003" [13:18:53] !log atsuko@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [13:18:53] !log atsuko@cumin1003 START - Cookbook sre.dns.wipe-cache cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [13:18:57] !log atsuko@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [13:18:57] !log atsuko@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1118 [13:18:59] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage [13:19:32] Neriah: I left a comment on Gerrit [13:19:49] (and then afterwards saw that the change was uploaded by someone else, so idk if you even feel responsible for answering it ^^) [13:19:52] https://gerrit.wikimedia.org/r/c/operations/mediawiki-config/+/1307836 [13:20:28] * Lucas_WMDE can deploy now btw [13:21:00] !log atsuko@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage [13:22:12] !log stran@deploy1003 Finished scap sync-world: Backport for [[gerrit:1308553|Deploy IRS to enwiki (T431316)]] (duration: 12m 12s) [13:22:15] T431316: Deploy IRS MVP on enwiki - https://phabricator.wikimedia.org/T431316 [13:22:25] stephanebisson: I'm done, sorry! [13:22:39] Tran, thanks [13:22:59] (03CR) 10TrainBranchBot: [C:03+2] "Approved by sbisson@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308622 (https://phabricator.wikimedia.org/T431540) (owner: 10Sbisson) [13:23:59] Lucas_WMDE: I'll leave this to dyepezg [13:24:00] I'm the one who scheduled the deployment because of his request in phab, but I think it's better to leave it to him [13:24:07] !log atsuko@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1118 [13:24:07] !log atsuko@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1118 [13:24:18] Neriah: okay, fair enough [13:24:44] thanks :) [13:24:57] !log atsuko@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage [13:26:34] !log blake@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1051.eqiad.wmnet with OS trixie [13:26:40] (03CR) 10Ssingh: hieradata: issue fixes (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1308150 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [13:27:15] (03Merged) 10jenkins-bot: Enable Article Guidance extension on itwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308622 (https://phabricator.wikimedia.org/T431540) (owner: 10Sbisson) [13:27:40] !log sbisson@deploy1003 Started scap sync-world: Backport for [[gerrit:1308622|Enable Article Guidance extension on itwiki (T431540)]] [13:27:43] T431540: Enable AG workflow in Italian Wiki - https://phabricator.wikimedia.org/T431540 [13:28:18] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. [13:28:25] (03PS4) 10CDobbins: hieradata: issue fixes [puppet] - 10https://gerrit.wikimedia.org/r/1308150 (https://phabricator.wikimedia.org/T401832) [13:28:34] (03CR) 10CDobbins: hieradata: issue fixes (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1308150 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [13:28:47] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. [13:29:29] (03CR) 10CDobbins: [V:03+1] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/8952/co" [puppet] - 10https://gerrit.wikimedia.org/r/1308150 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [13:29:54] !log sbisson@deploy1003 sbisson: Backport for [[gerrit:1308622|Enable Article Guidance extension on itwiki (T431540)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [13:30:31] !log sukhe@cumin1003 END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough [13:30:42] !log installing openssh security updates [13:30:43] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [13:30:57] !log sbisson@deploy1003 sbisson: Continuing with deployment [13:31:24] (03CR) 10Ssingh: "Thanks. Can you also compare https://puppet-compiler.wmflabs.org/output/1308150/8952/dns7002.wikimedia.org/fulldiff.html to see if we are " [puppet] - 10https://gerrit.wikimedia.org/r/1308150 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [13:34:14] !log atsuko@cumin1003 START - Cookbook sre.hosts.reimage for host cirrussearch1119.eqiad.wmnet with OS trixie [13:34:41] !log atsuko@cumin1003 START - Cookbook sre.hosts.move-vlan for host cirrussearch1119 [13:35:27] !log sbisson@deploy1003 Finished scap sync-world: Backport for [[gerrit:1308622|Enable Article Guidance extension on itwiki (T431540)]] (duration: 07m 46s) [13:35:30] T431540: Enable AG workflow in Italian Wiki - https://phabricator.wikimedia.org/T431540 [13:36:08] !log atsuko@cumin1003 START - Cookbook sre.dns.netbox [13:36:10] !log atsuko@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage [13:36:21] (03PS5) 10CDobbins: hieradata: issue fixes [puppet] - 10https://gerrit.wikimedia.org/r/1308150 (https://phabricator.wikimedia.org/T401832) [13:37:00] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1006.eqiad.wmnet with OS trixie [13:37:18] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12100079 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host thanos-be1006.eqiad.wmnet with OS trixie compl... [13:37:29] !log UTC afternoon backport+config window done [13:37:30] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [13:37:32] (03CR) 10CDobbins: [V:03+1] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/8953/co" [puppet] - 10https://gerrit.wikimedia.org/r/1308150 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [13:39:24] !log installing krb5 security updates [13:39:25] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [13:39:29] (03CR) 10Eevans: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1305988 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [13:39:45] (03CR) 10JMeybohm: [C:03+1] admin_ng: Restrict wikifunctions network access to DNS [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305799 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [13:40:25] !log atsuko@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage [13:40:37] !log atsuko@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" [13:40:41] !log atsuko@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003" [13:40:41] !log atsuko@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [13:40:41] !log atsuko@cumin1003 START - Cookbook sre.dns.wipe-cache cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [13:40:45] !log atsuko@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [13:40:45] !log atsuko@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1119 [13:41:31] (03CR) 10JMeybohm: [C:03+1] linked-artifacts: Add service catalog and mesh entries [puppet] - 10https://gerrit.wikimedia.org/r/1308604 (https://phabricator.wikimedia.org/T431108) (owner: 10Clément Goubert) [13:41:37] (03CR) 10JMeybohm: [C:03+1] linked-artifacts: service::catalog to production [puppet] - 10https://gerrit.wikimedia.org/r/1308605 (https://phabricator.wikimedia.org/T431108) (owner: 10Clément Goubert) [13:41:45] FIRING: SwiftLowContainerAvailability: Swift eqiad container availability low - https://wikitech.wikimedia.org/wiki/Swift/How_To - https://grafana.wikimedia.org/d/OPgmB1Eiz/swift?panelId=8&fullscreen&orgId=1&var-DC=eqiad - https://alerts.wikimedia.org/?q=alertname%3DSwiftLowContainerAvailability [13:41:53] !log atsuko@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1119 [13:41:53] !log atsuko@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1119 [13:45:47] !log atsuko@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1081.eqiad.wmnet with OS trixie [13:47:17] (03PS6) 10CDobbins: hieradata: issue fixes [puppet] - 10https://gerrit.wikimedia.org/r/1308150 (https://phabricator.wikimedia.org/T401832) [13:49:11] !log cwilliams@cumin1003 START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet [13:50:01] (03PS7) 10CDobbins: hieradata: issue fixes [puppet] - 10https://gerrit.wikimedia.org/r/1308150 (https://phabricator.wikimedia.org/T401832) [13:50:21] !log blake@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1051.eqiad.wmnet [13:50:22] !log blake@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1051.eqiad.wmnet [13:50:24] !log blake@cumin1003 END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1051.eqiad.wmnet [13:50:51] !log klausman@cumin2002 START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@eqiad [13:51:14] (03CR) 10CDobbins: [V:03+1] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/8954/co" [puppet] - 10https://gerrit.wikimedia.org/r/1308150 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [13:52:17] !log fceratto@cumin1003 START - Cookbook sre.mysql.depool [13:52:20] cwilliams@cumin1003 multiinstance_reboot (PID 2660818) is awaiting input [13:52:37] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.depool (exit_code=0) [13:52:38] (03PS8) 10CDobbins: hieradata: issue fixes [puppet] - 10https://gerrit.wikimedia.org/r/1308150 (https://phabricator.wikimedia.org/T401832) [13:52:48] !log cwilliams@cumin1003 END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet [13:53:27] !log atsuko@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage [13:53:27] (03CR) 10CDobbins: [V:03+1] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/8955/co" [puppet] - 10https://gerrit.wikimedia.org/r/1308150 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [13:53:45] (03CR) 10Blake: [C:03+2] "A node was successfully reimaged using this change: `test-cookbook -c 1308595 sre.k8s.renumber-node -t T421711 wikikube-worker1051.eqiad" [cookbooks] - 10https://gerrit.wikimedia.org/r/1308595 (https://phabricator.wikimedia.org/T431443) (owner: 10Blake) [13:54:11] !log installing libcap2 security updates [13:54:12] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [13:56:45] RESOLVED: SwiftLowContainerAvailability: Swift eqiad container availability low - https://wikitech.wikimedia.org/wiki/Swift/How_To - https://grafana.wikimedia.org/d/OPgmB1Eiz/swift?panelId=8&fullscreen&orgId=1&var-DC=eqiad - https://alerts.wikimedia.org/?q=alertname%3DSwiftLowContainerAvailability [13:56:50] (03CR) 10Blake: "recheck" [alerts] - 10https://gerrit.wikimedia.org/r/1307787 (https://phabricator.wikimedia.org/T430290) (owner: 10Blake) [13:56:59] (03PS2) 10Jforrester: wikifunctions: Raise executor process count for JS (but not Py) to 15 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306941 [13:56:59] (03PS1) 10Jforrester: wikifunctions: Upgrade evaluators from 2026-06-30-213833 to 2026-07-08-113600 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308647 (https://phabricator.wikimedia.org/T412768) [13:57:07] (03PS1) 10Jforrester: wikifunctions: Upgrade orchestrator from 2026-07-01-115510 to 2026-07-07-185549 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308648 (https://phabricator.wikimedia.org/T282918) [13:57:39] !log atsuko@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage [13:57:42] klausman@cumin2002 migrate-service-ipip (PID 637348) is awaiting input [13:58:11] !log cwilliams@cumin1003 START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet [13:59:43] !log klausman@cumin2002 START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs [14:00:05] Deploy window Wikifunctions Services UTC Afternoon (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260708T1400) [14:00:45] !log klausman@cumin2002 END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs [14:00:45] !log klausman@cumin2002 END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@eqiad [14:00:58] (03CR) 10Blake: [C:03+2] mw-pretrain: Remove default quota and ranges. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308161 (https://phabricator.wikimedia.org/T427668) (owner: 10Blake) [14:01:19] (03CR) 10Jforrester: [C:03+2] wikifunctions: Upgrade evaluators from 2026-06-30-213833 to 2026-07-08-113600 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308647 (https://phabricator.wikimedia.org/T412768) (owner: 10Jforrester) [14:02:15] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1001.eqiad.wmnet [14:04:55] FIRING: [2x] SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [14:05:18] (03Merged) 10jenkins-bot: renumber-node: Add an option to run homer inline. [cookbooks] - 10https://gerrit.wikimedia.org/r/1308595 (https://phabricator.wikimedia.org/T431443) (owner: 10Blake) [14:05:32] !log atsuko@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1118.eqiad.wmnet with OS trixie [14:07:09] !log cwilliams@cumin1003 START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet [14:08:49] !log elukey@cumin1003 START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet [14:10:14] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet [14:10:38] (03PS1) 10Brouberol: cloudnative-pg-crds: add missing Database CRD [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308649 [14:10:47] (03Merged) 10jenkins-bot: mw-pretrain: Remove default quota and ranges. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308161 (https://phabricator.wikimedia.org/T427668) (owner: 10Blake) [14:12:41] (03Merged) 10jenkins-bot: wikifunctions: Upgrade evaluators from 2026-06-30-213833 to 2026-07-08-113600 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308647 (https://phabricator.wikimedia.org/T412768) (owner: 10Jforrester) [14:13:07] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet [14:13:20] (03PS12) 10Federico Ceratto: sre.mysql: split pool/depool [cookbooks] - 10https://gerrit.wikimedia.org/r/1295480 (https://phabricator.wikimedia.org/T422361) [14:14:21] !log jforrester@deploy1003 helmfile [staging] START helmfile.d/services/wikifunctions: apply [14:14:39] !log sukhe@cumin1003 END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox) [14:14:48] !log fceratto@cumin1003 START - Cookbook sre.mysql.pool pool db2228: Pool test [14:15:08] !log jforrester@deploy1003 helmfile [staging] DONE helmfile.d/services/wikifunctions: apply [14:15:38] !log jforrester@deploy1003 helmfile [codfw] START helmfile.d/services/wikifunctions: apply [14:16:08] !log jforrester@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply [14:16:17] (03PS2) 10Brouberol: cloudnative-pg-crds: add missing Database CRD [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308649 (https://phabricator.wikimedia.org/T424493) [14:16:17] !log jforrester@deploy1003 helmfile [eqiad] START helmfile.d/services/wikifunctions: apply [14:16:36] (03CR) 10Jforrester: [C:03+2] wikifunctions: Upgrade orchestrator from 2026-07-01-115510 to 2026-07-07-185549 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308648 (https://phabricator.wikimedia.org/T282918) (owner: 10Jforrester) [14:16:46] !log jforrester@deploy1003 helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply [14:17:35] (03PS1) 10Scott French: wmnet: Revert codfw, eqsin, ulsfo etcd clients to codfw [dns] - 10https://gerrit.wikimedia.org/r/1308651 (https://phabricator.wikimedia.org/T430909) [14:18:46] !log cwilliams@cumin1003 START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet [14:18:46] (03PS1) 10Scott French: hieradata: Revert codfw PyBals at codfw etcd [puppet] - 10https://gerrit.wikimedia.org/r/1308653 (https://phabricator.wikimedia.org/T430909) [14:19:00] (03Merged) 10jenkins-bot: wikifunctions: Upgrade orchestrator from 2026-07-01-115510 to 2026-07-07-185549 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308648 (https://phabricator.wikimedia.org/T282918) (owner: 10Jforrester) [14:19:12] !log cwilliams@cumin1003 END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet [14:19:12] (03CR) 10Scott French: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308653 (https://phabricator.wikimedia.org/T430909) (owner: 10Scott French) [14:19:23] (03CR) 10CI reject: [V:04-1] hieradata: Revert codfw PyBals at codfw etcd [puppet] - 10https://gerrit.wikimedia.org/r/1308653 (https://phabricator.wikimedia.org/T430909) (owner: 10Scott French) [14:19:32] !log installing librabbitmq security updates [14:19:33] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:19:55] !log blake@deploy1003 helmfile [codfw] START helmfile.d/admin 'apply'. [14:20:23] !log jforrester@deploy1003 helmfile [staging] START helmfile.d/services/wikifunctions: apply [14:20:37] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/admin 'apply'. [14:20:43] !log jforrester@deploy1003 helmfile [staging] DONE helmfile.d/services/wikifunctions: apply [14:21:19] !log jforrester@deploy1003 helmfile [codfw] START helmfile.d/services/wikifunctions: apply [14:22:00] !log jforrester@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply [14:22:02] !log atsuko@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1119.eqiad.wmnet with OS trixie [14:22:08] !log jforrester@deploy1003 helmfile [eqiad] START helmfile.d/services/wikifunctions: apply [14:22:44] !log jforrester@deploy1003 helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply [14:22:51] (03PS1) 10JavierMonton: streams: webrequest - pageview - trending [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308656 (https://phabricator.wikimedia.org/T430675) [14:23:40] (03CR) 10Jforrester: [C:03+2] wikifunctions: Raise executor process count for JS (but not Py) to 15 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306941 (owner: 10Jforrester) [14:27:00] (03CR) 10Brouberol: [C:03+2] cloudnative-pg-crds: add missing Database CRD [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308649 (https://phabricator.wikimedia.org/T424493) (owner: 10Brouberol) [14:27:06] (03CR) 10Scott French: [C:03+2] wmnet: Revert codfw, eqsin, ulsfo etcd clients to codfw [dns] - 10https://gerrit.wikimedia.org/r/1308651 (https://phabricator.wikimedia.org/T430909) (owner: 10Scott French) [14:27:28] !log swfrench@dns1004 START - running authdns-update [14:28:05] (03PS2) 10Ssingh: hieradata: Revert codfw PyBals at codfw etcd [puppet] - 10https://gerrit.wikimedia.org/r/1308653 (https://phabricator.wikimedia.org/T430909) (owner: 10Scott French) [14:28:54] (03Merged) 10jenkins-bot: wikifunctions: Raise executor process count for JS (but not Py) to 15 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306941 (owner: 10Jforrester) [14:29:00] !log installing jackson-core security updates [14:29:01] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:29:28] !log swfrench@dns1004 END - running authdns-update [14:30:01] !log jforrester@deploy1003 helmfile [staging] START helmfile.d/services/wikifunctions: apply [14:30:04] Deploy window Wikifunctions Services UTC Afternoon (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260708T1400) [14:30:04] Deploy window Test Kitchen Experiment Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260708T1430) [14:30:06] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test [14:30:30] !log jforrester@deploy1003 helmfile [staging] DONE helmfile.d/services/wikifunctions: apply [14:30:38] !log jforrester@deploy1003 helmfile [codfw] START helmfile.d/services/wikifunctions: apply [14:30:56] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. [14:31:00] !log jforrester@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply [14:31:05] !log jforrester@deploy1003 helmfile [eqiad] START helmfile.d/services/wikifunctions: apply [14:31:13] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. [14:31:23] !log jforrester@deploy1003 helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply [14:31:39] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie [14:31:53] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12100459 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host thanos-be1007.eqiad.wmnet with OS trixie [14:32:01] (03CR) 10Filippo Giunchedi: [C:03+2] openstack: deprecate icinga check-neutron-conntrack [puppet] - 10https://gerrit.wikimedia.org/r/1302764 (https://phabricator.wikimedia.org/T328502) (owner: 10Filippo Giunchedi) [14:32:26] (03CR) 10Ssingh: [C:03+2] hieradata: Revert codfw PyBals at codfw etcd [puppet] - 10https://gerrit.wikimedia.org/r/1308653 (https://phabricator.wikimedia.org/T430909) (owner: 10Scott French) [14:32:38] !log cwilliams@cumin1003 START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet [14:32:53] !log cwilliams@cumin1003 END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet [14:34:14] !log switched codfw, eqsin, ulsfo etcd client SRV records back to codfw - T430909 [14:34:17] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:34:17] T430909: codfw: rack B3 maintenance - https://phabricator.wikimedia.org/T430909 [14:35:58] !log restart pybal on lvs2014 to revert back to conf2004 [14:35:59] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:37:48] !log restart pybal on lvs2013 to revert back to conf2004 [14:37:50] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:39:02] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.4 point update - https://phabricator.wikimedia.org/T420240#12100520 (10MoritzMuehlenhoff) [14:39:07] !log mvernon@cumin2003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1007.eqiad.wmnet with OS trixie [14:39:20] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12100523 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host thanos-be1007.eqiad.wmnet with OS trixie execu... [14:39:21] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.5 point update - https://phabricator.wikimedia.org/T427072#12100524 (10MoritzMuehlenhoff) [14:39:45] !log mvernon@cumin2003 START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [14:40:18] !log sudo cumin -b1 -s120 "P{lvs2011*} or P{lvs2012*}" "systemctl restart pybal.service" [14:40:21] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:41:00] !log uninstalling dhcpcd-base from trixie hosts which still have it installed T414341 [14:41:02] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:41:03] T414341: Avoid dhcpcd-base on trixie hosts - https://phabricator.wikimedia.org/T414341 [14:43:41] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. [14:43:57] (03PS2) 10Hashar: ci: add buildx on production Jenkins agents [puppet] - 10https://gerrit.wikimedia.org/r/1308659 (https://phabricator.wikimedia.org/T431582) [14:43:57] (03CR) 10Hashar: "@dduvall@wikimedia.org yesterday we moved the Jenkins production agents to Trixie with the Docker package from Debian and the build fails " [puppet] - 10https://gerrit.wikimedia.org/r/1308659 (https://phabricator.wikimedia.org/T431582) (owner: 10Hashar) [14:44:03] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. [14:44:11] 10ops-codfw, 06DC-Ops: move wikikube-worker2044 in same rack - https://phabricator.wikimedia.org/T431585 (10Jhancock.wm) 03NEW [14:46:48] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. [14:47:16] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. [14:49:39] !log mvernon@cumin2003 END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [14:49:47] PROBLEM - PyBal connections to etcd on lvs2011 is CRITICAL: CRITICAL: 0 connections established with conf2004.codfw.wmnet:4001 (min=12) https://wikitech.wikimedia.org/wiki/PyBal [14:49:56] uh? [14:50:47] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. [14:51:05] 10ops-codfw, 06SRE, 06DC-Ops: upgrade selected servers from 1G to 10G - https://phabricator.wikimedia.org/T429631#12100652 (10Jhancock.wm) [14:51:15] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. [14:52:28] !log restarted ulsfo confds, confirmed now connected to codfw backends except those using wikimedia.org SRV record - T430909 [14:52:31] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:52:32] T430909: codfw: rack B3 maintenance - https://phabricator.wikimedia.org/T430909 [14:54:17] RECOVERY - PyBal connections to etcd on lvs2011 is OK: OK: 12 connections established with conf2004.codfw.wmnet:4001 (min=12) https://wikitech.wikimedia.org/wiki/PyBal [14:54:25] that's more like it [14:55:19] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie [14:55:36] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12100663 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host thanos-be1007.eqiad.wmnet with OS trixie [14:55:56] (03PS1) 10Brouberol: cloudnative-pg: upgrade the image and restore watched namespaces [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308664 (https://phabricator.wikimedia.org/T424493) [14:58:47] (03CR) 10Hashar: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308194 (https://phabricator.wikimedia.org/T418521) (owner: 10Dzahn) [14:58:51] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.3 point update - https://phabricator.wikimedia.org/T414179#12100677 (10MoritzMuehlenhoff) [14:59:00] PROBLEM - PyBal connections to etcd on lvs2012 is CRITICAL: CRITICAL: 0 connections established with conf2004.codfw.wmnet:4001 (min=6) https://wikitech.wikimedia.org/wiki/PyBal [14:59:05] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.3 point update - https://phabricator.wikimedia.org/T414179#12100685 (10MoritzMuehlenhoff) 05Open→03Resolved a:03MoritzMuehlenhoff All done [14:59:14] (03CR) 10Btullis: [C:03+1] "Nice." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308664 (https://phabricator.wikimedia.org/T424493) (owner: 10Brouberol) [14:59:42] !log elukey@cumin1003 START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet [14:59:43] !log elukey@cumin1003 END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet [14:59:48] 06SRE, 06Infrastructure-Foundations: Avoid dhcpcd-base on trixie hosts - https://phabricator.wikimedia.org/T414341#12100690 (10MoritzMuehlenhoff) 05Open→03Resolved a:03MoritzMuehlenhoff We stopped installing dhcpcd-base on new trixie installations back in May and I've just removed it from the hosts w... [15:00:44] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, July 08 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-it" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308183 (https://phabricator.wikimedia.org/T431257) (owner: 10Clare Ming) [15:02:20] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.4 point update - https://phabricator.wikimedia.org/T420240#12100719 (10MoritzMuehlenhoff) [15:02:47] !log cwilliams@cumin1003 START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet [15:02:51] (03PS1) 10Elukey: sre.hosts.bmc-user-mgmt: fix netbox data access [cookbooks] - 10https://gerrit.wikimedia.org/r/1308665 [15:03:01] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1002.eqiad.wmnet [15:03:04] !log restarted eqsin, codfw confds - T430909 [15:03:07] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [15:03:08] T430909: codfw: rack B3 maintenance - https://phabricator.wikimedia.org/T430909 [15:03:15] !log sukhe@cumin1003 START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo (T430909) [15:03:30] !log elukey@cumin1003 START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet [15:03:31] !log restarted navtiming on webperf2003 - T430909 [15:03:31] !log elukey@cumin1003 END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet [15:03:34] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [15:03:42] !log blake@deploy1003 helmfile [codfw] START helmfile.d/admin 'apply'. [15:03:44] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/admin 'apply'. [15:04:08] (03PS2) 10Elukey: sre.hosts.bmc-user-mgmt: fix netbox data access [cookbooks] - 10https://gerrit.wikimedia.org/r/1308665 [15:04:14] !log elukey@cumin1003 START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet [15:04:14] !log blake@deploy1003 helmfile [codfw] START helmfile.d/admin 'apply'. [15:04:19] !log elukey@cumin1003 END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet [15:04:55] (03PS1) 10Btullis: datahub-next: use the uid claim for the DataHub username [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308666 (https://phabricator.wikimedia.org/T309382) [15:05:04] !log sukhe@cumin1003 END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo (T430909) [15:05:23] (03CR) 10Brouberol: [C:03+2] cloudnative-pg: upgrade the image and restore watched namespaces [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308664 (https://phabricator.wikimedia.org/T424493) (owner: 10Brouberol) [15:05:26] !log cwilliams@cumin1003 START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet [15:05:44] #wikimedia-serviceops [15:06:26] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/admin 'apply'. [15:06:27] !log sukhe@cumin1003 START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin (T430909) [15:06:37] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet [15:06:59] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/admin 'apply'. [15:08:11] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/admin 'apply'. [15:08:11] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. [15:08:35] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. [15:09:54] 14SRE-Sprint-Week-Sustainability-March2023, 06Traffic, 13Patch-For-Review, 07Sustainability (Incident Followup): Experiment with single backend CDN nodes - https://phabricator.wikimedia.org/T288106#12100744 (10BBlack) Additional metrics that would be informative: * Differential of network bandwidth on the... [15:10:41] !log sukhe@cumin1003 END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin (T430909) [15:10:47] T430909: codfw: rack B3 maintenance - https://phabricator.wikimedia.org/T430909 [15:13:57] (03CR) 10Ottomata: streams: webrequest - pageview - trending (035 comments) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308656 (https://phabricator.wikimedia.org/T430675) (owner: 10JavierMonton) [15:14:08] (03PS2) 10Btullis: datahub-next: use the uid claim for the DataHub username [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308666 (https://phabricator.wikimedia.org/T309382) [15:14:13] !log fceratto@cumin1003 START - Cookbook sre.mysql.depool depool db2228: Depool test [15:14:33] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: Depool test [15:14:40] !log cwilliams@cumin1003 START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet [15:14:42] !log fceratto@cumin1003 START - Cookbook sre.mysql.pool pool db2228: Pool test [15:15:01] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/mw-pretrain: apply [15:15:11] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage [15:15:22] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply [15:15:31] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply [15:15:34] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet [15:15:43] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply [15:17:04] (03PS3) 10Elukey: sre.hosts.bmc-user-mgmt: fix netbox data access [cookbooks] - 10https://gerrit.wikimedia.org/r/1308665 [15:17:04] (03PS1) 10Elukey: sre.hosts.provision: add a workaround for the ipmi user [cookbooks] - 10https://gerrit.wikimedia.org/r/1308668 [15:18:59] (03CR) 10Btullis: [C:03+2] datahub-next: use the uid claim for the DataHub username [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308666 (https://phabricator.wikimedia.org/T309382) (owner: 10Btullis) [15:19:58] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage [15:21:36] (03Merged) 10jenkins-bot: datahub-next: use the uid claim for the DataHub username [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308666 (https://phabricator.wikimedia.org/T309382) (owner: 10Btullis) [15:23:02] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply [15:23:23] (03CR) 10JHathaway: [C:03+1] sre.hosts.provision: add a workaround for the ipmi user [cookbooks] - 10https://gerrit.wikimedia.org/r/1308668 (owner: 10Elukey) [15:23:30] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply [15:23:59] (03CR) 10JHathaway: [C:03+1] sre.hosts.bmc-user-mgmt: fix netbox data access [cookbooks] - 10https://gerrit.wikimedia.org/r/1308665 (owner: 10Elukey) [15:26:02] 06SRE, 10observability: monitor postgresql replication status - https://phabricator.wikimedia.org/T116580#12100805 (10hnowlan) 05Open→03Resolved a:03hnowlan Replication monitoring has been in place for quite a while (needs to be moved to prometheus though, shhhh) [15:30:00] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test [15:35:38] (03PS1) 10Btullis: Revert "datahub-next: use the uid claim for the DataHub username" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308672 (https://phabricator.wikimedia.org/T309382) [15:35:52] 10ops-drmrs, 06SRE, 06Infrastructure-Foundations: FY2627 drmrs: upgrade infra links from 40G to 100G - https://phabricator.wikimedia.org/T431562#12100846 (10wiki_willy) Hey @cmooney - thanks for getting this on the radar. Totally makes sense to me, so I think we can just do this upgrade during the same time... [15:36:28] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1007.eqiad.wmnet with OS trixie [15:36:38] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12100863 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host thanos-be1007.eqiad.wmnet with OS trixie compl... [15:38:11] (03CR) 10Btullis: [C:03+2] Revert "datahub-next: use the uid claim for the DataHub username" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308672 (https://phabricator.wikimedia.org/T309382) (owner: 10Btullis) [15:39:11] !log elukey@cumin1003 START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [15:39:54] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [15:40:30] (03Merged) 10jenkins-bot: Revert "datahub-next: use the uid claim for the DataHub username" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308672 (https://phabricator.wikimedia.org/T309382) (owner: 10Btullis) [15:40:39] (03CR) 10Elukey: [C:03+2] sre.hosts.provision: add a workaround for the ipmi user [cookbooks] - 10https://gerrit.wikimedia.org/r/1308668 (owner: 10Elukey) [15:40:47] (03CR) 10Elukey: [C:03+2] sre.hosts.bmc-user-mgmt: fix netbox data access [cookbooks] - 10https://gerrit.wikimedia.org/r/1308665 (owner: 10Elukey) [15:40:48] (03CR) 10JavierMonton: streams: webrequest - pageview - trending (035 comments) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308656 (https://phabricator.wikimedia.org/T430675) (owner: 10JavierMonton) [15:41:45] FIRING: SwiftLowContainerAvailability: Swift eqiad container availability low - https://wikitech.wikimedia.org/wiki/Swift/How_To - https://grafana.wikimedia.org/d/OPgmB1Eiz/swift?panelId=8&fullscreen&orgId=1&var-DC=eqiad - https://alerts.wikimedia.org/?q=alertname%3DSwiftLowContainerAvailability [15:41:45] FIRING: SwiftLowObjectAvailability: Swift eqiad object availability low - https://wikitech.wikimedia.org/wiki/Swift/How_To - https://grafana.wikimedia.org/d/OPgmB1Eiz/swift?panelId=8&fullscreen&orgId=1&var-DC=eqiad - https://alerts.wikimedia.org/?q=alertname%3DSwiftLowObjectAvailability [15:42:02] !log jasmine@cumin2002 START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet [15:42:07] !log jasmine@cumin2002 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet [15:42:12] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply [15:42:26] (03CR) 10CWilliams: sre.mysql: split pool/depool (031 comment) [cookbooks] - 10https://gerrit.wikimedia.org/r/1295480 (https://phabricator.wikimedia.org/T422361) (owner: 10Federico Ceratto) [15:42:40] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply [15:42:43] !log jasmine@cumin2002 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet [15:42:44] !log jasmine@cumin2002 END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1164.eqiad.wmnet [15:44:26] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12100897 (10MatthewVernon) [15:48:47] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host cirrussearch1095.eqiad.wmnet with OS trixie [15:50:08] !log jasmine@cumin2002 START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet [15:50:13] !log jasmine@cumin2002 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet [15:50:16] !log jasmine@cumin2002 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet [15:50:42] !log jasmine@cumin2002 START - Cookbook sre.hosts.reimage for host wikikube-worker1164.eqiad.wmnet with OS trixie [15:51:16] !log jasmine@cumin2002 START - Cookbook sre.hosts.move-vlan for host wikikube-worker1164 [15:51:48] !log jasmine@cumin2002 START - Cookbook sre.dns.netbox [15:52:02] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host cirrussearch1093.eqiad.wmnet with OS trixie [15:54:42] 06SRE, 06Infrastructure-Foundations, 10netops: GSHUT (and other?) community matching/actions not working on SR-Linux - https://phabricator.wikimedia.org/T430810#12100957 (10cmooney) Found a few things testing this in container-lab: # We need to set "policy-result accept" in each statement or it will continu... [15:54:54] (03CR) 10Dzahn: "curious if there is a reason to limit it to production and not just do it in any realm, cloud and prod if on trixie or newer. unless we ha" [puppet] - 10https://gerrit.wikimedia.org/r/1308659 (https://phabricator.wikimedia.org/T431582) (owner: 10Hashar) [15:56:45] RESOLVED: SwiftLowContainerAvailability: Swift eqiad container availability low - https://wikitech.wikimedia.org/wiki/Swift/How_To - https://grafana.wikimedia.org/d/OPgmB1Eiz/swift?panelId=8&fullscreen&orgId=1&var-DC=eqiad - https://alerts.wikimedia.org/?q=alertname%3DSwiftLowContainerAvailability [15:56:45] RESOLVED: SwiftLowObjectAvailability: Swift eqiad object availability low - https://wikitech.wikimedia.org/wiki/Swift/How_To - https://grafana.wikimedia.org/d/OPgmB1Eiz/swift?panelId=8&fullscreen&orgId=1&var-DC=eqiad - https://alerts.wikimedia.org/?q=alertname%3DSwiftLowObjectAvailability [15:56:47] !log jasmine@cumin2002 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" [15:56:55] !log jasmine@cumin2002 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002" [15:56:55] !log jasmine@cumin2002 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [15:56:55] !log jasmine@cumin2002 START - Cookbook sre.dns.wipe-cache wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [15:56:59] !log jasmine@cumin2002 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [15:57:00] !log jasmine@cumin2002 START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1164 [15:57:06] (03CR) 10BCornwall: [V:03+1 C:03+2] cache::text: raise global_auth_ratelimit limits [puppet] - 10https://gerrit.wikimedia.org/r/1307220 (owner: 10BCornwall) [15:57:28] !log jasmine@cumin2002 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1164 [15:57:29] !log jasmine@cumin2002 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1164 [15:58:39] !log fceratto@cumin1003 START - Cookbook sre.mysql.depool depool pc1023: Depool test [15:58:40] !log fceratto@cumin1003 START - Cookbook sre.mysql.parsercache [15:58:49] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) [15:58:49] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test [15:59:00] !log fceratto@cumin1003 START - Cookbook sre.mysql.pool pool pc1023: Pool test [15:59:00] !log fceratto@cumin1003 START - Cookbook sre.mysql.parsercache [15:59:14] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) [15:59:14] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test [16:02:20] PROBLEM - Check if Pybal has been restarted after pybal.conf was changed on lvs2012 is CRITICAL: CRITICAL: Service pybal.service has not been restarted after /etc/pybal/pybal.conf was changed (gt 1h). https://wikitech.wikimedia.org/wiki/PyBal%23Pybal_service_has_not_been_restarted [16:03:17] huh [16:06:40] FIRING: SystemdUnitFailed: fetch-rings-thanos.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [16:08:50] !log bking@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage [16:09:06] !log elukey@cumin1003 START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet [16:09:11] !log elukey@cumin1003 END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet [16:09:22] FIRING: [2x] JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:10:25] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.4 point update - https://phabricator.wikimedia.org/T420240#12101019 (10MoritzMuehlenhoff) [16:10:54] (03PS1) 10Btullis: Fix the dummy secrets for the new an-test hosts [labs/private] - 10https://gerrit.wikimedia.org/r/1308679 (https://phabricator.wikimedia.org/T406746) [16:11:05] (03CR) 10Btullis: [V:03+2 C:03+2] Fix the dummy secrets for the new an-test hosts [labs/private] - 10https://gerrit.wikimedia.org/r/1308679 (https://phabricator.wikimedia.org/T406746) (owner: 10Btullis) [16:12:40] (03CR) 10Btullis: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308149 (https://phabricator.wikimedia.org/T406746) (owner: 10Btullis) [16:14:22] RESOLVED: [2x] JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:15:08] !log bking@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage [16:16:29] !log jasmine@cumin2002 START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage [16:16:44] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.5 point update - https://phabricator.wikimedia.org/T427072#12101039 (10MoritzMuehlenhoff) [16:18:51] !log bking@cumin2003 END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage [16:19:58] 10ops-eqiad, 06SRE, 06DC-Ops: Degraded RAID on backup1019 - https://phabricator.wikimedia.org/T431367#12101075 (10VRiley-WMF) Part has been shipped and awaiting arrival. [16:23:00] jouncebot: nowandnext [16:23:00] No deployments scheduled for the next 0 hour(s) and 36 minute(s) [16:23:00] In 0 hour(s) and 36 minute(s): MediaWiki infrastructure (UTC late) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260708T1700) [16:23:36] !log jasmine@cumin2002 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage [16:24:01] hnowlan: I'm going to start the process of deploying the patch now 😅 [16:24:04] (03PS1) 10Elukey: redfish: don't assume that the Allow header is always present [software/spicerack] - 10https://gerrit.wikimedia.org/r/1308680 (https://phabricator.wikimedia.org/T426180) [16:24:24] (03CR) 10BCornwall: [V:03+1 C:03+2] varnish: remove ancient Noise rule from text-frontend VCL [puppet] - 10https://gerrit.wikimedia.org/r/1215329 (owner: 10Dzahn) [16:25:00] (03CR) 10Ladsgroup: [C:03+2] swift: Load the file async [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1308156 (https://phabricator.wikimedia.org/T333445) (owner: 10Ladsgroup) [16:25:19] Amir1: cool :) [16:25:29] (03PS1) 10Btullis: datahub: disable REST API authorization [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308682 (https://phabricator.wikimedia.org/T402408) [16:26:25] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply [16:26:51] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply} [16:27:27] 10ops-codfw, 06SRE, 06DC-Ops: move wikikube-worker servers in same rack - https://phabricator.wikimedia.org/T431585#12101092 (10Jhancock.wm) [16:27:57] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage [16:29:21] Amir1: lmk when you're done please [16:29:32] sure thing [16:29:57] thanks <3 [16:35:04] (03Merged) 10jenkins-bot: swift: Load the file async [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1308156 (https://phabricator.wikimedia.org/T333445) (owner: 10Ladsgroup) [16:35:43] PROBLEM - Check unit status of push_cross_cluster_settings_9400 on cirrussearch1093 is CRITICAL: CRITICAL: Status of the systemd unit push_cross_cluster_settings_9400 https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [16:40:37] 10ops-codfw, 06SRE, 06DC-Ops: move wikikube-worker servers in same rack - https://phabricator.wikimedia.org/T431585#12101126 (10Jhancock.wm) [16:43:09] !log jasmine@cumin2002 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1164.eqiad.wmnet with OS trixie [16:44:10] the post-merge stuff is a bloodbath :( https://integration.wikimedia.org/zuul/ [16:44:39] FIRING: CirrusSearchNodeIndexingNotIncreasing: Elasticsearch instance cirrussearch1093-production-search-eqiad is not indexing - https://wikitech.wikimedia.org/wiki/Search/Elasticsearch_Administration#Indexing_hung_and_not_making_progress - https://grafana.wikimedia.org/d/JLK3I_siz/elasticsearch-indexing?orgId=1&from=now-3d&to=now&viewPanel=57 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchNodeIndexingNotIncreasing [16:44:40] Raine: wanna do your change. Mine is taking way too long. I need a break too [16:45:09] 10ops-eqiad, 06Data-Persistence, 06DBA, 06DC-Ops, and 3 others: es7 primary (es1039.eqiad.wmnet) crashed - https://phabricator.wikimedia.org/T430764#12101132 (10VRiley-WMF) Upone looking at a similar issue that was resolved, I have just updated the backplan firmware with I believe may have fixed the other... [16:45:21] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1093.eqiad.wmnet with OS trixie [16:46:36] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1095.eqiad.wmnet with OS trixie [16:46:42] Amir1: I'm going to be a while [16:46:50] switching deployment server [16:46:56] over 1h [16:46:57] is that okay? [16:47:43] yeah yeah [16:48:02] okay, off I go [16:48:08] <3 [16:49:37] RECOVERY - Check unit status of push_cross_cluster_settings_9400 on cirrussearch1093 is OK: OK: Status of the systemd unit push_cross_cluster_settings_9400 https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [16:49:42] !log kamila@deploy1003 Locking from deployment [MediaWiki]: switching deployment server [16:49:57] 10ops-eqiad, 06Data-Persistence, 06DBA, 06DC-Ops, and 3 others: es7 primary (es1039.eqiad.wmnet) crashed - https://phabricator.wikimedia.org/T430764#12101169 (10VRiley-WMF) [16:50:26] FIRING: [5x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-1/0/2 (Transport: cr1-eqiad:et-1/1/2 (Arelion, IC-374549) {#12267}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [16:51:49] 06SRE, 06Infrastructure-Foundations, 10netops: GSHUT (and other?) community matching/actions not working on SR-Linux - https://phabricator.wikimedia.org/T430810#12101183 (10cmooney) Also discussed this on the SR Linux discord and got some good info. My observation in containerlab was that when we set the lo... [16:52:52] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie [16:53:44] !log kamila@deploy1003 Unlocked for deployment [MediaWiki]: switching deployment server (duration: 04m 02s) [16:53:47] jasmine@cumin2002 renumber-node (PID 663124) is awaiting input [16:54:08] !log kamila@deploy1003 Locking from deployment [MediaWiki]: switching deployment server [16:55:17] (03CR) 10BCornwall: [C:03+2] cache::haproxy: Enable jwt for upload clusters [puppet] - 10https://gerrit.wikimedia.org/r/1305923 (https://phabricator.wikimedia.org/T400238) (owner: 10BCornwall) [16:56:03] !log kamila@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover [16:56:12] (03PS5) 10BCornwall: cache::haproxy: Enable jwt for upload clusters [puppet] - 10https://gerrit.wikimedia.org/r/1305923 (https://phabricator.wikimedia.org/T400238) [16:56:12] (03CR) 10Kamila Součková: [C:03+2] wmnet: update deployment CNAME record to deploy2003 [dns] - 10https://gerrit.wikimedia.org/r/1308134 (https://phabricator.wikimedia.org/T423714) (owner: 10Kamila Součková) [16:56:58] (03CR) 10Kamila Součková: [C:03+2] hiera: update deployment server to deploy2003 [puppet] - 10https://gerrit.wikimedia.org/r/1308147 (https://phabricator.wikimedia.org/T423714) (owner: 10Kamila Součková) [16:57:14] (03CR) 10BCornwall: [V:03+2 C:03+2] "Rebased" [puppet] - 10https://gerrit.wikimedia.org/r/1305923 (https://phabricator.wikimedia.org/T400238) (owner: 10BCornwall) [16:59:39] RESOLVED: CirrusSearchNodeIndexingNotIncreasing: Elasticsearch instance cirrussearch1093-production-search-eqiad is not indexing - https://wikitech.wikimedia.org/wiki/Search/Elasticsearch_Administration#Indexing_hung_and_not_making_progress - https://grafana.wikimedia.org/d/JLK3I_siz/elasticsearch-indexing?orgId=1&from=now-3d&to=now&viewPanel=57 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchNodeIndexingNotIncreasing [17:00:05] Deploy window MediaWiki infrastructure (UTC late) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260708T1700) [17:04:18] !log jasmine@cumin2002 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1164.eqiad.wmnet [17:04:21] !log jasmine@cumin2002 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1164.eqiad.wmnet [17:04:24] !log jasmine@cumin2002 END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1164.eqiad.wmnet [17:04:43] (03PS1) 10WMDE-Fisch: Enable sub-references on more group2 wikis (batch2) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308687 (https://phabricator.wikimedia.org/T430941) [17:05:39] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Thursday, July 09 UTC morning backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308687 (https://phabricator.wikimedia.org/T430941) (owner: 10WMDE-Fisch) [17:06:32] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, 06cloud-services-team (FY2025/2026-Q3-Q4): Ensure cloudvirt capacity is more evenly spread out among racks - https://phabricator.wikimedia.org/T424658#12101239 (10VRiley-WMF) @fgiunchedi I was curious to know if there is an estimated date for these moves? Will w... [17:07:46] PROBLEM - check if authdns-update was run after a change was merged to operations/dns.git on dns7002 is CRITICAL: Local zone files are NOT in sync with operations/dns.git (SHA: local is eb98ad1c9d808e845cd53c892dccb69d29409df7, dns.git is 6a4aba488be10973d4c2188880c341708e787e41) https://wikitech.wikimedia.org/wiki/DNS%23authdns_update_run [17:07:56] !log bking@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage [17:08:16] (03PS1) 10Jdlrobson: Enable section share on Farsi Wikipedia [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308671 (https://phabricator.wikimedia.org/T431514) [17:09:29] !log kamila@dns1005 START - running authdns-update [17:10:29] (03PS1) 10Reedy: CommonSettings: Remove $wgScoreSafeMode [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308688 (https://phabricator.wikimedia.org/T428484) [17:10:52] (03CR) 10Reedy: [C:04-2] "Needs to wait for 1.47.0-wmf.11" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308688 (https://phabricator.wikimedia.org/T428484) (owner: 10Reedy) [17:11:18] !log kamila@dns1005 END - running authdns-update [17:12:47] RECOVERY - check if authdns-update was run after a change was merged to operations/dns.git on dns7002 is OK: Local zone files and operations/dns.git are in sync https://wikitech.wikimedia.org/wiki/DNS%23authdns_update_run [17:15:20] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage [17:16:15] !log kamila@deploy1003 Unlocked for deployment [MediaWiki]: switching deployment server (duration: 22m 07s) [17:16:31] woohoo, no “connection is not using a post-quantum key exchange algorithm” warning on deploy2003 \o/ [17:18:12] !log kamila@cumin1003 START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet [17:19:04] Lucas_WMDE: whoops sorry, I just murdered it :D not done with it [17:19:04] (03PS1) 10Cathal Mooney: Nokia SR-Linux: fix issues with BGP policies matching communities [homer/public] - 10https://gerrit.wikimedia.org/r/1308689 (https://phabricator.wikimedia.org/T430810) [17:19:09] but yeah, no warning is good :D [17:19:24] Raine: sorry for being nosy and already logging onto the server :P [17:19:32] (I didn’t have anything to do on there, all good ^^) [17:19:52] (and yay for debian bookworm) [17:19:58] <3 [17:20:06] noooo how am I supposed to use my quantum computer to break into the deploy servers now :( [17:20:46] perryprog: of course that breaks someone's workflow https://xkcd.com/1172/ [17:20:54] ;-) [17:21:28] I was more thinking https://xkcd.com/538/ to be honest... [17:21:56] (03PS1) 10AOkoth: hiera: add host specific values for new phab host [puppet] - 10https://gerrit.wikimedia.org/r/1308690 (https://phabricator.wikimedia.org/T424055) [17:22:13] perryprog: also true :D [17:23:42] (03CR) 10Hashar: "I have amended your change in that regard :] This was done by John Bond at the time (T242910)." [puppet] - 10https://gerrit.wikimedia.org/r/1308194 (https://phabricator.wikimedia.org/T418521) (owner: 10Dzahn) [17:24:21] perryprog: don’t wrench react the people with deploy access :< [17:24:51] (03PS2) 10AOkoth: hiera: add host specific values for new phab host [puppet] - 10https://gerrit.wikimedia.org/r/1308690 (https://phabricator.wikimedia.org/T424055) [17:26:07] (03CR) 10Dzahn: [C:03+1] hiera: add host specific values for new phab host [puppet] - 10https://gerrit.wikimedia.org/r/1308690 (https://phabricator.wikimedia.org/T424055) (owner: 10AOkoth) [17:26:14] Lucas_WMDE: tbh if I wanted server access it'd be easier to just bribe taavi with a train pass or something [17:26:20] (03PS1) 10Kamila Součková: tcpircbot (logmsgbot): replace deploy2002 with deploy2003 [puppet] - 10https://gerrit.wikimedia.org/r/1289019 (https://phabricator.wikimedia.org/T426222) (owner: 10Dzahn) [17:26:20] (03CR) 10Kamila Součková: "oh, that's why scap log is failing! :D thanks again <3" [puppet] - 10https://gerrit.wikimedia.org/r/1289019 (https://phabricator.wikimedia.org/T426222) (owner: 10Dzahn) [17:26:23] (03CR) 10Kamila Součková: [C:03+2] tcpircbot (logmsgbot): replace deploy2002 with deploy2003 [puppet] - 10https://gerrit.wikimedia.org/r/1289019 (https://phabricator.wikimedia.org/T426222) (owner: 10Dzahn) [17:26:32] (03CR) 10Hashar: "PCC fails with:" [puppet] - 10https://gerrit.wikimedia.org/r/1308194 (https://phabricator.wikimedia.org/T418521) (owner: 10Dzahn) [17:28:01] (03PS1) 10Sakshi2: Fix typo in dockerfile.template to imrpove comments for user understanding [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1308691 (https://phabricator.wikimedia.org/T201491) [17:28:03] perryprog: :D [17:28:07] do you object to being !bashed? [17:28:11] nop [17:28:47] get !bash’ed nerd [17:28:48] !bash Lucas_WMDE: tbh if I wanted server access it'd be easier to just bribe taavi with a train pass or something [17:28:48] (03CR) 10Dzahn: "I am ok merging this without compiling it. But one other thing: to follow the process this needs an approval from the group owner. I will " [puppet] - 10https://gerrit.wikimedia.org/r/1308194 (https://phabricator.wikimedia.org/T418521) (owner: 10Dzahn) [17:28:48] taavi: Stored quip at https://bash.toolforge.org/quip/m0rGQp8B1kByGTxAbcHj [17:29:13] * Raine looks the other way because that's not how a responsible SRE should act (please do not try to bribe me with flights that would be mean) [17:29:45] (03CR) 10Dzahn: "@adancy@wikimedia.org You are one of the "group owners" of contint-docker. Can we have your approval to add this group on new dedicated je" [puppet] - 10https://gerrit.wikimedia.org/r/1308194 (https://phabricator.wikimedia.org/T418521) (owner: 10Dzahn) [17:29:53] * Raine (little planes are cute) [17:30:07] Raine: Cirrus Vision SF50 [17:30:23] awwwwwwWWWWW! [17:30:33] as a tall person I disagree about any positive description of a small plane [17:30:46] taavi: well that's a you problem, from my perspective :P [17:30:51] (my perspective is from below :D) [17:31:54] !log kamila@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet [17:32:25] (03Abandoned) 10Btullis: datahub: disable REST API authorization [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308682 (https://phabricator.wikimedia.org/T402408) (owner: 10Btullis) [17:32:29] (03CR) 10AOkoth: [C:03+2] hiera: add host specific values for new phab host [puppet] - 10https://gerrit.wikimedia.org/r/1308690 (https://phabricator.wikimedia.org/T424055) (owner: 10AOkoth) [17:33:02] (03CR) 10Dzahn: "seems good to me! especially since default is no limit at all" [puppet] - 10https://gerrit.wikimedia.org/r/1308081 (https://phabricator.wikimedia.org/T431398) (owner: 10Arnaudb) [17:33:08] (03CR) 10Dzahn: [C:03+1] gerrit: limit the volume of returnable pages [puppet] - 10https://gerrit.wikimedia.org/r/1308081 (https://phabricator.wikimedia.org/T431398) (owner: 10Arnaudb) [17:33:29] (03CR) 10RLazarus: [C:03+1] linked-artifacts: Add service catalog and mesh entries [puppet] - 10https://gerrit.wikimedia.org/r/1308604 (https://phabricator.wikimedia.org/T431108) (owner: 10Clément Goubert) [17:34:46] (03CR) 10Dzahn: [C:03+1] gerrit: remove unused bad_browser reference [puppet] - 10https://gerrit.wikimedia.org/r/1308351 (https://phabricator.wikimedia.org/T431398) (owner: 10Arnaudb) [17:34:56] (03CR) 10RLazarus: [C:03+1] "LGTM pending the Puppet change, thanks!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308609 (https://phabricator.wikimedia.org/T431108) (owner: 10Clément Goubert) [17:35:01] (03CR) 10Dzahn: [C:03+2] gerrit: remove unused bad_browser reference [puppet] - 10https://gerrit.wikimedia.org/r/1308351 (https://phabricator.wikimedia.org/T431398) (owner: 10Arnaudb) [17:35:59] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie [17:36:09] (03CR) 10Ahmon Dancy: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1308194 (https://phabricator.wikimedia.org/T418521) (owner: 10Dzahn) [17:42:10] (03CR) 10Ottomata: streams: webrequest - pageview - trending (033 comments) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308656 (https://phabricator.wikimedia.org/T430675) (owner: 10JavierMonton) [17:42:15] !log kamila@deploy2003 Rolling back deployment [17:42:49] !log kamila@deploy2003 Finished scap sync-world: Test deployment to validate deployment server switchover - T423714 (duration: 19m 50s) [17:42:52] T423714: Upgrade deploy* hosts to Debian Bookworm - https://phabricator.wikimedia.org/T423714 [17:42:57] !log kamila@deploy2003 Started scap sync-world: Test deployment to validate deployment server switchover - T423714 [17:44:28] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1099 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [17:44:28] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1101 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [17:44:28] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1088 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [17:44:28] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1108 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [17:44:28] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1078 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [17:44:29] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1121 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [17:44:29] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1111 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [17:44:30] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1117 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [17:44:34] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1095 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [17:44:38] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1072 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [17:44:42] PROBLEM - ElasticSearch health check for shards on 9643 on search.svc.eqiad.wmnet is CRITICAL: CRITICAL - elasticsearch https://search.svc.eqiad.wmnet:9643/_cluster/health error while fetching: HTTPSConnectionPool(host=search.svc.eqiad.wmnet, port=9643): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [17:44:44] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1081 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [17:44:46] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1102 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [17:44:52] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1073 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: (Connection aborted., RemoteDisconnected(Remote end closed connection without response)) https://wikitech.wikimedia.org/wiki/Search%23Administration [17:44:54] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1075 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [17:44:56] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1090 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [17:44:56] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1086 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [17:44:56] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1092 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [17:44:56] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1083 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [17:44:56] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1097 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [17:44:57] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1084 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [17:44:57] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1123 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [17:44:58] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1118 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Max retries exceeded with url: /_cluster/health (Caused by NewConnectionError(urllib3.connection.HTTPConnection object at 0x7f83f89a5400: Failed to establish a new connection: [Errno 111] Connection refused)) https://wikitec [17:44:58] dia.org/wiki/Search%23Administration [17:44:59] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1110 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [17:44:59] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1116 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [17:45:04] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1087 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [17:45:04] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1069 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [17:45:12] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1122 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [17:45:16] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1085 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [17:45:38] RECOVERY - ElasticSearch health check for shards on 9643 on search.svc.eqiad.wmnet is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: green, timed_out: False, number_of_nodes: 30, number_of_data_nodes: 30, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 5196, relocating_shards: 0, initializing_shards: 0, unassigned_shards: [17:45:38] ed_unassigned_shards: 0, number_of_pending_tasks: 1444, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 207058, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [17:45:38] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch1102 is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: green, timed_out: False, number_of_nodes: 30, number_of_data_nodes: 30, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 5196, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, delaye [17:45:39] gned_shards: 0, number_of_pending_tasks: 1444, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 207065, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [17:45:39] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch1072 is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: green, timed_out: False, number_of_nodes: 30, number_of_data_nodes: 30, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 5196, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, delaye [17:45:40] gned_shards: 0, number_of_pending_tasks: 1444, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 207092, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [17:45:40] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch1081 is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: green, timed_out: False, number_of_nodes: 30, number_of_data_nodes: 30, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 5196, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, delaye [17:45:41] gned_shards: 0, number_of_pending_tasks: 1445, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 207148, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [17:45:46] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch1075 is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: yellow, timed_out: False, number_of_nodes: 28, number_of_data_nodes: 28, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 4847, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 349, del [17:45:46] ssigned_shards: 349, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 93.2832948421863 https://wikitech.wikimedia.org/wiki/Search%23Administration [17:45:48] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch1086 is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: yellow, timed_out: False, number_of_nodes: 28, number_of_data_nodes: 28, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 4847, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 349, del [17:45:48] ssigned_shards: 349, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 93.2832948421863 https://wikitech.wikimedia.org/wiki/Search%23Administration [17:45:48] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch1084 is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: yellow, timed_out: False, number_of_nodes: 28, number_of_data_nodes: 28, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 4847, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 349, del [17:45:48] ssigned_shards: 349, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 93.2832948421863 https://wikitech.wikimedia.org/wiki/Search%23Administration [17:45:48] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch1090 is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: yellow, timed_out: False, number_of_nodes: 28, number_of_data_nodes: 28, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 4847, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 349, del [17:45:49] ssigned_shards: 349, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 93.2832948421863 https://wikitech.wikimedia.org/wiki/Search%23Administration [17:45:49] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch1083 is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: yellow, timed_out: False, number_of_nodes: 28, number_of_data_nodes: 28, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 4847, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 349, del [17:45:50] ssigned_shards: 349, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 93.2832948421863 https://wikitech.wikimedia.org/wiki/Search%23Administration [17:45:50] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch1092 is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: yellow, timed_out: False, number_of_nodes: 28, number_of_data_nodes: 28, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 4847, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 349, del [17:45:51] ssigned_shards: 349, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 93.2832948421863 https://wikitech.wikimedia.org/wiki/Search%23Administration [17:45:51] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch1097 is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: yellow, timed_out: False, number_of_nodes: 28, number_of_data_nodes: 28, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 4847, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 349, del [17:45:52] ssigned_shards: 349, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 93.2832948421863 https://wikitech.wikimedia.org/wiki/Search%23Administration [17:45:52] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch1123 is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: yellow, timed_out: False, number_of_nodes: 28, number_of_data_nodes: 28, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 4847, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 349, del [17:45:53] ssigned_shards: 349, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 93.2832948421863 https://wikitech.wikimedia.org/wiki/Search%23Administration [17:45:53] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch1110 is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: yellow, timed_out: False, number_of_nodes: 28, number_of_data_nodes: 28, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 4847, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 349, del [17:45:54] ssigned_shards: 349, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 93.2832948421863 https://wikitech.wikimedia.org/wiki/Search%23Administration [17:45:54] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch1116 is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: yellow, timed_out: False, number_of_nodes: 28, number_of_data_nodes: 28, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 4847, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 349, del [17:45:55] ssigned_shards: 349, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 93.2832948421863 https://wikitech.wikimedia.org/wiki/Search%23Administration [17:45:56] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch1087 is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: yellow, timed_out: False, number_of_nodes: 28, number_of_data_nodes: 28, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 4847, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 349, del [17:45:56] ssigned_shards: 349, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 93.2832948421863 https://wikitech.wikimedia.org/wiki/Search%23Administration [17:45:56] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch1069 is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: yellow, timed_out: False, number_of_nodes: 28, number_of_data_nodes: 28, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 4847, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 349, del [17:45:57] ssigned_shards: 349, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 93.2832948421863 https://wikitech.wikimedia.org/wiki/Search%23Administration [17:46:04] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch1122 is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: yellow, timed_out: False, number_of_nodes: 28, number_of_data_nodes: 28, discovered_master: True, active_primary_shards: 1732, active_shards: 4847, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 349, delayed_unassigned_shards: 349, numbe [17:46:04] ding_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 93.2832948421863 https://wikitech.wikimedia.org/wiki/Search%23Administration [17:46:08] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch1085 is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: yellow, timed_out: False, number_of_nodes: 28, number_of_data_nodes: 28, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 4847, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 349, del [17:46:09] ssigned_shards: 349, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 93.2832948421863 https://wikitech.wikimedia.org/wiki/Search%23Administration [17:46:20] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch1078 is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: yellow, timed_out: False, number_of_nodes: 28, number_of_data_nodes: 28, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 4847, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 349, del [17:46:20] ssigned_shards: 349, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 93.2832948421863 https://wikitech.wikimedia.org/wiki/Search%23Administration [17:46:20] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch1088 is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: yellow, timed_out: False, number_of_nodes: 28, number_of_data_nodes: 28, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 4847, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 349, del [17:46:20] ssigned_shards: 349, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 93.2832948421863 https://wikitech.wikimedia.org/wiki/Search%23Administration [17:46:20] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch1099 is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: yellow, timed_out: False, number_of_nodes: 28, number_of_data_nodes: 28, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 4847, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 349, del [17:46:21] ssigned_shards: 349, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 93.2832948421863 https://wikitech.wikimedia.org/wiki/Search%23Administration [17:46:21] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch1117 is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: yellow, timed_out: False, number_of_nodes: 28, number_of_data_nodes: 28, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 4847, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 349, del [17:46:22] ssigned_shards: 349, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 93.2832948421863 https://wikitech.wikimedia.org/wiki/Search%23Administration [17:46:22] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch1101 is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: yellow, timed_out: False, number_of_nodes: 28, number_of_data_nodes: 28, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 4847, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 349, del [17:46:23] ssigned_shards: 349, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 93.2832948421863 https://wikitech.wikimedia.org/wiki/Search%23Administration [17:46:23] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch1121 is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: yellow, timed_out: False, number_of_nodes: 28, number_of_data_nodes: 28, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 4847, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 349, del [17:46:24] ssigned_shards: 349, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 93.2832948421863 https://wikitech.wikimedia.org/wiki/Search%23Administration [17:46:24] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch1108 is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: yellow, timed_out: False, number_of_nodes: 28, number_of_data_nodes: 28, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 4847, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 349, del [17:46:25] ssigned_shards: 349, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 93.2832948421863 https://wikitech.wikimedia.org/wiki/Search%23Administration [17:46:25] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch1111 is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: yellow, timed_out: False, number_of_nodes: 28, number_of_data_nodes: 28, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 4847, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 349, del [17:46:26] ssigned_shards: 349, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 93.2832948421863 https://wikitech.wikimedia.org/wiki/Search%23Administration [17:46:26] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch1095 is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: yellow, timed_out: False, number_of_nodes: 28, number_of_data_nodes: 28, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 4847, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 349, del [17:46:27] ssigned_shards: 349, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 93.2832948421863 https://wikitech.wikimedia.org/wiki/Search%23Administration [17:46:50] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie [17:47:56] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch1118 is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: yellow, timed_out: False, number_of_nodes: 28, number_of_data_nodes: 28, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 5009, relocating_shards: 1, initializing_shards: 0, unassigned_shards: 187, del [17:47:56] ssigned_shards: 187, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 96.40107775211702 https://wikitech.wikimedia.org/wiki/Search%23Administration [17:47:58] 10ops-drmrs, 06SRE, 06Infrastructure-Foundations, 10netops: Move TATA IP Transit circuit 1243318 cross-connect from rack B12 to rack B13 - https://phabricator.wikimedia.org/T431605 (10cmooney) 03NEW p:05Triage→03High [17:48:08] 10ops-drmrs, 06SRE, 06Infrastructure-Foundations, 10netops: Move TATA IP Transit circuit 1243318 cross-connect from rack B12 to rack B13 - https://phabricator.wikimedia.org/T431605#12101446 (10cmooney) [17:49:20] 10ops-eqiad, 06SRE, 06cloud-services-team, 10Cloud-VPS, 06DC-Ops: New thermal paste for cloudvirt1071.eqiad.wmnet - https://phabricator.wikimedia.org/T431429#12101448 (10VRiley-WMF) I am currently updating the firmware for bios and iDRAC. My next step will be to apply the thermal paste. [17:50:38] (03PS1) 10AOkoth: fpm: add templates for version 8.4 [puppet] - 10https://gerrit.wikimedia.org/r/1308695 (https://phabricator.wikimedia.org/T424055) [17:51:07] new deployment server looking good, I think :D [17:52:14] Woohoo! [17:52:16] Congrats [17:53:49] thanks ^^ thanks to everyone who made switchovers easy etc :D [17:54:02] and swfrench-wmf in particular for help with all the bookworm updates <3 [17:54:33] * swfrench-wmf realizes this was happening [17:54:41] hi :D [17:54:44] very nice, Raine! :) [17:54:50] congrats [17:55:11] so, we'll be on deploy2003 for about a week, and if nothing exciting happens during that time, I will reimage deploy1003 with bookworm too and switch back [17:55:11] Raine: do you want me to file a task for the inevitable trixie upgrade? :D [17:55:27] taavi: sure, make it depend on not running php on the deploy servers please :P [17:55:39] PROBLEM - Check unit status of push_cross_cluster_settings_9600 on cirrussearch1118 is CRITICAL: CRITICAL: Status of the systemd unit push_cross_cluster_settings_9600 https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [17:55:43] !log bking@cumin2003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch1102.eqiad.wmnet with OS trixie [17:55:44] (we did consider going straight to trixie, but, not feasible because of that) [17:56:25] FIRING: [2x] SystemdUnitFailed: fetch-rings-thanos.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [17:56:34] 10ops-drmrs, 06SRE, 06Infrastructure-Foundations, 10netops: Move TATA IP Transit circuit 1243318 cross-connect from rack B12 to rack B13 - https://phabricator.wikimedia.org/T431605#12101471 (10cmooney) [17:57:11] Raine: I think all php executions that scap runs are done under docker now. [17:57:41] dancy: yes, but there's a few mwscript things that don't work under k8s [17:57:53] ah [17:57:55] (03PS1) 10Majavah: hieradata: Replace cloudinfra-db04 with cloudinfra-db-5 [puppet] - 10https://gerrit.wikimedia.org/r/1308696 (https://phabricator.wikimedia.org/T402005) [17:58:19] most importantly importing large image datasets I think, lemme find the task... [17:58:42] hey, at least we now have a team to own that media-related workflow :D [17:58:58] (03CR) 10Majavah: [C:03+2] hieradata: Replace cloudinfra-db04 with cloudinfra-db-5 [puppet] - 10https://gerrit.wikimedia.org/r/1308696 (https://phabricator.wikimedia.org/T402005) (owner: 10Majavah) [17:59:22] (03CR) 10Dzahn: [C:04-1] "looking at the 8.3 files they are just symbolic links to the general template for all PHP8 versions so this should probably be as well - u" [puppet] - 10https://gerrit.wikimedia.org/r/1308695 (https://phabricator.wikimedia.org/T424055) (owner: 10AOkoth) [17:59:37] (03PS2) 10AOkoth: fpm: add templates for version 8.4 [puppet] - 10https://gerrit.wikimedia.org/r/1308695 (https://phabricator.wikimedia.org/T424055) [17:59:40] https://phabricator.wikimedia.org/T377497 [18:00:42] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie [18:01:03] (wazzup stashbot ) [18:01:05] (03CR) 10AOkoth: "Ack." [puppet] - 10https://gerrit.wikimedia.org/r/1308695 (https://phabricator.wikimedia.org/T424055) (owner: 10AOkoth) [18:01:27] !log kamila@deploy2003 Finished scap sync-world: Test deployment to validate deployment server switchover - T423714 (duration: 18m 29s) [18:01:29] T423714: Upgrade deploy* hosts to Debian Bookworm - https://phabricator.wikimedia.org/T423714 [18:01:53] oh url :D fine fine, T377497 [18:01:54] T377497: Functional replacement for importImages.php on Kubernetes - https://phabricator.wikimedia.org/T377497 [18:02:17] (03Abandoned) 10Jdlrobson: Improve responsiveness of toolbelt [skins/Vector] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1308227 (https://phabricator.wikimedia.org/T429518) (owner: 10Jdlrobson) [18:02:18] taavi: yup and I still have a hard time believing it :D [18:05:06] (03PS3) 10AOkoth: fpm: add templates for version 8.4 [puppet] - 10https://gerrit.wikimedia.org/r/1308695 (https://phabricator.wikimedia.org/T424055) [18:05:10] FIRING: [2x] SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [18:05:39] RECOVERY - Check unit status of push_cross_cluster_settings_9600 on cirrussearch1118 is OK: OK: Status of the systemd unit push_cross_cluster_settings_9600 https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [18:05:52] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host cirrussearch1068.eqiad.wmnet with OS trixie [18:07:11] PROBLEM - Check unit status of push_cross_cluster_settings_9600 on cirrussearch1073 is CRITICAL: CRITICAL: Status of the systemd unit push_cross_cluster_settings_9600 https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [18:08:28] Raine: are you happy if I start doing helmfile stuff? [18:08:49] rzl: go ahead, I'm all done :-) [18:08:55] thanks! [18:10:58] !log rzl@deploy2003 helmfile [staging-eqiad] START helmfile.d/admin 'apply'. [18:11:26] (03CR) 10Dzahn: [C:03+2] phabricator: drop diffusion.ssh-host config [puppet] - 10https://gerrit.wikimedia.org/r/1305044 (https://phabricator.wikimedia.org/T429367) (owner: 10Aklapper) [18:12:36] (03CR) 10Dzahn: [C:03+2] ci::docker: ensure jenkins agent user is in docker group also in prod [puppet] - 10https://gerrit.wikimedia.org/r/1308194 (https://phabricator.wikimedia.org/T418521) (owner: 10Dzahn) [18:13:47] !log rzl@deploy2003 helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. [18:13:55] !log rzl@deploy2003 helmfile [staging-codfw] START helmfile.d/admin 'apply'. [18:15:03] !log rzl@deploy2003 helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. [18:15:27] !log bking@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage [18:18:26] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host cirrussearch1122.eqiad.wmnet with OS trixie [18:18:58] !log bking@cumin2003 START - Cookbook sre.hosts.move-vlan for host cirrussearch1122 [18:19:04] !log bking@cumin2003 START - Cookbook sre.dns.netbox [18:19:23] !log rzl@deploy2003 helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. [18:20:14] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage [18:20:34] (03CR) 10Dzahn: [C:03+2] ci: add contint2003 to list of hosts allowed to git/9148 [puppet] - 10https://gerrit.wikimedia.org/r/1308245 (https://phabricator.wikimedia.org/T418521) (owner: 10Dzahn) [18:21:04] !log rzl@deploy2003 helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. [18:21:11] !log rzl@deploy2003 helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. [18:21:43] !log bking@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage [18:21:59] !log rzl@deploy2003 helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. [18:24:39] 10ops-drmrs, 06SRE, 06Infrastructure-Foundations, 10netops: Move TATA IP Transit circuit 1243318 cross-connect from rack B12 to rack B13 - https://phabricator.wikimedia.org/T431605#12101591 (10cmooney) [18:25:31] !log bking@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" [18:25:36] !log bking@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003" [18:25:36] !log bking@cumin2003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [18:25:37] !log bking@cumin2003 START - Cookbook sre.dns.wipe-cache cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [18:25:40] !log bking@cumin2003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [18:25:41] !log bking@cumin2003 START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1122 [18:26:04] !log bking@cumin2003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1122 [18:26:04] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1122 [18:27:40] (03CR) 10Cathal Mooney: [C:03+2] Nokia: Create ACL prefix-list entry for BGP peers and use in CPM acl [homer/public] - 10https://gerrit.wikimedia.org/r/1284695 (https://phabricator.wikimedia.org/T425703) (owner: 10Cathal Mooney) [18:28:20] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage [18:29:58] (03Merged) 10jenkins-bot: Nokia: Create ACL prefix-list entry for BGP peers and use in CPM acl [homer/public] - 10https://gerrit.wikimedia.org/r/1284695 (https://phabricator.wikimedia.org/T425703) (owner: 10Cathal Mooney) [18:32:40] (03Abandoned) 10JHathaway: Puppet 8: Replace legacy facts [puppet] - 10https://gerrit.wikimedia.org/r/1305987 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [18:34:13] (03CR) 10Ssingh: "Looks good! I am saving the +1 for when we actually merge and try again, which will be Monday 13 July." [puppet] - 10https://gerrit.wikimedia.org/r/1308150 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [18:35:05] (03CR) 10RLazarus: [C:03+2] Remove ipoid chart and service definitions [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306784 (https://phabricator.wikimedia.org/T416623) (owner: 10Dreamy Jazz) [18:37:15] (03CR) 10Ssingh: [C:04-1] Revert^2 "varnish: Add CSP report-only header value" (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1307826 (owner: 10BCornwall) [18:38:38] !log bking@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage [18:40:51] (03CR) 10Jhancock.wm: [C:03+2] preseed.yaml: Use new (EFI) Cassandra JBOD preseed [puppet] - 10https://gerrit.wikimedia.org/r/1308188 (https://phabricator.wikimedia.org/T416538) (owner: 10Eevans) [18:44:27] (03Merged) 10jenkins-bot: Remove ipoid chart and service definitions [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306784 (https://phabricator.wikimedia.org/T416623) (owner: 10Dreamy Jazz) [18:44:29] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage [18:46:01] !log rzl@deploy2003 helmfile [staging-codfw] START helmfile.d/admin 'apply'. [18:46:50] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in module airflow [puppet] - 10https://gerrit.wikimedia.org/r/1308716 (https://phabricator.wikimedia.org/T372666) [18:47:18] !log rzl@deploy2003 helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. [18:47:38] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie [18:47:40] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in module bigtop [puppet] - 10https://gerrit.wikimedia.org/r/1308717 (https://phabricator.wikimedia.org/T372666) [18:48:03] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in module confluent [puppet] - 10https://gerrit.wikimedia.org/r/1308718 (https://phabricator.wikimedia.org/T372666) [18:48:43] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in module elasticsearch [puppet] - 10https://gerrit.wikimedia.org/r/1308719 (https://phabricator.wikimedia.org/T372666) [18:48:58] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1068.eqiad.wmnet with OS trixie [18:49:00] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in module kafkatee [puppet] - 10https://gerrit.wikimedia.org/r/1308720 (https://phabricator.wikimedia.org/T372666) [18:49:18] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in module opensearch [puppet] - 10https://gerrit.wikimedia.org/r/1308721 (https://phabricator.wikimedia.org/T372666) [18:49:22] FIRING: [5x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-1/0/2 (Transport: cr1-eqiad:et-1/1/2 (Arelion, IC-374549) {#12267}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [18:50:31] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host cirrussearch1100.eqiad.wmnet with OS trixie [18:51:00] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in module statistics [puppet] - 10https://gerrit.wikimedia.org/r/1308722 (https://phabricator.wikimedia.org/T372666) [18:51:29] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in module zookeeper [puppet] - 10https://gerrit.wikimedia.org/r/1308724 (https://phabricator.wikimedia.org/T372666) [18:51:52] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in profile analytics [puppet] - 10https://gerrit.wikimedia.org/r/1308725 (https://phabricator.wikimedia.org/T372666) [18:52:06] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in profile bigtop [puppet] - 10https://gerrit.wikimedia.org/r/1308727 (https://phabricator.wikimedia.org/T372666) [18:52:16] !log rzl@deploy2003 helmfile [staging-eqiad] START helmfile.d/admin 'apply'. [18:52:53] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in profile druid [puppet] - 10https://gerrit.wikimedia.org/r/1308730 (https://phabricator.wikimedia.org/T372666) [18:53:00] !log rzl@deploy2003 helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. [18:53:09] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in profile elasticsearch [puppet] - 10https://gerrit.wikimedia.org/r/1308732 (https://phabricator.wikimedia.org/T372666) [18:54:22] FIRING: [5x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-1/0/2 (Transport: cr1-eqiad:et-1/1/2 (Arelion, IC-374549) {#12267}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [18:55:45] !log rzl@deploy2003 helmfile [codfw] START helmfile.d/admin 'apply'. [18:57:22] !log rzl@deploy2003 helmfile [codfw] DONE helmfile.d/admin 'apply'. [18:58:44] !log rzl@deploy2003 helmfile [eqiad] START helmfile.d/admin 'apply'. [18:59:10] !log rolling out update to BGP ACL on Nokia Switches eqiad, codfw & ulsfo T425703 [18:59:11] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [19:00:14] (03CR) 10Dzahn: "feel free to add me back when there is something new here" [puppet] - 10https://gerrit.wikimedia.org/r/1178874 (https://phabricator.wikimedia.org/T378028) (owner: 10AOkoth) [19:00:36] !log rzl@deploy2003 helmfile [eqiad] DONE helmfile.d/admin 'apply'. [19:02:45] PROBLEM - OSPF status on cr1-codfw is CRITICAL: OSPFv2: 4/5 UP : OSPFv3: 4/4 UP : 5 v2 P2P interfaces vs. 4 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [19:03:08] (03CR) 10RLazarus: [C:03+2] deployment_server: absent ipoid kubernetes service [puppet] - 10https://gerrit.wikimedia.org/r/1306779 (https://phabricator.wikimedia.org/T416623) (owner: 10Dreamy Jazz) [19:03:41] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1122.eqiad.wmnet with OS trixie [19:04:22] FIRING: [5x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-1/0/2 (Transport: cr1-eqiad:et-1/1/2 (Arelion, IC-374549) {#12267}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [19:05:54] !log bking@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage [19:08:14] !log cmooney@cumin1003 START - Cookbook sre.network.tls for network device asw1-b3-magru [19:08:36] !log cmooney@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b3-magru [19:09:09] !log cmooney@cumin1003 START - Cookbook sre.network.tls for network device asw1-b4-magru [19:09:30] !log cmooney@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b4-magru [19:09:33] !log cmooney@cumin1003 START - Cookbook sre.network.tls for network device lsw1-c3-codfw [19:09:43] !log cmooney@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-codfw [19:09:47] !log cmooney@cumin1003 START - Cookbook sre.network.tls for network device lsw1-c6-codfw [19:09:57] !log cmooney@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-codfw [19:10:00] !log cmooney@cumin1003 START - Cookbook sre.network.tls for network device lsw1-d6-codfw [19:10:02] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage [19:10:10] !log cmooney@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d6-codfw [19:10:13] !log cmooney@cumin1003 START - Cookbook sre.network.tls for network device lsw1-d8-codfw [19:10:23] !log cmooney@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-codfw [19:10:26] !log cmooney@cumin1003 START - Cookbook sre.network.tls for network device cr1-magru [19:10:47] !log cmooney@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-magru [19:10:50] !log cmooney@cumin1003 START - Cookbook sre.network.tls for network device ssw1-d1-codfw [19:11:00] !log cmooney@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw [19:13:38] 10ops-esams, 06SRE, 06Commons, 06DC-Ops, and 3 others: ESAMS and others serving older revisions of overwritten files - https://phabricator.wikimedia.org/T425216#12101800 (10ssingh) Thanks for reporting @AlexisJazz. We discussed this in Traffic. @Bblack suggested that since it's difficult to debug one-off c... [19:14:22] 10ops-eqiad, 06SRE, 06cloud-services-team, 10Cloud-VPS, 06DC-Ops: New thermal paste for cloudvirt1071.eqiad.wmnet - https://phabricator.wikimedia.org/T431429#12101802 (10VRiley-WMF) I have applied the thermal paste and the BIOS and iDRAC are up to date with firmware. @Andrew would you be able to look thi... [19:18:43] !log cmooney@cumin1003 START - Cookbook sre.network.tls for network device lsw1-c2-codfw [19:18:53] !log cmooney@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-codfw [19:18:56] !log cmooney@cumin1003 START - Cookbook sre.network.tls for network device lsw1-d1-codfw [19:19:06] !log cmooney@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-codfw [19:19:09] !log cmooney@cumin1003 START - Cookbook sre.network.tls for network device lsw1-d3-codfw [19:19:19] !log cmooney@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw [19:19:22] !log cmooney@cumin1003 START - Cookbook sre.network.tls for network device lsw1-d7-codfw [19:19:32] !log cmooney@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-codfw [19:19:35] !log cmooney@cumin1003 START - Cookbook sre.network.tls for network device cr2-magru [19:19:56] !log cmooney@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-magru [19:19:57] !log gerrit - replacing private key for registerEmail verification [19:19:58] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [19:19:59] !log cmooney@cumin1003 START - Cookbook sre.network.tls for network device ssw1-d8-codfw [19:20:09] !log cmooney@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-codfw [19:21:47] RECOVERY - OSPF status on cr1-codfw is OK: OSPFv2: 4/4 UP : OSPFv3: 4/4 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [19:29:04] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1100.eqiad.wmnet with OS trixie [19:29:23] FIRING: [5x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-1/0/2 (Transport: cr1-eqiad:et-1/1/2 (Arelion, IC-374549) {#12267}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [19:30:23] !log ran homer on lsw1-d8-codfw, adding wikikube-worker2331 to cluster [19:30:25] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [19:31:08] (03CR) 10Jasmine: [C:03+2] k8s: add wikikube-worker2331 [puppet] - 10https://gerrit.wikimedia.org/r/1289022 (https://phabricator.wikimedia.org/T426688) (owner: 10Jasmine) [19:32:05] 10ops-eqiad, 06Data-Persistence, 06DBA, 06DC-Ops, and 3 others: es7 primary (es1039.eqiad.wmnet) crashed - https://phabricator.wikimedia.org/T430764#12101865 (10VRiley-WMF) Scratch that, sorry. I just logged back into it and it seems to throw the same error. Let me make a few more changes on my end. [19:35:55] (03PS2) 10Danielyepezgarces: extwiki: Rename wgSitename to Güiquipedia [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1307836 (https://phabricator.wikimedia.org/T431334) [19:36:57] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in profile hive [puppet] - 10https://gerrit.wikimedia.org/r/1308747 (https://phabricator.wikimedia.org/T372666) [19:37:17] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in profile kafka [puppet] - 10https://gerrit.wikimedia.org/r/1308748 (https://phabricator.wikimedia.org/T372666) [19:37:29] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in profile kafkatee [puppet] - 10https://gerrit.wikimedia.org/r/1308749 (https://phabricator.wikimedia.org/T372666) [19:38:11] !log cmooney@cumin1003 START - Cookbook sre.network.tls for network device lsw1-c5-codfw [19:38:25] !log cmooney@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-codfw [19:38:28] !log cmooney@cumin1003 START - Cookbook sre.network.tls for network device lsw1-c7-codfw [19:38:39] !log cmooney@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-codfw [19:38:41] !log cmooney@cumin1003 START - Cookbook sre.network.tls for network device lsw1-d5-codfw [19:38:52] !log cmooney@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d5-codfw [19:40:21] (03PS3) 10Danielyepezgarces: extwiki: Rename wgSitename to Güiquipedia [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1307836 (https://phabricator.wikimedia.org/T431334) [19:40:44] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host cirrussearch1098.eqiad.wmnet with OS trixie [19:40:54] (03PS4) 10Danielyepezgarces: extwiki: Rename wgSitename to Güiquipedia [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1307836 (https://phabricator.wikimedia.org/T431334) [19:42:59] (03CR) 10RLazarus: [C:03+2] deployment_server: remove ipoid users [puppet] - 10https://gerrit.wikimedia.org/r/1306782 (https://phabricator.wikimedia.org/T416623) (owner: 10Dreamy Jazz) [19:43:14] (03PS9) 10Clément Goubert: deployment_server: remove ipoid users [puppet] - 10https://gerrit.wikimedia.org/r/1306782 (https://phabricator.wikimedia.org/T416623) (owner: 10Dreamy Jazz) [19:43:23] !log jasmine@cumin2002 conftool action : set/weight=10; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc [19:43:43] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in profile opensearch [puppet] - 10https://gerrit.wikimedia.org/r/1308750 (https://phabricator.wikimedia.org/T372666) [19:43:55] !log jasmine@cumin2002 conftool action : set/pooled=yes; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc [19:44:05] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in profile presto [puppet] - 10https://gerrit.wikimedia.org/r/1308751 (https://phabricator.wikimedia.org/T372666) [19:44:31] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in profile query_service [puppet] - 10https://gerrit.wikimedia.org/r/1308752 (https://phabricator.wikimedia.org/T372666) [19:44:43] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in profile zookeeper [puppet] - 10https://gerrit.wikimedia.org/r/1308753 (https://phabricator.wikimedia.org/T372666) [19:45:19] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in role analytics_cluster [puppet] - 10https://gerrit.wikimedia.org/r/1308754 (https://phabricator.wikimedia.org/T372666) [19:45:44] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in role analytics_test_cluster [puppet] - 10https://gerrit.wikimedia.org/r/1308755 (https://phabricator.wikimedia.org/T372666) [19:45:59] (03CR) 10RLazarus: [C:03+2] deployment_server: remove ipoid users [puppet] - 10https://gerrit.wikimedia.org/r/1306782 (https://phabricator.wikimedia.org/T416623) (owner: 10Dreamy Jazz) [19:46:17] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in role druid [puppet] - 10https://gerrit.wikimedia.org/r/1308756 (https://phabricator.wikimedia.org/T372666) [19:46:29] !log restarting gerrit on gerrit-spare.wikimedia.org (gerrit2002) [19:46:30] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [19:46:33] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in role wdqs [puppet] - 10https://gerrit.wikimedia.org/r/1308757 (https://phabricator.wikimedia.org/T372666) [19:46:57] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in role kafka [puppet] - 10https://gerrit.wikimedia.org/r/1308759 (https://phabricator.wikimedia.org/T372666) [19:47:19] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in misc kafka [puppet] - 10https://gerrit.wikimedia.org/r/1308760 (https://phabricator.wikimedia.org/T372666) [19:48:22] !log restarting gerrit on gerrit-replica.wikimedia.org (gerrit1003) [19:48:23] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [19:48:52] !log jasmine@cumin2002 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2331.codfw.wmnet [19:48:55] !log jasmine@cumin2002 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2331.codfw.wmnet [19:49:01] jouncebot: nowandenext [19:49:06] jouncebot: nowandnext [19:49:07] No deployments scheduled for the next 0 hour(s) and 10 minute(s) [19:49:07] In 0 hour(s) and 10 minute(s): UTC late backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260708T2000) [19:49:31] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in profile hadoop [puppet] - 10https://gerrit.wikimedia.org/r/1308761 (https://phabricator.wikimedia.org/T372666) [19:50:10] (03CR) 10JHathaway: "for sure, coming your way!" [puppet] - 10https://gerrit.wikimedia.org/r/1305987 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [19:50:43] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308761 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [19:52:07] jhathaway: I would like to interrupt your flow for a minute :) [19:52:17] gotta restart gerrit, heh [19:52:37] please do, one can only look at puppet facts so long in one day :) [19:52:42] ok :) [19:52:58] !log restarting gerrit on gerrit.wikimedia.org (gerrit2003) [19:52:58] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [19:54:29] it always takes "just long enough to start worrying and then it's back" [19:54:36] !log cmooney@cumin1003 START - Cookbook sre.network.tls for network device lsw1-c1-codfw [19:54:46] FIRING: GerritHAProxyBackendUnavailable: Gerrit backend is unavilable for tcp-proxy (HAProxy) gerrit_ssh - https://wikitech.wikimedia.org/wiki/Gerrit/Operations#GerritHAProxyBackendUnavailable - grafana.wikimedia.org/d/459365f6-df37-48d6-8142-82b22c1875e7/gerrit-tcp-proxy?viewPanel=panel-15 - https://alerts.wikimedia.org/?q=alertname%3DGerritHAProxyBackendUnavailable [19:54:47] !log cmooney@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c1-codfw [19:54:49] !log cmooney@cumin1003 START - Cookbook sre.network.tls for network device lsw1-d4-codfw [19:55:00] !log cmooney@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-codfw [19:55:12] and back. thanks! [19:55:46] FIRING: [2x] GerritHAProxyServiceUnavailable: Gerrit tcp-proxy (HAProxy) service gerrit_ssh is DOWN in codfw - https://wikitech.wikimedia.org/wiki/Gerrit/Operations#GerritHAProxyServiceUnavailable - grafana.wikimedia.org/d/459365f6-df37-48d6-8142-82b22c1875e7/gerrit-tcp-proxy?viewPanel=panel-15 - https://alerts.wikimedia.org/?q=alertname%3DGerritHAProxyServiceUnavailable [19:56:01] !log bking@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage [19:59:22] FIRING: [5x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-1/0/2 (Transport: cr1-eqiad:et-1/1/2 (Arelion, IC-374549) {#12267}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [19:59:46] FIRING: GerritHigh5xxRatio: High rate of HTTP 5xx responses on gerrit2003 - TODO - https://grafana.wikimedia.org/d/nX8li17Sk/gerrit-overview-sre-collab - https://alerts.wikimedia.org/?q=alertname%3DGerritHigh5xxRatio [19:59:46] RESOLVED: [4x] GerritHAProxyBackendUnavailable: Gerrit backend is unavilable for tcp-proxy (HAProxy) gerrit_ssh - https://wikitech.wikimedia.org/wiki/Gerrit/Operations#GerritHAProxyBackendUnavailable - grafana.wikimedia.org/d/459365f6-df37-48d6-8142-82b22c1875e7/gerrit-tcp-proxy?viewPanel=panel-15 - https://alerts.wikimedia.org/?q=alertname%3DGerritHAProxyBackendUnavailable [19:59:49] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308761 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [19:59:51] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308753 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [19:59:53] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308760 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [19:59:55] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308759 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [19:59:57] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308757 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [19:59:58] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308756 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [19:59:59] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308755 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [20:00:04] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308752 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [20:00:05] RoanKattouw, urbanecm, TheresNoTime, kindrobot, and cjming: It is that lovely time of the day again! You are hereby commanded to deploy UTC late backport window. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260708T2000). [20:00:05] cjming: A patch you scheduled for UTC late backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [20:00:08] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308754 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [20:00:12] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308751 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [20:00:16] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308750 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [20:00:19] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage [20:00:24] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308749 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [20:00:28] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308748 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [20:00:32] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308747 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [20:00:36] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308732 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [20:00:40] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308730 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [20:00:44] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308727 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [20:00:46] RESOLVED: [2x] GerritHAProxyServiceUnavailable: Gerrit tcp-proxy (HAProxy) service gerrit_ssh is DOWN in codfw - https://wikitech.wikimedia.org/wiki/Gerrit/Operations#GerritHAProxyServiceUnavailable - grafana.wikimedia.org/d/459365f6-df37-48d6-8142-82b22c1875e7/gerrit-tcp-proxy?viewPanel=panel-15 - https://alerts.wikimedia.org/?q=alertname%3DGerritHAProxyServiceUnavailable [20:00:48] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308725 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [20:00:52] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308724 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [20:01:00] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308722 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [20:01:04] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308721 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [20:01:09] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308720 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [20:01:13] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308719 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [20:01:17] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308718 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [20:01:25] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308717 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [20:02:41] i will deploy my patch [20:04:41] (03PS4) 10Clare Ming: Move Test Kitchen config from CommonSettings.php [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308183 (https://phabricator.wikimedia.org/T431257) [20:04:46] RESOLVED: GerritHigh5xxRatio: High rate of HTTP 5xx responses on gerrit2003 - TODO - https://grafana.wikimedia.org/d/nX8li17Sk/gerrit-overview-sre-collab - https://alerts.wikimedia.org/?q=alertname%3DGerritHigh5xxRatio [20:07:04] (03CR) 10TrainBranchBot: [C:03+2] "Approved by cjming@deploy2003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308183 (https://phabricator.wikimedia.org/T431257) (owner: 10Clare Ming) [20:07:07] (03PS1) 10Ahmon Dancy: buildkitd: Bump buildkit image to wmf-v0.31.1 [puppet] - 10https://gerrit.wikimedia.org/r/1308767 (https://phabricator.wikimedia.org/T429988) [20:08:57] (03Merged) 10jenkins-bot: Move Test Kitchen config from CommonSettings.php [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308183 (https://phabricator.wikimedia.org/T431257) (owner: 10Clare Ming) [20:09:53] !log cjming@deploy2003 Started scap sync-world: Backport for [[gerrit:1308183|Move Test Kitchen config from CommonSettings.php (T431257)]] [20:09:58] T431257: Move $wgTestKitchenLoggedInExperimentEventIntakeServiceUrl to InitialiseSettings.php - https://phabricator.wikimedia.org/T431257 [20:11:51] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308716 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [20:13:02] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host cirrussearch1073.eqiad.wmnet with OS trixie [20:18:07] PROBLEM - OpenSearch health check for shards on 9200 on cirrussearch1094 is CRITICAL: CRITICAL - elasticsearch http://localhost:9200/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9200): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [20:20:34] i haven't deployed in a while - does it usually take this long to build the K8s images? it feels like it's been stuck there for nearly 10 minutes [20:20:58] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1098.eqiad.wmnet with OS trixie [20:22:24] 06SRE, 06Traffic: Investigate / Fix upload.wikimedia.org lack of Cache-Control headers - https://phabricator.wikimedia.org/T431621#12102056 (10ssingh) [20:23:34] oh - i guess it took 10 m 52s [20:28:42] !log cjming@deploy2003 cjming: Backport for [[gerrit:1308183|Move Test Kitchen config from CommonSettings.php (T431257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [20:28:45] T431257: Move $wgTestKitchenLoggedInExperimentEventIntakeServiceUrl to InitialiseSettings.php - https://phabricator.wikimedia.org/T431257 [20:28:49] cjming: Looks like that change resulted in l10n rebuild [20:29:40] (03CR) 10RLazarus: [C:03+2] admin_ng: Remove obsolete coredns 1.8.7-2 tag, unset everywhere [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305544 (owner: 10RLazarus) [20:30:08] !log cjming@deploy2003 cjming: Continuing with deployment [20:30:34] dancy: thanks - ya i guess those take alot longer [20:31:45] (03CR) 10SBassett: [C:03+2] chore: Deploy new version of security hall-of-fame [deployment-charts] - 10https://gerrit.wikimedia.org/r/1307235 (https://phabricator.wikimedia.org/T429821) (owner: 10Rscout) [20:32:25] !log bking@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage [20:34:53] (03CR) 10Dzahn: [C:03+2] buildkitd: Bump buildkit image to wmf-v0.31.1 [puppet] - 10https://gerrit.wikimedia.org/r/1308767 (https://phabricator.wikimedia.org/T429988) (owner: 10Ahmon Dancy) [20:35:45] FIRING: WidespreadPuppetFailure: Puppet has failed in drmrs - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [20:37:20] (03CR) 10Dzahn: "@sukhe is this really ok? could it permanently have the low value?" [dns] - 10https://gerrit.wikimedia.org/r/1306459 (https://phabricator.wikimedia.org/T425441) (owner: 10Arnaudb) [20:37:56] (03PS4) 10JHathaway: Puppet 8: Replace legacy facts [puppet] - 10https://gerrit.wikimedia.org/r/1305988 (https://phabricator.wikimedia.org/T372666) [20:37:58] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage [20:38:15] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1305988 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [20:38:30] (03CR) 10CI reject: [V:04-1] Puppet 8: Replace legacy facts [puppet] - 10https://gerrit.wikimedia.org/r/1305988 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [20:38:41] (03Merged) 10jenkins-bot: admin_ng: Remove obsolete coredns 1.8.7-2 tag, unset everywhere [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305544 (owner: 10RLazarus) [20:39:01] (03CR) 10JHathaway: "@eevans@wikimedia.org manually rebased, happy to break up into modules if desired." [puppet] - 10https://gerrit.wikimedia.org/r/1305988 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [20:40:02] 10ops-eqiad, 06SRE, 06DC-Ops, 10decommission-hardware: decommission cloudcontrol1008-dev.eqiad.wmnet - https://phabricator.wikimedia.org/T429527#12102105 (10Andrew) [20:40:03] 10ops-eqiad, 06SRE, 06cloud-services-team, 06DC-Ops: Do something with cloudcontrol100[8-10]-dev - https://phabricator.wikimedia.org/T406630#12102104 (10Andrew) [20:40:07] 10ops-eqiad, 06SRE, 06cloud-services-team, 06DC-Ops: Do something with cloudcontrol100[8-10]-dev - https://phabricator.wikimedia.org/T406630#12102106 (10Andrew) 05Open→03Resolved [20:40:58] 06SRE, 06Traffic: Investigate / Fix upload.wikimedia.org lack of Cache-Control headers - https://phabricator.wikimedia.org/T431621#12102110 (10ssingh) Thanks for filing the task! I think we will need to investigate a bit more. For //a// data point to start with: Image from the front page: https://upload.wiki... [20:42:03] !log jhancock@cumin2002 START - Cookbook sre.hosts.reimage for host restbase2039.codfw.wmnet with OS bullseye [20:42:17] 10ops-codfw, 06SRE, 06Data-Persistence, 06DC-Ops: FY2526 Q3:rack/setup/install restbase2039 - https://phabricator.wikimedia.org/T416538#12102111 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by jhancock@cumin2002 for host restbase2039.codfw.wmnet with OS bullseye [20:42:33] (03Merged) 10jenkins-bot: chore: Deploy new version of security hall-of-fame [deployment-charts] - 10https://gerrit.wikimedia.org/r/1307235 (https://phabricator.wikimedia.org/T429821) (owner: 10Rscout) [20:42:55] !log cjming@deploy2003 Finished scap sync-world: Backport for [[gerrit:1308183|Move Test Kitchen config from CommonSettings.php (T431257)]] (duration: 33m 02s) [20:42:58] T431257: Move $wgTestKitchenLoggedInExperimentEventIntakeServiceUrl to InitialiseSettings.php - https://phabricator.wikimedia.org/T431257 [20:43:16] (03CR) 10Dzahn: "I see the latest comment says there was still one question open. Any updates on that one?" [puppet] - 10https://gerrit.wikimedia.org/r/1290731 (https://phabricator.wikimedia.org/T425441) (owner: 10Arnaudb) [20:43:25] https://test.wikipedia.org/wiki/Main_Page is an internal error, with trace which is fun [20:43:33] wow - that was 36 minutes [20:48:24] !log deploy2003 - kill 1102 (stunnel4) ; systemctl start stunnel4 (T418262) [20:48:27] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [20:48:27] T418262: deploy2003 implementation tracking - https://phabricator.wikimedia.org/T418262 [20:49:39] FIRING: CirrusSearchNodeIndexingNotIncreasing: Elasticsearch instance cirrussearch1094-production-search-eqiad is not indexing - https://wikitech.wikimedia.org/wiki/Search/Elasticsearch_Administration#Indexing_hung_and_not_making_progress - https://grafana.wikimedia.org/d/JLK3I_siz/elasticsearch-indexing?orgId=1&from=now-3d&to=now&viewPanel=57 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchNodeIndexingNotIncreasing [20:50:47] 10ops-eqiad, 06Data-Persistence, 06DBA, 06DC-Ops, and 3 others: es7 primary (es1039.eqiad.wmnet) crashed - https://phabricator.wikimedia.org/T430764#12102163 (10VRiley-WMF) Dell had me update the firmware for the drives and the backplan as well to be sure. The system is currently showing as healthy, but I... [20:54:22] FIRING: [5x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-1/0/2 (Transport: cr1-eqiad:et-1/1/2 (Arelion, IC-374549) {#12267}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [20:58:20] (03PS1) 10CDobbins: varnish: update CSP report-only header [puppet] - 10https://gerrit.wikimedia.org/r/1308772 [20:58:51] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1073.eqiad.wmnet with OS trixie [21:00:00] !log jhancock@cumin2002 START - Cookbook sre.hosts.downtime for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage [21:00:04] Deploy window Wikifunctions Services UTC Late (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260708T2100) [21:00:45] RESOLVED: WidespreadPuppetFailure: Puppet has failed in drmrs - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [21:01:14] ^ fixing puppet on 1(!) host pushes this under the baseline. that is what "widespread" means :p [21:03:46] (03CR) 10RLazarus: [C:03+2] coredns: Parameterize `name` and `k8s-app` [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305545 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [21:04:03] FIRING: MediaWikiEditFailures: Elevated MediaWiki edit failures (session_loss) for cluster - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://grafana.wikimedia.org/d/000000208/edit-count?orgId=1&viewPanel=13 - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiEditFailures [21:04:13] !log jhancock@cumin2002 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage [21:05:10] whoa, that session_loss spike is real [21:05:19] (03CR) 10CDobbins: "upload: 0 tests failed, 0 tests skipped, 21 tests passed" [puppet] - 10https://gerrit.wikimedia.org/r/1308772 (owner: 10CDobbins) [21:06:27] I don't see any immediate causes looking at sessionstore, except some traffic changes that look associated with cjming's deployment? [21:06:58] a few minutes' difference there, so I don't know if it's actually causative or not [21:09:22] FIRING: [5x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-1/0/2 (Transport: cr1-eqiad:et-1/1/2 (Arelion, IC-374549) {#12267}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [21:13:02] (03PS1) 10Btullis: datahub: raise mae-consumer Kafka listener concurrency [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308775 (https://phabricator.wikimedia.org/T402408) [21:13:11] (03Merged) 10jenkins-bot: coredns: Parameterize `name` and `k8s-app` [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305545 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [21:13:28] (03CR) 10RLazarus: [C:03+2] coredns: Add an internal_only value [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305546 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [21:13:38] (03PS3) 10RLazarus: coredns: Add an internal_only value [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305546 (https://phabricator.wikimedia.org/T427864) [21:13:39] (03CR) 10CI reject: [V:04-1] coredns: Add an internal_only value [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305546 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [21:14:49] (03CR) 10RLazarus: "recheck" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305546 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [21:15:27] (03PS5) 10Jasmine: sre.k8s.renumber-node: Increase timeout to 600s for puppet runs on deploy hosts [cookbooks] - 10https://gerrit.wikimedia.org/r/1307872 (https://phabricator.wikimedia.org/T430226) [21:16:31] (03PS6) 10Jasmine: sre.k8s.renumber-node: Increase timeout to 600s for puppet runs on deploy hosts [cookbooks] - 10https://gerrit.wikimedia.org/r/1307872 (https://phabricator.wikimedia.org/T430226) [21:18:48] (03PS5) 10JHathaway: Puppet 8: Replace legacy facts [puppet] - 10https://gerrit.wikimedia.org/r/1305988 (https://phabricator.wikimedia.org/T372666) [21:19:22] (03CR) 10CI reject: [V:04-1] Puppet 8: Replace legacy facts [puppet] - 10https://gerrit.wikimedia.org/r/1305988 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:19:47] cjming: on the phab ticket T431257 you and another person said that config variable needs to stay where it is - but then it was moved today. Is that right? [21:19:47] T431257: Move $wgTestKitchenLoggedInExperimentEventIntakeServiceUrl to InitialiseSettings.php - https://phabricator.wikimedia.org/T431257 [21:20:28] (03CR) 10JHathaway: "thanks Hugh, also happy to break it up per module" [puppet] - 10https://gerrit.wikimedia.org/r/1305989 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:20:46] cjming: oh, nevermind. I read that wrong. [21:21:06] !log jhancock@cumin2002 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" [21:21:27] !log jhancock@cumin2002 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002" [21:21:29] !log jhancock@cumin2002 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host restbase2039.codfw.wmnet with OS bullseye [21:21:37] 10ops-codfw, 06SRE, 06Data-Persistence, 06DC-Ops: FY2526 Q3:rack/setup/install restbase2039 - https://phabricator.wikimedia.org/T416538#12102252 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by jhancock@cumin2002 for host restbase2039.codfw.wmnet with OS bullseye completed: - restbase... [21:22:14] PROBLEM - Check unit status of push_cross_cluster_settings_9200 on cirrussearch1094 is CRITICAL: CRITICAL: Status of the systemd unit push_cross_cluster_settings_9200 https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [21:22:16] 10ops-codfw, 06SRE, 06Data-Persistence, 06DC-Ops: FY2526 Q3:rack/setup/install restbase2039 - https://phabricator.wikimedia.org/T416538#12102255 (10Jhancock.wm) 05Open→03Resolved a:03Jhancock.wm [21:22:22] 10ops-codfw, 06SRE, 06Data-Persistence, 06DC-Ops: FY2526 Q3:rack/setup/install restbase2039 - https://phabricator.wikimedia.org/T416538#12102260 (10Jhancock.wm) @Eevans all done. ty! [21:22:42] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host cirrussearch1094.eqiad.wmnet with OS trixie [21:24:34] (03CR) 10Btullis: [C:03+2] datahub: raise mae-consumer Kafka listener concurrency [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308775 (https://phabricator.wikimedia.org/T402408) (owner: 10Btullis) [21:24:42] (03Merged) 10jenkins-bot: coredns: Add an internal_only value [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305546 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [21:26:54] (03PS6) 10JHathaway: Puppet 8: Replace legacy facts [puppet] - 10https://gerrit.wikimedia.org/r/1305988 (https://phabricator.wikimedia.org/T372666) [21:27:01] (03Merged) 10jenkins-bot: datahub: raise mae-consumer Kafka listener concurrency [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308775 (https://phabricator.wikimedia.org/T402408) (owner: 10Btullis) [21:27:35] !log rzl@deploy2003 helmfile [staging-codfw] START helmfile.d/admin 'apply'. [21:28:55] (03CR) 10JHathaway: "recheck" [puppet] - 10https://gerrit.wikimedia.org/r/1305984 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:28:56] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1305984 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:29:03] RESOLVED: MediaWikiEditFailures: Elevated MediaWiki edit failures (session_loss) for cluster - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://grafana.wikimedia.org/d/000000208/edit-count?orgId=1&viewPanel=13 - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiEditFailures [21:29:20] !log rzl@deploy2003 helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. [21:31:03] PROBLEM - PyBal backends health check on lvs2014 is CRITICAL: PYBAL CRITICAL - CRITICAL - k8s-ingress-staging_30443: Servers kubestage2003.codfw.wmnet, kubestage2004.codfw.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [21:31:13] PROBLEM - PyBal backends health check on lvs2013 is CRITICAL: PYBAL CRITICAL - CRITICAL - k8s-ingress-staging_30443: Servers kubestage2002.codfw.wmnet, kubestage2001.codfw.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [21:31:33] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308760 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:31:51] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308757 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:32:04] FIRING: MediaWikiEditFailures: Elevated MediaWiki edit failures (session_loss) for cluster - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://grafana.wikimedia.org/d/000000208/edit-count?orgId=1&viewPanel=13 - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiEditFailures [21:32:22] is there concern that my patch is causing issues? i can revert if needed [21:33:20] cjming: I'm curious about that spike in edit failures, click through to the grafana link -- but I don't have any more context than that alert, haven't dug for logs or anything [21:33:44] I also don't know that it's related to your patch at all, it just started right after so I'm naturally suspicious :) [21:35:54] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply [21:35:59] all that patch did was move the logged-in experiment intake service url from CommonSettings to InitializeSettings [21:36:27] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply} [21:37:30] maybe for some logged-in experiments, it caused a hiccup? [21:37:45] !log bking@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage [21:38:25] (03PS1) 10Bking: query_service: Update blazegraph exporter for newer Python library [puppet] - 10https://gerrit.wikimedia.org/r/1308778 (https://phabricator.wikimedia.org/T431057) [21:39:26] !log rzl@deploy2003 helmfile [staging-eqiad] START helmfile.d/admin 'apply'. [21:41:40] (03CR) 10Ryan Kemper: [C:03+1] query_service: Update blazegraph exporter for newer Python library [puppet] - 10https://gerrit.wikimedia.org/r/1308778 (https://phabricator.wikimedia.org/T431057) (owner: 10Bking) [21:42:02] (03CR) 10Bking: [C:03+2] query_service: Update blazegraph exporter for newer Python library [puppet] - 10https://gerrit.wikimedia.org/r/1308778 (https://phabricator.wikimedia.org/T431057) (owner: 10Bking) [21:42:07] !log rzl@deploy2003 helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. [21:44:10] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage [21:44:21] 06SRE, 06ServiceOps new, 07Service-deployment-requests: New Service Request: Headless VE - https://phabricator.wikimedia.org/T431497#12102324 (10Scott_French) Given ongoing discussion on the Headless VE service and its path to production, I'm parking this for the moment in Needs Info / Blocked for the moment... [21:46:31] (03CR) 10Eevans: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1305988 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:47:04] RESOLVED: MediaWikiEditFailures: Elevated MediaWiki edit failures (session_loss) for cluster - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://grafana.wikimedia.org/d/000000208/edit-count?orgId=1&viewPanel=13 - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiEditFailures [21:48:02] 10ops-eqiad, 06SRE, 10Data-Persistence-Backup, 06DC-Ops: db1245 crashed - https://phabricator.wikimedia.org/T431115#12102326 (10VRiley-WMF) Worked with the server and here is what I have responded to Dell with Here is the information you requested # Please confirm if there were any recent changes/update... [21:48:03] FIRING: MediaWikiEditFailures: Elevated MediaWiki edit failures (session_loss) for cluster - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://grafana.wikimedia.org/d/000000208/edit-count?orgId=1&viewPanel=13 - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiEditFailures [21:49:21] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS trixie [21:49:49] !log bking@cumin2003 START - Cookbook sre.hosts.move-vlan for host wcqs2002 [21:49:59] !log bking@cumin2003 START - Cookbook sre.dns.netbox [21:52:18] RESOLVED: MediaWikiEditFailures: Elevated MediaWiki edit failures (session_loss) for cluster - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://grafana.wikimedia.org/d/000000208/edit-count?orgId=1&viewPanel=13 - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiEditFailures [21:54:02] !log bking@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" [21:54:07] !log bking@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003" [21:54:07] !log bking@cumin2003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [21:54:08] !log bking@cumin2003 START - Cookbook sre.dns.wipe-cache wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [21:54:11] !log bking@cumin2003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [21:54:12] !log bking@cumin2003 START - Cookbook sre.network.configure-switch-interfaces for host wcqs2002 [21:54:29] !log bking@cumin2003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2002 [21:54:29] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2002 [21:55:19] !log urbanecm@deploy2003 helmfile [eqiad] START helmfile.d/services/mw-experimental: apply [21:55:23] !log cmooney@cumin1003 START - Cookbook sre.network.tls for network device lsw1-c4-codfw [21:55:32] !log cmooney@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-codfw [21:55:54] !log cmooney@cumin1003 START - Cookbook sre.network.tls for network device lsw1-d2-codfw [21:55:59] !log urbanecm@deploy2003 helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply [21:56:03] FIRING: MediaWikiEditFailures: Elevated MediaWiki edit failures (session_loss) for cluster - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://grafana.wikimedia.org/d/000000208/edit-count?orgId=1&viewPanel=13 - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiEditFailures [21:56:04] !log cmooney@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-codfw [21:56:40] FIRING: SystemdUnitFailed: fetch-rings-thanos.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [21:56:57] (03PS1) 10Urbanecm: mw-experimental: Fix MOTD typo [puppet] - 10https://gerrit.wikimedia.org/r/1308779 [22:00:05] Deploy window Readers deployment window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260708T2200) [22:01:08] !log Make https://test.wikipedia.org/w/index.php?title=MediaWiki:GrowthExperimentsSuggestedEdits.json&diff=prev&oldid=750552 with GrowthExperiments disabled (via mw-experimental), then run `\MediaWiki\MediaWikiServices::getInstance()->get('CommunityConfiguration.ProviderFactory')->newProvider('GrowthSuggestedEdits')->getStore()->invalidate()` (T431625) [22:01:11] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [22:01:12] T431625: https://test.wikipedia.org/ has been replaced by a TypeError - https://phabricator.wikimedia.org/T431625 [22:01:16] (03PS1) 10Dzahn: gerrit: replace RSA ssh host key with new ed25519 key [puppet] - 10https://gerrit.wikimedia.org/r/1282395 (https://phabricator.wikimedia.org/T240266) [22:01:16] (03CR) 10Dzahn: [C:04-1] "Hashar said " As soon as we add an OpenSSH key, that file will be ignored " https://phabricator.wikimedia.org/T240266#12088955" [puppet] - 10https://gerrit.wikimedia.org/r/1282395 (https://phabricator.wikimedia.org/T240266) (owner: 10Dzahn) [22:05:10] FIRING: [2x] SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [22:05:24] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1094.eqiad.wmnet with OS trixie [22:06:15] !log rzl@deploy2003 helmfile [codfw] START helmfile.d/admin 'apply'. [22:06:40] (03PS1) 10Btullis: datahub: tune the production OpenSearch cluster [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308780 (https://phabricator.wikimedia.org/T402408) [22:09:06] !log rzl@deploy2003 helmfile [codfw] DONE helmfile.d/admin 'apply'. [22:09:45] (03PS1) 10Urbanecm: fix: Return consistent tuple from filterLimitReachedAddLink [extensions/GrowthExperiments] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1308781 (https://phabricator.wikimedia.org/T431625) [22:09:52] jouncebot: nowandnext [22:09:52] For the next 0 hour(s) and 50 minute(s): Readers deployment window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260708T2200) [22:09:52] In 7 hour(s) and 50 minute(s): MediaWiki infrastructure (UTC early) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260709T0600) [22:09:52] In 7 hour(s) and 50 minute(s): Primary database switchover (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260709T0600) [22:10:01] (03CR) 10Btullis: [C:03+2] datahub: tune the production OpenSearch cluster [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308780 (https://phabricator.wikimedia.org/T402408) (owner: 10Btullis) [22:10:23] (03PS1) 10Urbanecm: fix: Return consistent tuple from filterLimitReachedAddLink [extensions/GrowthExperiments] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1308782 (https://phabricator.wikimedia.org/T431625) [22:10:31] (03CR) 10Urbanecm: [C:03+2] fix: Return consistent tuple from filterLimitReachedAddLink [extensions/GrowthExperiments] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1308782 (https://phabricator.wikimedia.org/T431625) (owner: 10Urbanecm) [22:10:34] (03CR) 10Urbanecm: [C:03+2] fix: Return consistent tuple from filterLimitReachedAddLink [extensions/GrowthExperiments] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1308781 (https://phabricator.wikimedia.org/T431625) (owner: 10Urbanecm) [22:11:13] (03CR) 10TrainBranchBot: [C:03+2] "Approved by urbanecm@deploy2003 using scap backport" [extensions/GrowthExperiments] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1308782 (https://phabricator.wikimedia.org/T431625) (owner: 10Urbanecm) [22:11:13] (03CR) 10TrainBranchBot: [C:03+2] "Approved by urbanecm@deploy2003 using scap backport" [extensions/GrowthExperiments] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1308781 (https://phabricator.wikimedia.org/T431625) (owner: 10Urbanecm) [22:12:50] (03Merged) 10jenkins-bot: datahub: tune the production OpenSearch cluster [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308780 (https://phabricator.wikimedia.org/T402408) (owner: 10Btullis) [22:13:41] !log rzl@deploy2003 helmfile [eqiad] START helmfile.d/admin 'apply'. [22:13:49] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply [22:13:50] !log bking@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage [22:14:27] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply} [22:14:52] (03CR) 10RLazarus: [C:03+2] "Sure is, thanks!" [puppet] - 10https://gerrit.wikimedia.org/r/1308779 (owner: 10Urbanecm) [22:16:03] RESOLVED: MediaWikiEditFailures: Elevated MediaWiki edit failures (session_loss) for cluster - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://grafana.wikimedia.org/d/000000208/edit-count?orgId=1&viewPanel=13 - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiEditFailures [22:17:09] !log rzl@deploy2003 helmfile [eqiad] DONE helmfile.d/admin 'apply'. [22:19:22] FIRING: [3x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-1/0/2 (Transport: cr1-eqiad:et-1/1/2 (Arelion, IC-374549) {#12267}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [22:19:49] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage [22:19:50] !log rzl@deploy2003 helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. [22:21:16] !log rzl@deploy2003 helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. [22:21:18] FIRING: MediaWikiEditFailures: Elevated MediaWiki edit failures (session_loss) for cluster - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://grafana.wikimedia.org/d/000000208/edit-count?orgId=1&viewPanel=13 - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiEditFailures [22:22:33] !log rzl@deploy2003 helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. [22:22:40] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for rscout - https://phabricator.wikimedia.org/T431633 (10Rscout) 03NEW [22:26:37] !log rzl@deploy2003 helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. [22:27:19] (03Merged) 10jenkins-bot: fix: Return consistent tuple from filterLimitReachedAddLink [extensions/GrowthExperiments] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1308782 (https://phabricator.wikimedia.org/T431625) (owner: 10Urbanecm) [22:27:22] (03Merged) 10jenkins-bot: fix: Return consistent tuple from filterLimitReachedAddLink [extensions/GrowthExperiments] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1308781 (https://phabricator.wikimedia.org/T431625) (owner: 10Urbanecm) [22:27:47] !log urbanecm@deploy2003 Started scap sync-world: Backport for [[gerrit:1308782|fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781|fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] [22:27:50] T431625: https://test.wikipedia.org/ has been replaced by a TypeError - https://phabricator.wikimedia.org/T431625 [22:29:08] !log rzl@deploy2003 helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. [22:30:09] !log rzl@deploy2003 helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. [22:30:47] !log rzl@deploy2003 helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. [22:32:33] !log rzl@deploy2003 helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. [22:33:28] !log rzl@deploy2003 helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. [22:33:31] !log urbanecm@deploy2003 urbanecm: Backport for [[gerrit:1308782|fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781|fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [22:33:35] T431625: https://test.wikipedia.org/ has been replaced by a TypeError - https://phabricator.wikimedia.org/T431625 [22:34:06] !log urbanecm@deploy2003 urbanecm: Continuing with deployment [22:35:34] !log rzl@deploy2003 helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. [22:36:14] !log rzl@deploy2003 helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. [22:37:47] !log rzl@deploy2003 helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. [22:40:13] !log rzl@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. [22:40:41] !log urbanecm@deploy2003 Finished scap sync-world: Backport for [[gerrit:1308782|fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781|fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] (duration: 12m 55s) [22:40:44] T431625: https://test.wikipedia.org/ has been replaced by a TypeError - https://phabricator.wikimedia.org/T431625 [22:42:04] !log rzl@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. [22:45:17] (03CR) 10RLazarus: [C:03+2] admin_ng: Install coredns-internalonly in wikikube [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305547 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [22:50:16] (03CR) 10RLazarus: admin_ng: Install coredns-internalonly in wikikube [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305547 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [22:50:18] (03CR) 10RLazarus: [C:03+2] admin_ng: Install coredns-internalonly in wikikube [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305547 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [22:51:18] RESOLVED: MediaWikiEditFailures: Elevated MediaWiki edit failures (session_loss) for cluster - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://grafana.wikimedia.org/d/000000208/edit-count?orgId=1&viewPanel=13 - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiEditFailures [22:58:38] (03CR) 10Andrea Denisse: [C:03+1] "LGTM, thank you!" [puppet] - 10https://gerrit.wikimedia.org/r/1305718 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [22:59:42] (03CR) 10Andrea Denisse: [C:03+1] "LGTM, thank you!!" [puppet] - 10https://gerrit.wikimedia.org/r/1305769 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [23:03:37] (03CR) 10RLazarus: [C:03+2] "recheck" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305547 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [23:12:50] (03PS2) 10JHathaway: Puppet 8: Replace legacy facts [puppet] - 10https://gerrit.wikimedia.org/r/1305989 (https://phabricator.wikimedia.org/T372666) [23:15:21] (03CR) 10CI reject: [V:04-1] Puppet 8: Replace legacy facts [puppet] - 10https://gerrit.wikimedia.org/r/1305989 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [23:15:21] (03CR) 10Cwhite: [C:03+1] "LGTM!" [puppet] - 10https://gerrit.wikimedia.org/r/1308721 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [23:15:51] !log rzl@deploy2003 helmfile [staging-codfw] START helmfile.d/admin 'apply'. [23:16:32] !log rzl@deploy2003 helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. [23:20:14] (03CR) 10Cwhite: "CI seems to indicate that there are (possibly unrelated) failing tests in monitoring::host." [puppet] - 10https://gerrit.wikimedia.org/r/1305989 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [23:23:23] (03CR) 10Cwhite: [C:03+1] "LGTM!" [puppet] - 10https://gerrit.wikimedia.org/r/1308750 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [23:32:13] (03PS1) 10RLazarus: coredns-internalonly: Delete nodePort [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308788 (https://phabricator.wikimedia.org/T427864) [23:40:14] !log cmooney@cumin1003 START - Cookbook sre.dns.netbox [23:42:36] (03PS1) 10TrainBranchBot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1308790 [23:42:36] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1308790 (owner: 10TrainBranchBot) [23:43:06] (03CR) 10CI reject: [V:04-1] coredns-internalonly: Delete nodePort [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308788 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [23:46:36] !log cmooney@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" [23:46:44] !log cmooney@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003" [23:46:44] !log cmooney@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [23:47:09] RESOLVED: CirrusSearchNodeIndexingNotIncreasing: Elasticsearch instance cirrussearch1094-production-search-eqiad is not indexing - https://wikitech.wikimedia.org/wiki/Search/Elasticsearch_Administration#Indexing_hung_and_not_making_progress - https://grafana.wikimedia.org/d/JLK3I_siz/elasticsearch-indexing?orgId=1&from=now-3d&to=now&viewPanel=57 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchNodeIndexingNotIncreasing [23:49:59] (03CR) 10CI reject: [V:04-1] Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1308790 (owner: 10TrainBranchBot) [23:51:36] 06SRE, 06Data-Engineering, 06Data-Platform-SRE (2026-07-03 - 2026-07-31), 07Sustainability (Incident Followup): The webrequest_sampled_live data pipeline and its query tools have become mission-critical and require re-engineering for resilience - https://phabricator.wikimedia.org/T431112#12102634 (10Kappaka... [23:52:15] 06SRE, 06Data-Engineering, 06Data-Platform-SRE (2026-07-03 - 2026-07-31), 07Sustainability (Incident Followup): The webrequest_sampled_live data pipeline and its query tools have become mission-critical and require re-engineering for resilience - https://phabricator.wikimedia.org/T431112#12102635 (10Kap... [23:52:47] !log ladsgroup@deploy2003:~$ mwscript-k8s --follow -- extensions/ORES/maintenance/PurgeScoreCache.php --wiki=simplewiki --model damaging --old (T431159) [23:52:50] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [23:52:50] T431159: ORES flagging not working on simplewiki - https://phabricator.wikimedia.org/T431159