[00:06:04] (03PS2) 10Jasmine: sre.k8s.renumber-node: Increase timeout on deploy hosts puppet run to 600s [cookbooks] - 10https://gerrit.wikimedia.org/r/1307872 (https://phabricator.wikimedia.org/T430226) [00:13:26] (03PS3) 10Jasmine: sre.k8s.renumber-node: Increase timeout on 600s for puppet run on deploy hosts [cookbooks] - 10https://gerrit.wikimedia.org/r/1307872 (https://phabricator.wikimedia.org/T430226) [00:14:22] (03PS4) 10Jasmine: sre.k8s.renumber-node: Increase timeout to 600s for puppet runs on deploy hosts [cookbooks] - 10https://gerrit.wikimedia.org/r/1307872 (https://phabricator.wikimedia.org/T430226) [00:25:42] PROBLEM - Dell PowerEdge or Supermicro Broadcom RAID Controller on backup1019 is CRITICAL: communication: 0 OK : controller: 1 Needs Attention : physical_disk: 0 OK : virtual_disk: 1 Dgrd : bbu: 0 OK : enclosure: 0 OK : CLI Version = 007.1910.0000.0000 Oct 08, 2021 https://wikitech.wikimedia.org/wiki/PERCCli%23Monitoring [00:25:44] ACKNOWLEDGEMENT - Dell PowerEdge or Supermicro Broadcom RAID Controller on backup1019 is CRITICAL: communication: 0 OK : controller: 1 Needs Attention : physical_disk: 0 OK : virtual_disk: 1 Dgrd : bbu: 0 OK : enclosure: 0 OK : CLI Version = 007.1910.0000.0000 Oct 08, 2021 nagiosadmin RAID handler auto-ack: https://phabricator.wikimedia.org/T431367 https://wikitech.wikimedia.org/wiki/PERCCli%23Monitoring [00:25:56] 10ops-eqiad, 06SRE, 06DC-Ops: Degraded RAID on backup1019 - https://phabricator.wikimedia.org/T431367 (10ops-monitoring-bot) 03NEW [00:37:04] 06SRE, 10SRE-Access-Requests: Grant Access to fr-tech groups for laurabarluzzi - https://phabricator.wikimedia.org/T431338#12092763 (10Dzahn) [00:37:33] 06SRE, 10SRE-Access-Requests: Grant Access to fr-tech groups for laurabarluzzi - https://phabricator.wikimedia.org/T431338#12092764 (10Dzahn) [00:39:21] FIRING: [9x] SLOBudgetBurn: Search update lag is below 95% target in eqiad - https://alerts.wikimedia.org/?q=alertname%3DSLOBudgetBurn [00:41:51] 06SRE, 10SRE-Access-Requests: Grant Access to fr-tech groups for laurabarluzzi - https://phabricator.wikimedia.org/T431338#12092767 (10Dzahn) Thanks so much for these clarifying additions, Aklapper. @Laurabarluzzi @Dwisehaupt Alright, now that it's clear that this is a shell access request, (as opposed to L... [00:42:45] 06SRE, 10SRE-Access-Requests: Grant Access to fr-tech groups for laurabarluzzi - https://phabricator.wikimedia.org/T431338#12092768 (10Dzahn) [00:44:21] FIRING: [9x] SLOBudgetBurn: Search update lag is below 95% target in eqiad - https://alerts.wikimedia.org/?q=alertname%3DSLOBudgetBurn [00:47:43] 06SRE, 10SRE-Access-Requests: Grant Access to fr-tech groups for laurabarluzzi - https://phabricator.wikimedia.org/T431338#12092770 (10Dzahn) @Laurabarluzzi Hi there, please take a look at the ticket now. I pre-filled your user name, email, approving party and reason. What we will need next are the remaini... [00:49:21] FIRING: [9x] SLOBudgetBurn: Search update lag is below 95% target in eqiad - https://alerts.wikimedia.org/?q=alertname%3DSLOBudgetBurn [00:53:59] 06SRE, 10SRE-Access-Requests: Grant Access to fr-tech groups for laurabarluzzi - https://phabricator.wikimedia.org/T431338#12092772 (10Dzahn) @Dwisehaupt @Damilare Hi, can we clarify the requested groups please? If I look strictly at "what Damilare has" as it was stated in the ticket then this means `fr-tech-... [00:54:21] RESOLVED: [9x] SLOBudgetBurn: Search update lag is below 95% target in eqiad - https://alerts.wikimedia.org/?q=alertname%3DSLOBudgetBurn [00:54:38] 06SRE, 10SRE-Access-Requests: Grant Access to fr-tech groups for laurabarluzzi - https://phabricator.wikimedia.org/T431338#12092773 (10Dzahn) [00:57:37] 06SRE, 10SRE-Access-Requests: Grant Access to fr-tech groups for laurabarluzzi - https://phabricator.wikimedia.org/T431338#12092774 (10Dzahn) 05Open→03In progress Ok, we should be good to go here now. Rotating SRE clinic duty may continue on this ticket as well. One more thing: Please get a manager from f... [00:58:32] 06SRE, 10SRE-Access-Requests: Grant shell access (fr-tech-devs) for laurabarluzzi - https://phabricator.wikimedia.org/T431338#12092780 (10Dzahn) [01:05:55] (03CR) 10Dzahn: "this will be at the beginning of the maintenance on July 7th 2026" [puppet] - 10https://gerrit.wikimedia.org/r/1307850 (https://phabricator.wikimedia.org/T418109) (owner: 10Dzahn) [01:11:17] (03PS1) 10TrainBranchBot: Branch commit for wmf/1.47.0-wmf.10 [core] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1307878 (https://phabricator.wikimedia.org/T430829) [01:11:19] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/1.47.0-wmf.10 [core] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1307878 (https://phabricator.wikimedia.org/T430829) (owner: 10TrainBranchBot) [01:12:15] (03PS1) 10TrainBranchBot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1307879 [01:12:15] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1307879 (owner: 10TrainBranchBot) [01:18:29] (03Merged) 10jenkins-bot: Branch commit for wmf/1.47.0-wmf.10 [core] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1307878 (https://phabricator.wikimedia.org/T430829) (owner: 10TrainBranchBot) [01:19:09] (03CR) 10Dzahn: [V:03+1] "the main switch for maintenance window July 7th" [puppet] - 10https://gerrit.wikimedia.org/r/1306446 (https://phabricator.wikimedia.org/T418521) (owner: 10Dzahn) [01:20:27] (03Merged) 10jenkins-bot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1307879 (owner: 10TrainBranchBot) [01:28:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [01:41:27] (03PS1) 10Dzahn: Revert^2 "integration: switch integration-agent-docker VMs to Java 21" [puppet] - 10https://gerrit.wikimedia.org/r/1307882 [01:43:14] (03PS1) 10Dzahn: Revert^2 "integration: switch integration-agent-docker VMs to Java 21" [puppet] - 10https://gerrit.wikimedia.org/r/1307883 [02:00:05] Deploy window Automatic branching of MediaWiki, extensions, skins, and vendor – see Heterogeneous deployment/Train deploys (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260707T0200) [02:00:36] !log mwpresync@deploy1003 Started scap build-images: Publishing wmf/next image [02:07:27] !log mwpresync@deploy1003 Finished scap build-images: Publishing wmf/next image (duration: 06m 51s) [02:09:22] FIRING: [2x] JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:14:22] RESOLVED: [2x] JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [03:00:05] Deploy window Automatic deployment of MediaWiki, extensions, skins, and vendor to testwikis only – see Heterogeneous deployment/Train deploys (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260707T0300) [03:01:50] (03PS1) 10TrainBranchBot: testwikis to 1.47.0-wmf.10 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1307887 (https://phabricator.wikimedia.org/T430829) [03:01:53] (03CR) 10TrainBranchBot: [C:03+2] "Initiated by mwpresync@deploy1003" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1307887 (https://phabricator.wikimedia.org/T430829) (owner: 10TrainBranchBot) [03:02:56] (03Merged) 10jenkins-bot: testwikis to 1.47.0-wmf.10 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1307887 (https://phabricator.wikimedia.org/T430829) (owner: 10TrainBranchBot) [03:03:20] !log mwpresync@deploy1003 Started scap sync-world: testwikis to 1.47.0-wmf.10 refs T430829 [03:03:23] T430829: 1.47.0-wmf.10 deployment blockers - https://phabricator.wikimedia.org/T430829 [03:03:44] FIRING: RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [03:08:44] RESOLVED: RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [03:40:24] !log mwpresync@deploy1003 Finished scap sync-world: testwikis to 1.47.0-wmf.10 refs T430829 (duration: 37m 04s) [03:40:27] T430829: 1.47.0-wmf.10 deployment blockers - https://phabricator.wikimedia.org/T430829 [04:00:04] Deploy window Automatic removal of all obsolete MediaWiki versions from the deployment and bare metal servers (except the most-recent obsolete version) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260707T0400) [04:02:46] !log mwpresync@deploy1003 Pruned MediaWiki: 1.47.0-wmf.7 (duration: 02m 38s) [05:27:38] (03CR) 10Ayounsi: [C:03+1] "Thanks !" [puppet] - 10https://gerrit.wikimedia.org/r/1307852 (https://phabricator.wikimedia.org/T430930) (owner: 10Scott French) [05:28:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [05:30:10] (03PS1) 10Marostegui: tables-catalog.yaml: Add oathauth_user_handles [puppet] - 10https://gerrit.wikimedia.org/r/1308006 (https://phabricator.wikimedia.org/T431124) [05:32:19] (03CR) 10CI reject: [V:04-1] tables-catalog.yaml: Add oathauth_user_handles [puppet] - 10https://gerrit.wikimedia.org/r/1308006 (https://phabricator.wikimedia.org/T431124) (owner: 10Marostegui) [05:37:55] (03PS2) 10Marostegui: tables-catalog.yaml: Add oathauth_user_handles [puppet] - 10https://gerrit.wikimedia.org/r/1308006 (https://phabricator.wikimedia.org/T431124) [05:59:22] FIRING: [2x] CoreRouterInterfaceDown: Core router interface down - cr2-codfw:xe-0/0/1:1 (Transport: cr2-eqord:xe-0/1/0 (Arelion, IC-314534 29ms 10Gbps wave) {#10694_12249-2}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [06:00:04] Deploy window MediaWiki infrastructure (UTC early) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260707T0600) [06:00:04] marostegui, Amir1, and federico3: Time to do the Primary database switchover deploy. Don't look at me like that. You signed up for it. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260707T0600). [06:04:56] (03PS1) 10Muehlenhoff: Record LDAP access for segt [puppet] - 10https://gerrit.wikimedia.org/r/1308007 [06:07:20] FIRING: CirrusSearchSaneitizerFixRateTooHigh: MediaWiki CirrusSearch Saneitizer is fixing an abnormally high number of documents in eqiad - https://wikitech.wikimedia.org/wiki/Search/CirrusStreamingUpdater#San(e)itizing - https://grafana.wikimedia.org/d/JLK3I_siz/elasticsearch-indexing?viewPanel=59&orgId=1&from=now-6M&to=now&var-search_cluster=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchSaneitizerFixRateTooHigh [06:08:21] (03CR) 10Muehlenhoff: [C:03+2] Record LDAP access for segt [puppet] - 10https://gerrit.wikimedia.org/r/1308007 (owner: 10Muehlenhoff) [06:12:49] 06SRE, 06Data-Engineering, 06Data-Platform-SRE (2026-07-03 - 2026-07-31), 07Sustainability (Incident Followup): The webrequest_sampled_live data pipeline and its query tools have become mission-critical and require re-engineering for resilience - https://phabricator.wikimedia.org/T431112#12093090 (10elukey... [06:13:43] 06SRE, 06Infrastructure-Foundations, 13Patch-For-Review, 06Release-Engineering-Team (Radar), 07User-notice: Sunsetting mirrors.wikimedia.org - https://phabricator.wikimedia.org/T416707#12093097 (10MoritzMuehlenhoff) 05Open→03Resolved a:03MoritzMuehlenhoff mirrors.wikimedia.org has been deregist... [06:15:52] !log installing php8.4 security updates [06:15:53] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [06:19:21] !log installing php8.2 security updates [06:19:21] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [06:21:06] !log root@cumin1003 START - Cookbook sre.mysql.depool depool pc1024: Security updates [06:29:22] RESOLVED: [2x] CoreRouterInterfaceDown: Core router interface down - cr2-codfw:xe-0/0/1:1 (Transport: cr2-eqord:xe-0/1/0 (Arelion, IC-314534 29ms 10Gbps wave) {#10694_12249-2}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [06:31:20] !log root@cumin1003 END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool pc1024: Security updates [06:42:09] !log install nginx security updates [06:42:10] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [06:44:58] (03PS1) 10Kevin Bazira: ml-services: deploy tts isvc in experimental ns [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308012 (https://phabricator.wikimedia.org/T430536) [06:48:19] !log root@cumin1003 START - Cookbook sre.mysql.depool depool pc1024: Security updates [06:48:20] !log root@cumin1003 START - Cookbook sre.mysql.parsercache [06:48:27] !log root@cumin1003 END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) [06:48:28] !log root@cumin1003 END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1024: Security updates [06:53:48] PROBLEM - MariaDB Replica IO: pc4 on pc2024 is CRITICAL: CRITICAL slave_io_state Slave_IO_Running: No, Errno: 2003, Errmsg: error reconnecting to master repl2024@pc1024.eqiad.wmnet:3306 - retry-time: 60 maximum-retries: 100000 message: Cant connect to server on pc1024.eqiad.wmnet (111 Connection refused) https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [06:55:48] RECOVERY - MariaDB Replica IO: pc4 on pc2024 is OK: OK slave_io_state Slave_IO_Running: Yes https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [06:55:50] !log elukey@cumin1003 START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet [06:55:51] !log elukey@cumin1003 END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet [06:57:13] (03PS15) 10Elukey: Add sre.hosts.bmc-user-mgmt.py [cookbooks] - 10https://gerrit.wikimedia.org/r/1302859 (https://phabricator.wikimedia.org/T426180) [07:00:05] Amir1, urbanecm, and awight: How many deployers does it take to do UTC morning backport window deploy? (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260707T0700). [07:00:05] No Gerrit patches in the queue for this window AFAICS. [07:02:06] (03PS1) 10Muehlenhoff: Remove ml-cache Cumin alias [puppet] - 10https://gerrit.wikimedia.org/r/1308013 (https://phabricator.wikimedia.org/T430654) [07:02:47] (03PS2) 10Muehlenhoff: Remove role:mirror and related Puppet classes [puppet] - 10https://gerrit.wikimedia.org/r/1307383 (https://phabricator.wikimedia.org/T416707) [07:06:15] (03PS1) 10Giuseppe Lavagetto: UI warnings for known fingerprints [software/hiddenparma/deploy] - 10https://gerrit.wikimedia.org/r/1308015 [07:06:35] (03CR) 10Giuseppe Lavagetto: [V:03+2 C:03+2] UI warnings for known fingerprints [software/hiddenparma/deploy] - 10https://gerrit.wikimedia.org/r/1308015 (owner: 10Giuseppe Lavagetto) [07:10:52] !log oblivian@cumin1003 START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" [07:10:53] !log oblivian@cumin1003 START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 [07:11:06] !log root@cumin1003 START - Cookbook sre.mysql.pool pool pc1024: Security updates [07:11:06] !log root@cumin1003 START - Cookbook sre.mysql.parsercache [07:11:19] !log root@cumin1003 END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) [07:11:19] !log root@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1024: Security updates [07:11:50] !log oblivian@cumin1003 END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003 [07:11:51] !log oblivian@cumin1003 END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003" [07:12:43] (03CR) 10Filippo Giunchedi: [C:03+1] profile::benthos: add x-provenance handling for webrequest [puppet] - 10https://gerrit.wikimedia.org/r/1307169 (https://phabricator.wikimedia.org/T427068) (owner: 10Elukey) [07:15:39] (03Abandoned) 10Hashar: Revert^2 "integration: switch integration-agent-docker VMs to Java 21" [puppet] - 10https://gerrit.wikimedia.org/r/1307883 (owner: 10Dzahn) [07:15:40] (03Abandoned) 10Hashar: Revert^2 "integration: switch integration-agent-docker VMs to Java 21" [puppet] - 10https://gerrit.wikimedia.org/r/1307882 (owner: 10Dzahn) [07:15:57] (03PS2) 10Hashar: Revert^2 "integration: switch integration-agent-docker VMs to Java 21" [puppet] - 10https://gerrit.wikimedia.org/r/1307851 (owner: 10Dzahn) [07:16:16] (03CR) 10Hashar: [C:03+1] Revert^2 "integration: switch integration-agent-docker VMs to Java 21" [puppet] - 10https://gerrit.wikimedia.org/r/1307851 (owner: 10Dzahn) [07:21:30] Hi! I'll have a private code change to deploy. Is something being deployed now? [07:21:46] FIRING: [4x] GerritHAProxyBackendUnavailable: Gerrit backend is unavilable for tcp-proxy (HAProxy) gerrit_ssh - https://wikitech.wikimedia.org/wiki/Gerrit/Operations#GerritHAProxyBackendUnavailable - grafana.wikimedia.org/d/459365f6-df37-48d6-8142-82b22c1875e7/gerrit-tcp-proxy?viewPanel=panel-15 - https://alerts.wikimedia.org/?q=alertname%3DGerritHAProxyBackendUnavailable [07:21:51] FIRING: [2x] GerritHAProxyServiceUnavailable: Gerrit tcp-proxy (HAProxy) service gerrit_ssh is DOWN in codfw - https://wikitech.wikimedia.org/wiki/Gerrit/Operations#GerritHAProxyServiceUnavailable - grafana.wikimedia.org/d/459365f6-df37-48d6-8142-82b22c1875e7/gerrit-tcp-proxy?viewPanel=panel-15 - https://alerts.wikimedia.org/?q=alertname%3DGerritHAProxyServiceUnavailable [07:22:40] FIRING: [4x] ProbeDown: Service gerrit2003:29418 has failed probes (tcp_gerrit_ssh_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [07:23:04] I'm assuming the window is free, proceeding [07:26:46] RESOLVED: [4x] GerritHAProxyBackendUnavailable: Gerrit backend is unavilable for tcp-proxy (HAProxy) gerrit_ssh - https://wikitech.wikimedia.org/wiki/Gerrit/Operations#GerritHAProxyBackendUnavailable - grafana.wikimedia.org/d/459365f6-df37-48d6-8142-82b22c1875e7/gerrit-tcp-proxy?viewPanel=panel-15 - https://alerts.wikimedia.org/?q=alertname%3DGerritHAProxyBackendUnavailable [07:26:57] RESOLVED: [2x] GerritHAProxyServiceUnavailable: Gerrit tcp-proxy (HAProxy) service gerrit_ssh is DOWN in codfw - https://wikitech.wikimedia.org/wiki/Gerrit/Operations#GerritHAProxyServiceUnavailable - grafana.wikimedia.org/d/459365f6-df37-48d6-8142-82b22c1875e7/gerrit-tcp-proxy?viewPanel=panel-15 - https://alerts.wikimedia.org/?q=alertname%3DGerritHAProxyServiceUnavailable [07:27:31] RESOLVED: [4x] ProbeDown: Service gerrit2003:29418 has failed probes (tcp_gerrit_ssh_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [07:29:45] (03CR) 10Bartosz Wójtowicz: [C:03+1] "LGTM, exciting!!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308012 (https://phabricator.wikimedia.org/T430536) (owner: 10Kevin Bazira) [07:34:51] (03CR) 10Kevin Bazira: [C:03+2] ml-services: deploy tts isvc in experimental ns [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308012 (https://phabricator.wikimedia.org/T430536) (owner: 10Kevin Bazira) [07:35:50] (03CR) 10Kosta Harlan: [C:03+1] maintain-views: Make change_tag_def and change_tag partially public [puppet] - 10https://gerrit.wikimedia.org/r/1307100 (https://phabricator.wikimedia.org/T386456) (owner: 10Dreamy Jazz) [07:37:08] (03Merged) 10jenkins-bot: ml-services: deploy tts isvc in experimental ns [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308012 (https://phabricator.wikimedia.org/T430536) (owner: 10Kevin Bazira) [07:40:07] !log kevinbazira@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . [07:41:31] !log root@cumin1003 START - Cookbook sre.mysql.depool depool pc1015: Security updates [07:41:31] !log root@cumin1003 START - Cookbook sre.mysql.parsercache [07:41:38] !log root@cumin1003 END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) [07:41:38] !log root@cumin1003 END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1015: Security updates [07:42:51] !log Deployed private patch for Suggested Ivestigations [07:42:52] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [07:43:46] FIRING: GerritDiskSpaceExhaustionIncoming: Gerrit disk space runway on gerrit2003:/srv is too low - https://wikitech.wikimedia.org/wiki/Gerrit/Operations#TODO - https://grafana.wikimedia.org/d/rYdddlPWk/node-exporter-collaboration-services?var-job=node&var-nodename=gerrit2003&var-node=gerrit2003%3A9100&refresh=1m - https://alerts.wikimedia.org/?q=alertname%3DGerritDiskSpaceExhaustionIncoming [07:47:40] PROBLEM - MariaDB Replica IO: pc5 on pc2015 is CRITICAL: CRITICAL slave_io_state Slave_IO_Running: No, Errno: 2003, Errmsg: error reconnecting to master repl2024@pc1015.eqiad.wmnet:3306 - retry-time: 60 maximum-retries: 100000 message: Cant connect to server on pc1015.eqiad.wmnet (111 Connection refused) https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [07:52:40] RECOVERY - MariaDB Replica IO: pc5 on pc2015 is OK: OK slave_io_state Slave_IO_Running: Yes https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [07:56:05] (03PS1) 10Ayounsi: Add depool policy for restbase/sessionstore/cassandra_dev [puppet] - 10https://gerrit.wikimedia.org/r/1308030 (https://phabricator.wikimedia.org/T327300) [07:58:32] (03PS2) 10Klausman: home/klausman: add tmuxp recipe for dealing with admin_ng [puppet] - 10https://gerrit.wikimedia.org/r/1308032 [07:58:38] (03CR) 10Klausman: [V:03+2 C:03+2] home/klausman: add tmuxp recipe for dealing with admin_ng [puppet] - 10https://gerrit.wikimedia.org/r/1308032 (owner: 10Klausman) [08:00:34] (03PS1) 10Jgiannelos: mwparsoid: Mount /srv/mediawiki from mw-experimental [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308029 [08:01:53] (03CR) 10Muehlenhoff: [C:03+2] Remove role:mirror and related Puppet classes [puppet] - 10https://gerrit.wikimedia.org/r/1307383 (https://phabricator.wikimedia.org/T416707) (owner: 10Muehlenhoff) [08:09:31] (03PS1) 10Arnaudb: gerrit: predictive diskspace alert threshold adjustment [alerts] - 10https://gerrit.wikimedia.org/r/1308034 (https://phabricator.wikimedia.org/T431382) [08:09:39] !log root@cumin1003 START - Cookbook sre.mysql.pool pool pc1015: Security updates [08:09:39] !log root@cumin1003 START - Cookbook sre.mysql.parsercache [08:09:51] !log root@cumin1003 END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) [08:09:52] !log root@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1015: Security updates [08:12:00] (03CR) 10Arnaudb: [C:03+2] gerrit: predictive diskspace alert threshold adjustment [alerts] - 10https://gerrit.wikimedia.org/r/1308034 (https://phabricator.wikimedia.org/T431382) (owner: 10Arnaudb) [08:13:46] RESOLVED: GerritDiskSpaceExhaustionIncoming: Gerrit disk space runway on gerrit2003:/srv is too low - https://wikitech.wikimedia.org/wiki/Gerrit/Operations#TODO - https://grafana.wikimedia.org/d/rYdddlPWk/node-exporter-collaboration-services?var-job=node&var-nodename=gerrit2003&var-node=gerrit2003%3A9100&refresh=1m - https://alerts.wikimedia.org/?q=alertname%3DGerritDiskSpaceExhaustionIncoming [08:14:15] (03PS1) 10Bartosz Wójtowicz: ml-services: Update qwen3-14b MODEL_NAME env var. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308036 (https://phabricator.wikimedia.org/T431268) [08:17:14] (03PS16) 10Elukey: Add sre.hosts.bmc-user-mgmt.py [cookbooks] - 10https://gerrit.wikimedia.org/r/1302859 (https://phabricator.wikimedia.org/T426180) [08:17:27] !log elukey@cumin1003 START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet [08:17:29] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet [08:24:15] (03PS17) 10Elukey: Add sre.hosts.bmc-user-mgmt.py [cookbooks] - 10https://gerrit.wikimedia.org/r/1302859 (https://phabricator.wikimedia.org/T426180) [08:24:22] !log elukey@cumin1003 START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet [08:24:24] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet [08:25:01] (03PS18) 10Elukey: Add sre.hosts.bmc-user-mgmt.py [cookbooks] - 10https://gerrit.wikimedia.org/r/1302859 (https://phabricator.wikimedia.org/T426180) [08:25:06] !log elukey@cumin1003 START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet [08:25:09] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet [08:27:52] (03PS19) 10Elukey: Add sre.hosts.bmc-user-mgmt.py [cookbooks] - 10https://gerrit.wikimedia.org/r/1302859 (https://phabricator.wikimedia.org/T426180) [08:27:56] !log elukey@cumin1003 START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet [08:27:59] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet [08:29:31] (03PS20) 10Elukey: Add sre.hosts.bmc-user-mgmt.py [cookbooks] - 10https://gerrit.wikimedia.org/r/1302859 (https://phabricator.wikimedia.org/T426180) [08:29:37] !log elukey@cumin1003 START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet [08:29:41] !log elukey@cumin1003 END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet [08:30:20] (03CR) 10Filippo Giunchedi: [C:03+2] dumps-nfs: re-use dumps.w.o address [puppet] - 10https://gerrit.wikimedia.org/r/1307338 (https://phabricator.wikimedia.org/T411248) (owner: 10Filippo Giunchedi) [08:31:42] (03CR) 10Ayounsi: [C:03+1] "that's great !" [homer/public] - 10https://gerrit.wikimedia.org/r/1284695 (https://phabricator.wikimedia.org/T425703) (owner: 10Cathal Mooney) [08:35:26] PROBLEM - PyBal IPVS diff check on lvs1020 is CRITICAL: (CRITICAL: Mismatch between IPVS and PyBal https://wikitech.wikimedia.org/wiki/PyBal [08:35:45] that's me ^ [08:37:36] (03PS1) 10Fabfur: cache_misc: apply traffic classification [puppet] - 10https://gerrit.wikimedia.org/r/1308040 [08:38:11] (03CR) 10CI reject: [V:04-1] cache_misc: apply traffic classification [puppet] - 10https://gerrit.wikimedia.org/r/1308040 (owner: 10Fabfur) [08:38:35] should be recovering at the next check [08:40:04] !log root@cumin1003 START - Cookbook sre.mysql.depool depool pc1016: Security updates [08:40:04] !log root@cumin1003 START - Cookbook sre.mysql.parsercache [08:40:12] !log root@cumin1003 END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) [08:40:12] !log root@cumin1003 END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates [08:40:24] RECOVERY - PyBal IPVS diff check on lvs1020 is OK: OK: no difference between hosts in IPVS/PyBal https://wikitech.wikimedia.org/wiki/PyBal [08:42:26] !log switch dumps-nfs address to be shared with rsync/http - T411248 [08:42:29] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [08:42:30] T411248: Plan to make clouddumps more resilient and easier to operate - https://phabricator.wikimedia.org/T411248 [08:42:34] (03CR) 10Filippo Giunchedi: [C:03+2] wikimedia.org: move dumps-nfs to dumps-lb [dns] - 10https://gerrit.wikimedia.org/r/1307339 (https://phabricator.wikimedia.org/T411248) (owner: 10Filippo Giunchedi) [08:42:52] (03CR) 10JMeybohm: "This does only change summary, description and dashboard links but not the actual alert expression(s). Or am I missing something?" [alerts] - 10https://gerrit.wikimedia.org/r/1307787 (https://phabricator.wikimedia.org/T430290) (owner: 10Blake) [08:43:05] (03PS2) 10Majavah: dynamicproxy: Update schema to only support a single backend [puppet] - 10https://gerrit.wikimedia.org/r/1307714 (https://phabricator.wikimedia.org/T429960) [08:43:26] !log filippo@dns1006 START - running authdns-update [08:44:59] (03PS2) 10Fabfur: cache_misc: apply traffic classification [puppet] - 10https://gerrit.wikimedia.org/r/1308040 [08:45:07] !log filippo@dns1006 END - running authdns-update [08:46:28] PROBLEM - MariaDB Replica IO: pc6 on pc2016 is CRITICAL: CRITICAL slave_io_state Slave_IO_Running: No, Errno: 2003, Errmsg: error reconnecting to master repl2024@pc1016.eqiad.wmnet:3306 - retry-time: 60 maximum-retries: 100000 message: Cant connect to server on pc1016.eqiad.wmnet (111 Connection refused) https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [08:49:04] (03CR) 10Filippo Giunchedi: "In addition to what Tobias said re: preseed-test, to ensure a bigger-than-standard LV size I would recommend lvm::logical_volume with size" [puppet] - 10https://gerrit.wikimedia.org/r/1307368 (https://phabricator.wikimedia.org/T431017) (owner: 10Dpogorzelski) [08:50:28] RECOVERY - MariaDB Replica IO: pc6 on pc2016 is OK: OK slave_io_state Slave_IO_Running: Yes https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [08:50:31] !log javiermonton@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply [08:50:54] !log javiermonton@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply [08:51:36] (03CR) 10Filippo Giunchedi: [C:03+1] memcached: Allow disabling Icinga checks [puppet] - 10https://gerrit.wikimedia.org/r/1307763 (https://phabricator.wikimedia.org/T384305) (owner: 10Majavah) [08:52:15] (03CR) 10Filippo Giunchedi: [C:03+1] Disable Memcached Icinga checks for cloudcontrols/cloudwebs [puppet] - 10https://gerrit.wikimedia.org/r/1307764 (https://phabricator.wikimedia.org/T384305) (owner: 10Majavah) [08:52:52] (03CR) 10Filippo Giunchedi: [C:03+1] P:wmcs::novaproxy: Install HAProxy service [puppet] - 10https://gerrit.wikimedia.org/r/1307782 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [08:53:13] (03CR) 10Majavah: [V:03+1 C:03+2] memcached: Allow disabling Icinga checks [puppet] - 10https://gerrit.wikimedia.org/r/1307763 (https://phabricator.wikimedia.org/T384305) (owner: 10Majavah) [08:53:24] (03CR) 10Majavah: [C:03+2] Disable Memcached Icinga checks for cloudcontrols/cloudwebs [puppet] - 10https://gerrit.wikimedia.org/r/1307764 (https://phabricator.wikimedia.org/T384305) (owner: 10Majavah) [08:53:37] (03CR) 10Filippo Giunchedi: [C:03+1] P:wmcs::novaproxy: Handle HTTP redirects with HAProxy [puppet] - 10https://gerrit.wikimedia.org/r/1307784 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [08:55:20] (03CR) 10Filippo Giunchedi: [C:03+1] dynamicproxy: Update schema to only support a single backend [puppet] - 10https://gerrit.wikimedia.org/r/1307714 (https://phabricator.wikimedia.org/T429960) (owner: 10Majavah) [08:56:48] !log root@cumin1003 START - Cookbook sre.mysql.depool depool pc1016: Security updates [08:56:49] !log root@cumin1003 START - Cookbook sre.mysql.parsercache [08:56:54] !log root@cumin1003 END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) [08:56:54] !log root@cumin1003 END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates [09:01:09] (03CR) 10Majavah: [C:03+2] P:wmcs::novaproxy: Install HAProxy service [puppet] - 10https://gerrit.wikimedia.org/r/1307782 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [09:02:28] PROBLEM - MariaDB Replica IO: pc6 on pc2016 is CRITICAL: CRITICAL slave_io_state Slave_IO_Running: No, Errno: 2003, Errmsg: error reconnecting to master repl2024@pc1016.eqiad.wmnet:3306 - retry-time: 60 maximum-retries: 100000 message: Cant connect to server on pc1016.eqiad.wmnet (111 Connection refused) https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [09:05:28] RECOVERY - MariaDB Replica IO: pc6 on pc2016 is OK: OK slave_io_state Slave_IO_Running: Yes https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [09:14:58] !log cwilliams@cumin1003 START - Cookbook sre.mysql.parsercache [09:14:59] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) [09:19:31] (03CR) 10Hnowlan: [V:03+1] "PCC SUCCESS (DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/8934/console" [puppet] - 10https://gerrit.wikimedia.org/r/1307104 (https://phabricator.wikimedia.org/T407117) (owner: 10Hnowlan) [09:19:42] (03PS9) 10Blake: bgp: Remove :0 from instance and remote_instance names. [alerts] - 10https://gerrit.wikimedia.org/r/1307787 (https://phabricator.wikimedia.org/T430290) [09:20:17] (03PS10) 10Blake: bgp: Remove port from instance and remote_instance names. Relabel instance into hostname. [alerts] - 10https://gerrit.wikimedia.org/r/1307787 (https://phabricator.wikimedia.org/T430290) [09:20:18] (03CR) 10Hnowlan: [V:03+1 C:03+2] cassandra: remove obsolete expiry check [puppet] - 10https://gerrit.wikimedia.org/r/1307104 (https://phabricator.wikimedia.org/T407117) (owner: 10Hnowlan) [09:21:36] (03CR) 10Hnowlan: [C:03+1] logstash: add parameters needed for security plugin [puppet] - 10https://gerrit.wikimedia.org/r/1305769 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [09:21:46] (03PS4) 10Majavah: P:wmcs::novaproxy: Handle HTTP redirects with HAProxy [puppet] - 10https://gerrit.wikimedia.org/r/1307784 (https://phabricator.wikimedia.org/T429930) [09:21:46] (03PS1) 10Majavah: P:wmcs::novaproxy: Fix Prometheus node variable name [puppet] - 10https://gerrit.wikimedia.org/r/1308047 [09:21:55] !log root@cumin1003 START - Cookbook sre.mysql.pool pool pc1016: Security updates [09:21:55] !log root@cumin1003 START - Cookbook sre.mysql.parsercache [09:22:09] !log root@cumin1003 END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) [09:22:09] !log root@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1016: Security updates [09:22:28] (03CR) 10Hnowlan: [C:03+2] restbase: move nrpe check to prom blackbox check [puppet] - 10https://gerrit.wikimedia.org/r/1306661 (https://phabricator.wikimedia.org/T407141) (owner: 10Hnowlan) [09:22:59] (03CR) 10CI reject: [V:04-1] bgp: Remove port from instance and remote_instance names. Relabel instance into hostname. [alerts] - 10https://gerrit.wikimedia.org/r/1307787 (https://phabricator.wikimedia.org/T430290) (owner: 10Blake) [09:23:39] (03PS1) 10Hnowlan: Revert "restbase: move nrpe check to prom blackbox check" [puppet] - 10https://gerrit.wikimedia.org/r/1308048 [09:24:11] (03PS2) 10Majavah: P:wmcs::novaproxy: Fix Prometheus node variable name [puppet] - 10https://gerrit.wikimedia.org/r/1308047 (https://phabricator.wikimedia.org/T429930) [09:24:14] (03PS5) 10Majavah: P:wmcs::novaproxy: Handle HTTP redirects with HAProxy [puppet] - 10https://gerrit.wikimedia.org/r/1307784 (https://phabricator.wikimedia.org/T429930) [09:25:18] (03CR) 10CWilliams: [C:03+2] sre.mysql.pool: --skip-safety-checks fails for parsercache hosts [cookbooks] - 10https://gerrit.wikimedia.org/r/1307815 (https://phabricator.wikimedia.org/T429878) (owner: 10CWilliams) [09:25:20] (03PS11) 10Blake: bgp: Remove port from instance and remote_instance names. Relabel instance into hostname. [alerts] - 10https://gerrit.wikimedia.org/r/1307787 (https://phabricator.wikimedia.org/T430290) [09:27:29] (03CR) 10Majavah: [C:03+2] P:wmcs::novaproxy: Fix Prometheus node variable name [puppet] - 10https://gerrit.wikimedia.org/r/1308047 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [09:27:40] (03CR) 10Majavah: [C:03+2] P:wmcs::novaproxy: Handle HTTP redirects with HAProxy [puppet] - 10https://gerrit.wikimedia.org/r/1307784 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [09:28:30] !log root@cumin1003 START - Cookbook sre.mysql.depool depool db1153: Security updates [09:28:30] !log root@cumin1003 START - Cookbook sre.mysql.parsercache [09:28:36] (03Merged) 10jenkins-bot: sre.mysql.pool: --skip-safety-checks fails for parsercache hosts [cookbooks] - 10https://gerrit.wikimedia.org/r/1307815 (https://phabricator.wikimedia.org/T429878) (owner: 10CWilliams) [09:28:38] !log root@cumin1003 END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) [09:28:38] !log root@cumin1003 END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Security updates [09:28:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:30:30] (03PS3) 10Fabfur: cache_misc: apply traffic classification [puppet] - 10https://gerrit.wikimedia.org/r/1308040 [09:34:39] PROBLEM - MariaDB Replica IO: ms3 on db2252 is CRITICAL: CRITICAL slave_io_state Slave_IO_Running: No, Errno: 2003, Errmsg: error reconnecting to master repl2024@db1153.eqiad.wmnet:3306 - retry-time: 60 maximum-retries: 100000 message: Cant connect to server on db1153.eqiad.wmnet (111 Connection refused) https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [09:34:46] (03CR) 10Blake: "Thanks for the catch, I'm seeing other alerts rewriting instance into hostname, which makes sense, as the Spicerack docs note that "...the" [alerts] - 10https://gerrit.wikimedia.org/r/1307787 (https://phabricator.wikimedia.org/T430290) (owner: 10Blake) [09:37:38] (03PS1) 10Kevin Bazira: ml-services: deploy cope-a-9b isvc image that emits vLLM metrics [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308049 (https://phabricator.wikimedia.org/T431136) [09:41:21] (03PS1) 10Kevin Bazira: ml-services: deploy gpt-oss-safeguard-20b isvc image that emits vLLM metrics [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308050 (https://phabricator.wikimedia.org/T431136) [09:42:41] PROBLEM - MariaDB Replica Lag: ms3 on db2252 is CRITICAL: CRITICAL slave_sql_lag Replication lag: 643.25 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [09:45:12] !log atsuko@cumin1003 START - Cookbook sre.hosts.reimage for host cirrussearch1116.eqiad.wmnet with OS trixie [09:45:22] (03PS1) 10Kevin Bazira: ml-services: deploy embeddings isvc image that emits vLLM metrics [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308051 (https://phabricator.wikimedia.org/T431136) [09:45:32] (03CR) 10Brouberol: [C:03+1] topolvm: customise the imported chart for WMF [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305974 (https://phabricator.wikimedia.org/T429331) (owner: 10Btullis) [09:45:50] !log atsuko@cumin1003 START - Cookbook sre.hosts.reimage for host cirrussearch1117.eqiad.wmnet with OS trixie [09:47:01] (03CR) 10Fabfur: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308040 (owner: 10Fabfur) [09:47:19] (03Abandoned) 10Hnowlan: Revert "restbase: move nrpe check to prom blackbox check" [puppet] - 10https://gerrit.wikimedia.org/r/1308048 (owner: 10Hnowlan) [09:48:22] (03PS1) 10Kevin Bazira: ml-services: deploy qwen36 isvc image that emits vLLM metrics [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308052 (https://phabricator.wikimedia.org/T431136) [09:49:08] !log fceratto@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Security updates [09:49:16] (03PS1) 10Majavah: haproxy: Make logging work with default config [puppet] - 10https://gerrit.wikimedia.org/r/1308053 [09:51:56] (03CR) 10Dpogorzelski: [C:03+1] ml-services: deploy cope-a-9b isvc image that emits vLLM metrics [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308049 (https://phabricator.wikimedia.org/T431136) (owner: 10Kevin Bazira) [09:52:07] (03CR) 10Dpogorzelski: [C:03+1] ml-services: deploy gpt-oss-safeguard-20b isvc image that emits vLLM metrics [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308050 (https://phabricator.wikimedia.org/T431136) (owner: 10Kevin Bazira) [09:52:17] (03CR) 10Dpogorzelski: [C:03+1] ml-services: deploy embeddings isvc image that emits vLLM metrics [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308051 (https://phabricator.wikimedia.org/T431136) (owner: 10Kevin Bazira) [09:52:28] (03CR) 10Dpogorzelski: [C:03+1] ml-services: deploy qwen36 isvc image that emits vLLM metrics [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308052 (https://phabricator.wikimedia.org/T431136) (owner: 10Kevin Bazira) [09:52:55] (03CR) 10Kevin Bazira: [C:03+2] ml-services: deploy cope-a-9b isvc image that emits vLLM metrics [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308049 (https://phabricator.wikimedia.org/T431136) (owner: 10Kevin Bazira) [09:52:56] (03CR) 10Majavah: [V:03+1] "PCC SUCCESS (NOOP 7): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/8939/console" [puppet] - 10https://gerrit.wikimedia.org/r/1308053 (owner: 10Majavah) [09:53:54] 06SRE, 06Data-Engineering, 06Data-Platform-SRE (2026-07-03 - 2026-07-31), 07Sustainability (Incident Followup): The webrequest_sampled_live data pipeline and its query tools have become mission-critical and require re-engineering for resilience - https://phabricator.wikimedia.org/T431112#12093585 (10hnowlan... [09:54:18] (03CR) 10Brouberol: [C:03+1] topolvm: tighten controller RBAC [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305975 (https://phabricator.wikimedia.org/T429331) (owner: 10Btullis) [09:54:42] (03CR) 10Brouberol: [C:03+1] topolvm: scrape controller/node metrics via prometheus.io annotations [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306222 (https://phabricator.wikimedia.org/T429331) (owner: 10Btullis) [09:55:05] (03Merged) 10jenkins-bot: ml-services: deploy cope-a-9b isvc image that emits vLLM metrics [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308049 (https://phabricator.wikimedia.org/T431136) (owner: 10Kevin Bazira) [09:55:25] (03PS2) 10Majavah: haproxy: Make logging work with default config [puppet] - 10https://gerrit.wikimedia.org/r/1308053 [09:55:30] (03CR) 10Brouberol: [C:03+1] topolvm-crds: add the TopoLVM CRD for version 0.38.1 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305976 (https://phabricator.wikimedia.org/T429331) (owner: 10Btullis) [09:57:25] !log atsuko@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage [09:57:36] (03CR) 10Brouberol: admin_ng: define the topolvm CSI releases (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306223 (https://phabricator.wikimedia.org/T429331) (owner: 10Btullis) [09:57:49] (03CR) 10Brouberol: [C:03+1] admin_ng: enable the topolvm CSI driver on dse-k8s [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305978 (https://phabricator.wikimedia.org/T429331) (owner: 10Btullis) [09:58:21] (03CR) 10Brouberol: [C:03+1] topolvm: import the upstream chart version 15.7.1 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305973 (https://phabricator.wikimedia.org/T429331) (owner: 10Btullis) [09:58:26] !log atsuko@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage [09:58:34] !log kevinbazira@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . [09:58:51] (03CR) 10Majavah: [V:03+1] "PCC SUCCESS (NOOP 7 CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/" [puppet] - 10https://gerrit.wikimedia.org/r/1308053 (owner: 10Majavah) [09:59:35] (03CR) 10Btullis: admin_ng: define the topolvm CSI releases (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306223 (https://phabricator.wikimedia.org/T429331) (owner: 10Btullis) [09:59:58] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.5 point update - https://phabricator.wikimedia.org/T427072#12093614 (10MoritzMuehlenhoff) [10:00:04] Deploy window MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260707T1000) [10:03:34] !log atsuko@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage [10:03:43] RECOVERY - MariaDB Replica IO: ms3 on db2252 is OK: OK slave_io_state Slave_IO_Running: Yes https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [10:04:44] (03CR) 10Klausman: [C:03+1] "This is covered by change 1306682, which in turn is currently on hold since we may re-purpose some of those machines for DPE use (rather t" [puppet] - 10https://gerrit.wikimedia.org/r/1308013 (https://phabricator.wikimedia.org/T430654) (owner: 10Muehlenhoff) [10:05:09] (03CR) 10Majavah: [C:03+2] dynamicproxy: Update schema to only support a single backend [puppet] - 10https://gerrit.wikimedia.org/r/1307714 (https://phabricator.wikimedia.org/T429960) (owner: 10Majavah) [10:07:12] !log atsuko@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage [10:07:29] (03PS2) 10Hnowlan: smart: emit BBU data from megaraid hosts as gauge [puppet] - 10https://gerrit.wikimedia.org/r/1307087 (https://phabricator.wikimedia.org/T430149) [10:07:35] FIRING: CirrusSearchSaneitizerFixRateTooHigh: MediaWiki CirrusSearch Saneitizer is fixing an abnormally high number of documents in eqiad - https://wikitech.wikimedia.org/wiki/Search/CirrusStreamingUpdater#San(e)itizing - https://grafana.wikimedia.org/d/JLK3I_siz/elasticsearch-indexing?viewPanel=59&orgId=1&from=now-6M&to=now&var-search_cluster=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchSaneitizerFixRateTooHigh [10:07:43] RECOVERY - MariaDB Replica Lag: ms3 on db2252 is OK: OK slave_sql_lag Replication lag: 0.19 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [10:10:01] (03CR) 10Hnowlan: smart: emit BBU data from megaraid hosts as gauge (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1307087 (https://phabricator.wikimedia.org/T430149) (owner: 10Hnowlan) [10:10:29] (03CR) 10Kevin Bazira: [C:03+2] ml-services: deploy gpt-oss-safeguard-20b isvc image that emits vLLM metrics [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308050 (https://phabricator.wikimedia.org/T431136) (owner: 10Kevin Bazira) [10:11:42] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host thanos-fe2004.codfw.wmnet with OS trixie [10:11:57] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12093666 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host thanos-fe2004.codfw.wmnet with OS trixie [10:12:11] !log mvernon@cumin2003 START - Cookbook sre.hosts.move-vlan for host thanos-fe2004 [10:12:58] !log mvernon@cumin2003 START - Cookbook sre.dns.netbox [10:13:16] (03Merged) 10jenkins-bot: ml-services: deploy gpt-oss-safeguard-20b isvc image that emits vLLM metrics [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308050 (https://phabricator.wikimedia.org/T431136) (owner: 10Kevin Bazira) [10:14:26] !log fceratto@cumin1003 START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet [10:14:26] !log fceratto@cumin1003 END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet [10:14:59] !log fceratto@cumin1003 START - Cookbook sre.hosts.remove-downtime for db1153.eqiad.wmnet [10:14:59] !log fceratto@cumin1003 END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1153.eqiad.wmnet [10:15:01] (03CR) 10Hnowlan: [C:03+2] smart: emit BBU data from megaraid hosts as gauge [puppet] - 10https://gerrit.wikimedia.org/r/1307087 (https://phabricator.wikimedia.org/T430149) (owner: 10Hnowlan) [10:15:36] !log fceratto@cumin1003 START - Cookbook sre.mysql.pool pool db2252: Repooling after reboot [10:15:37] !log fceratto@cumin1003 START - Cookbook sre.mysql.parsercache [10:15:50] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) [10:15:50] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Repooling after reboot [10:15:58] !log kevinbazira@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . [10:19:03] mvernon@cumin2003 reimage (PID 3752577) is awaiting input [10:20:52] !log mvernon@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" [10:20:56] !log mvernon@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003" [10:20:57] !log mvernon@cumin2003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [10:20:57] !log mvernon@cumin2003 START - Cookbook sre.dns.wipe-cache thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [10:20:59] (03PS1) 10Majavah: dynamicproxy: Include singular 'backend' key with API responses [puppet] - 10https://gerrit.wikimedia.org/r/1308060 (https://phabricator.wikimedia.org/T429960) [10:21:00] !log mvernon@cumin2003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [10:21:01] (03PS1) 10Majavah: dynamicproxy: Allow using singular 'backend' for writing data [puppet] - 10https://gerrit.wikimedia.org/r/1308061 (https://phabricator.wikimedia.org/T429960) [10:21:01] !log mvernon@cumin2003 START - Cookbook sre.network.configure-switch-interfaces for host thanos-fe2004 [10:21:04] (03PS1) 10Majavah: openstack: wmcs-webproxy: Use singular 'backend' for API requests [puppet] - 10https://gerrit.wikimedia.org/r/1308062 (https://phabricator.wikimedia.org/T429960) [10:21:06] (03PS1) 10Majavah: openstack: novastats: Update proxyleaks for singular backend object [puppet] - 10https://gerrit.wikimedia.org/r/1308063 (https://phabricator.wikimedia.org/T429960) [10:21:09] (03PS1) 10Majavah: openstack: wmf_sink: Update for singular proxy 'backend' object [puppet] - 10https://gerrit.wikimedia.org/r/1308064 (https://phabricator.wikimedia.org/T429960) [10:22:32] !log atsuko@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1116.eqiad.wmnet with OS trixie [10:22:57] jouncebot: nowandnext [10:22:57] For the next 0 hour(s) and 37 minute(s): MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260707T1000) [10:22:57] In 1 hour(s) and 37 minute(s): Mobileapps/RESTBase/Wikifeeds (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260707T1200) [10:24:05] mvernon@cumin2003 reimage (PID 3752577) is awaiting input [10:25:59] !log mvernon@cumin2003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host thanos-fe2004 [10:26:00] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host thanos-fe2004 [10:26:37] !log cgoubert@deploy1003 Started deploy [restbase/deploy@8a25036]: 1306049: Add isvwiki to RESTBase | https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - T429936 [10:26:40] T429936: Add isvwiki to RESTBase - https://phabricator.wikimedia.org/T429936 [10:26:56] !log atsuko@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1117.eqiad.wmnet with OS trixie [10:27:22] !log cgoubert@deploy1003 Finished deploy [restbase/deploy@8a25036]: 1306049: Add isvwiki to RESTBase | https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - T429936 (duration: 00m 45s) [10:27:28] !log cgoubert@deploy1003 Started deploy [restbase/deploy@2fc37d4]: 1306049: Add isvwiki to RESTBase | https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - T429936 [10:30:36] (03CR) 10Kevin Bazira: [C:03+2] ml-services: deploy embeddings isvc image that emits vLLM metrics [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308051 (https://phabricator.wikimedia.org/T431136) (owner: 10Kevin Bazira) [10:33:08] (03Merged) 10jenkins-bot: ml-services: deploy embeddings isvc image that emits vLLM metrics [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308051 (https://phabricator.wikimedia.org/T431136) (owner: 10Kevin Bazira) [10:33:40] (03PS1) 10Brouberol: Add https://standaarden.overheid.nl/tooi/sparql to WDQS allowlist [puppet] - 10https://gerrit.wikimedia.org/r/1308067 (https://phabricator.wikimedia.org/T407109) [10:33:42] (03PS1) 10Brouberol: Add https://linked.rism.io/api to WDQS allowlist [puppet] - 10https://gerrit.wikimedia.org/r/1308068 (https://phabricator.wikimedia.org/T430264) [10:33:45] (03PS1) 10Brouberol: Add https://lgbtdb.wikibase.cloud/query/sparql to WDQS allowlist [puppet] - 10https://gerrit.wikimedia.org/r/1308069 (https://phabricator.wikimedia.org/T430523) [10:35:47] !log kevinbazira@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . [10:37:52] (03PS3) 10Btullis: topolvm: import the upstream chart version 15.7.1 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305973 (https://phabricator.wikimedia.org/T429331) [10:37:52] (03PS6) 10Btullis: topolvm: customise the imported chart for WMF [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305974 (https://phabricator.wikimedia.org/T429331) [10:37:52] (03PS8) 10Btullis: topolvm: tighten controller RBAC [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305975 (https://phabricator.wikimedia.org/T429331) [10:37:53] (03PS8) 10Btullis: topolvm: scrape controller/node metrics via prometheus.io annotations [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306222 (https://phabricator.wikimedia.org/T429331) [10:37:54] (03PS9) 10Btullis: topolvm-crds: add the TopoLVM CRD for version 0.38.1 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305976 (https://phabricator.wikimedia.org/T429331) [10:37:55] (03PS9) 10Btullis: admin_ng: define the topolvm CSI releases [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306223 (https://phabricator.wikimedia.org/T429331) [10:37:59] (03PS10) 10Btullis: admin_ng: enable the topolvm CSI driver on dse-k8s [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305978 (https://phabricator.wikimedia.org/T429331) [10:39:05] (03PS10) 10Btullis: admin_ng: define the topolvm CSI releases [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306223 (https://phabricator.wikimedia.org/T429331) [10:39:05] (03PS11) 10Btullis: admin_ng: enable the topolvm CSI driver on dse-k8s [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305978 (https://phabricator.wikimedia.org/T429331) [10:39:46] (03CR) 10Btullis: [C:03+2] "Acknowledged" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305973 (https://phabricator.wikimedia.org/T429331) (owner: 10Btullis) [10:40:17] (03CR) 10Clément Goubert: [C:03+1] sre.k8s.renumber-node: Increase timeout to 600s for puppet runs on deploy hosts (031 comment) [cookbooks] - 10https://gerrit.wikimedia.org/r/1307872 (https://phabricator.wikimedia.org/T430226) (owner: 10Jasmine) [10:40:36] (03CR) 10Btullis: [C:03+2] topolvm: customise the imported chart for WMF [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305974 (https://phabricator.wikimedia.org/T429331) (owner: 10Btullis) [10:40:51] (03CR) 10Btullis: [C:03+2] topolvm: tighten controller RBAC [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305975 (https://phabricator.wikimedia.org/T429331) (owner: 10Btullis) [10:41:01] (03CR) 10Btullis: [C:03+2] topolvm: scrape controller/node metrics via prometheus.io annotations [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306222 (https://phabricator.wikimedia.org/T429331) (owner: 10Btullis) [10:41:11] (03CR) 10Btullis: [C:03+2] topolvm-crds: add the TopoLVM CRD for version 0.38.1 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305976 (https://phabricator.wikimedia.org/T429331) (owner: 10Btullis) [10:41:28] (03CR) 10Btullis: admin_ng: define the topolvm CSI releases (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306223 (https://phabricator.wikimedia.org/T429331) (owner: 10Btullis) [10:41:56] (03Merged) 10jenkins-bot: topolvm: import the upstream chart version 15.7.1 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305973 (https://phabricator.wikimedia.org/T429331) (owner: 10Btullis) [10:42:46] (03Merged) 10jenkins-bot: topolvm: customise the imported chart for WMF [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305974 (https://phabricator.wikimedia.org/T429331) (owner: 10Btullis) [10:42:51] (03Merged) 10jenkins-bot: topolvm: tighten controller RBAC [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305975 (https://phabricator.wikimedia.org/T429331) (owner: 10Btullis) [10:42:55] (03Merged) 10jenkins-bot: topolvm: scrape controller/node metrics via prometheus.io annotations [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306222 (https://phabricator.wikimedia.org/T429331) (owner: 10Btullis) [10:42:58] (03CR) 10Kevin Bazira: [C:03+2] ml-services: deploy qwen36 isvc image that emits vLLM metrics [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308052 (https://phabricator.wikimedia.org/T431136) (owner: 10Kevin Bazira) [10:43:15] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage [10:43:39] (03Merged) 10jenkins-bot: topolvm-crds: add the TopoLVM CRD for version 0.38.1 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305976 (https://phabricator.wikimedia.org/T429331) (owner: 10Btullis) [10:44:13] !log cgoubert@deploy1003 Finished deploy [restbase/deploy@2fc37d4]: 1306049: Add isvwiki to RESTBase | https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - T429936 (duration: 16m 44s) [10:44:16] T429936: Add isvwiki to RESTBase - https://phabricator.wikimedia.org/T429936 [10:44:21] (03PS1) 10VadymTS1: Init Wikiquote Minangkabau [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308070 (https://phabricator.wikimedia.org/T429922) [10:45:09] (03Merged) 10jenkins-bot: ml-services: deploy qwen36 isvc image that emits vLLM metrics [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308052 (https://phabricator.wikimedia.org/T431136) (owner: 10Kevin Bazira) [10:46:23] !log atsuko@cumin1003 START - Cookbook sre.hosts.reimage for host cirrussearch1085.eqiad.wmnet with OS trixie [10:46:41] !log atsuko@cumin1003 START - Cookbook sre.hosts.reimage for host cirrussearch1086.eqiad.wmnet with OS trixie [10:47:32] (03PS1) 10Hnowlan: team-sre/hardware: add BBU check [alerts] - 10https://gerrit.wikimedia.org/r/1308071 (https://phabricator.wikimedia.org/T430149) [10:48:16] !log kevinbazira@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . [10:48:33] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage [10:49:49] (03CR) 10Btullis: [C:03+1] Add https://standaarden.overheid.nl/tooi/sparql to WDQS allowlist [puppet] - 10https://gerrit.wikimedia.org/r/1308067 (https://phabricator.wikimedia.org/T407109) (owner: 10Brouberol) [10:50:00] (03CR) 10Btullis: [C:03+1] Add https://linked.rism.io/api to WDQS allowlist [puppet] - 10https://gerrit.wikimedia.org/r/1308068 (https://phabricator.wikimedia.org/T430264) (owner: 10Brouberol) [10:50:11] (03CR) 10Btullis: [C:03+1] Add https://lgbtdb.wikibase.cloud/query/sparql to WDQS allowlist [puppet] - 10https://gerrit.wikimedia.org/r/1308069 (https://phabricator.wikimedia.org/T430523) (owner: 10Brouberol) [11:00:31] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, July 08 UTC morning backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308070 (https://phabricator.wikimedia.org/T429922) (owner: 10VadymTS1) [11:00:55] (03CR) 10Muehlenhoff: "Ok, then let me go ahead and merge this one (since there is currently an alert about broken aliases)" [puppet] - 10https://gerrit.wikimedia.org/r/1308013 (https://phabricator.wikimedia.org/T430654) (owner: 10Muehlenhoff) [11:00:57] (03CR) 10Muehlenhoff: [C:03+2] Remove ml-cache Cumin alias [puppet] - 10https://gerrit.wikimedia.org/r/1308013 (https://phabricator.wikimedia.org/T430654) (owner: 10Muehlenhoff) [11:02:19] !log atsuko@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage [11:03:17] !log atsuko@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage [11:03:53] (03CR) 10Clément Goubert: "This affects a majority of the old RESTBase endpoints, namely:" [puppet] - 10https://gerrit.wikimedia.org/r/1306905 (https://phabricator.wikimedia.org/T428909) (owner: 10Clément Goubert) [11:04:01] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2004.codfw.wmnet with OS trixie [11:10:18] !log atsuko@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage [11:11:18] (03CR) 10Clément Goubert: "Disregard the list of old RESTBase endpoints, actually, those are under `/api/rest_v1/`. The point stands that it's confusing as heck :D" [puppet] - 10https://gerrit.wikimedia.org/r/1306905 (https://phabricator.wikimedia.org/T428909) (owner: 10Clément Goubert) [11:12:48] (03CR) 10Clément Goubert: [C:03+2] trafficserver::backend: Remove X-W-D for /w/rest.php [puppet] - 10https://gerrit.wikimedia.org/r/1306905 (https://phabricator.wikimedia.org/T428909) (owner: 10Clément Goubert) [11:14:32] !log atsuko@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage [11:24:20] (03PS3) 10Clément Goubert: mwparsoid: Mount /srv/mediawiki from mw-experimental [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308029 (https://phabricator.wikimedia.org/T431387) (owner: 10Jgiannelos) [11:24:20] (03CR) 10Clément Goubert: [C:03+1] "LGTM, we'll revisit making it default to `false` based on frequency of local code editing" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308029 (https://phabricator.wikimedia.org/T431387) (owner: 10Jgiannelos) [11:31:07] (03PS1) 10Arnaudb: gerrit: limit the volume of returnable pages [puppet] - 10https://gerrit.wikimedia.org/r/1308081 (https://phabricator.wikimedia.org/T431398) [11:31:07] (03CR) 10Arnaudb: "see T431398 for more details" [puppet] - 10https://gerrit.wikimedia.org/r/1308081 (https://phabricator.wikimedia.org/T431398) (owner: 10Arnaudb) [11:32:49] !log atsuko@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1085.eqiad.wmnet with OS trixie [11:33:37] (03PS1) 10Bartosz Wójtowicz: ml-services: Remove experimental qwen3-14b deployment. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308085 (https://phabricator.wikimedia.org/T431268) [11:35:09] !log mvernon@cumin2003 conftool action : set/pooled=inactive; selector: name=thanos-fe2004.codfw.wmnet [11:35:42] (03PS1) 10Klausman: role::ml_k8s::master: Fold IPIP update back into main file [puppet] - 10https://gerrit.wikimedia.org/r/1308080 (https://phabricator.wikimedia.org/T420438) [11:35:44] !log atsuko@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1086.eqiad.wmnet with OS trixie [11:36:32] !log mvernon@cumin2003 conftool action : set/pooled=yes; selector: name=thanos-fe2004.codfw.wmnet [11:37:55] (03CR) 10Btullis: [C:03+1] "Looks good to me." [puppet] - 10https://gerrit.wikimedia.org/r/1308080 (https://phabricator.wikimedia.org/T420438) (owner: 10Klausman) [11:38:26] (03CR) 10Klausman: [C:03+2] role::ml_k8s::master: Fold IPIP update back into main file [puppet] - 10https://gerrit.wikimedia.org/r/1308080 (https://phabricator.wikimedia.org/T420438) (owner: 10Klausman) [11:39:16] !log cmooney@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin [11:40:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 0% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [11:40:57] FIRING: ProbeDown: Service mw-web:4450 has failed probes (http_mw-web_ip4) #page - https://wikitech.wikimedia.org/wiki/Runbook#mw-web:4450 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [11:41:52] !log cmooney@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin [11:42:02] !log cr3-eqsin, begin traffic drain to reset PIC and upgrade JunOS T429386 [11:42:04] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [11:42:05] T429386: cr2-esams rpd failure after enabling bgp 'graceful-shutdown' (June 2026) - https://phabricator.wikimedia.org/T429386 [11:42:14] (03CR) 10Ladsgroup: [C:03+1] tables-catalog.yaml: Add oathauth_user_handles [puppet] - 10https://gerrit.wikimedia.org/r/1308006 (https://phabricator.wikimedia.org/T431124) (owner: 10Marostegui) [11:42:43] !ack [11:42:44] 8151 (ACKED) ProbeDown sre (10.2.2.75 ip4 mw-web:4450 probes/service http_mw-web_ip4 eqiad) [11:43:15] FIRING: MediaWikiLatencyExceeded: p75 latency high: eqiad mw-web releases routed via main (k8s) 2.5s - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook#Average_latency_exceeded - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=55&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiLatencyExceeded [11:44:58] FIRING: ProbeDown: Service text-https:443 has failed probes (http_text-https_ip6) - https://wikitech.wikimedia.org/wiki/Runbook#text-https:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [11:45:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 0.2078% idle #page - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [11:45:57] RESOLVED: ProbeDown: Service mw-web:4450 has failed probes (http_mw-web_ip4) #page - https://wikitech.wikimedia.org/wiki/Runbook#mw-web:4450 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [11:46:13] FIRING: HaproxyUnavailable: HAProxy (cache_text) has reduced HTTP availability #page - https://wikitech.wikimedia.org/wiki/HAProxy#HAProxy_for_edge_caching - https://grafana.wikimedia.org/d/000000479/frontend-traffic?viewPanel=13 - https://alerts.wikimedia.org/?q=alertname%3DHaproxyUnavailable [11:46:22] !incidents [11:46:22] 8152 (UNACKED) PHPFPMTooBusy sre (mw-web main eqiad) [11:46:23] 8153 (UNACKED) HaproxyUnavailable cache_text global sre (thanos-rule@main) [11:46:23] 8151 (RESOLVED) ProbeDown sre (10.2.2.75 ip4 mw-web:4450 probes/service http_mw-web_ip4 eqiad) [11:46:23] 8149 (RESOLVED) TransitPeeringTransportOutSaturation network sre (cr1-codfw:9804 Transport: cr4-ulsfo:xe-0/1/1 (Lumen, 442550294) {#12252_12295-1} xe-1/1/1:0 gnmi codfw) [11:46:23] 8148 (RESOLVED) TransitPeeringTransportOutSaturation network sre (cr1-codfw:9804 Transport: cr4-ulsfo:xe-0/1/1 (Lumen, 442550294) {#12252_12295-1} xe-1/1/1:0 gnmi codfw) [11:46:30] !ack [11:46:31] 8152 (ACKED) PHPFPMTooBusy sre (mw-web main eqiad) [11:46:31] 8153 (ACKED) HaproxyUnavailable cache_text global sre (thanos-rule@main) [11:48:30] (03PS1) 10Btullis: datahub: promote production to DataHub v1.6.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308086 (https://phabricator.wikimedia.org/T402408) [11:49:27] !log blake@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on wikikube-worker1160.eqiad.wmnet with reason: Verifying matchers for silence [11:49:57] RESOLVED: ProbeDown: Service text-https:443 has failed probes (http_text-https_ip6) - https://wikitech.wikimedia.org/wiki/Runbook#text-https:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [11:50:21] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host thanos-fe2005.codfw.wmnet with OS trixie [11:51:13] RESOLVED: HaproxyUnavailable: HAProxy (cache_text) has reduced HTTP availability #page - https://wikitech.wikimedia.org/wiki/HAProxy#HAProxy_for_edge_caching - https://grafana.wikimedia.org/d/000000479/frontend-traffic?viewPanel=13 - https://alerts.wikimedia.org/?q=alertname%3DHaproxyUnavailable [11:53:08] (03CR) 10Brouberol: [C:03+1] admin_ng: define the topolvm CSI releases [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306223 (https://phabricator.wikimedia.org/T429331) (owner: 10Btullis) [11:53:59] (03CR) 10Btullis: [C:03+2] admin_ng: define the topolvm CSI releases [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306223 (https://phabricator.wikimedia.org/T429331) (owner: 10Btullis) [11:54:23] (03CR) 10Brouberol: [C:03+2] Add https://standaarden.overheid.nl/tooi/sparql to WDQS allowlist [puppet] - 10https://gerrit.wikimedia.org/r/1308067 (https://phabricator.wikimedia.org/T407109) (owner: 10Brouberol) [11:54:30] !incidents [11:54:30] 8152 (ACKED) PHPFPMTooBusy sre (mw-web main eqiad) [11:54:31] 8153 (RESOLVED) HaproxyUnavailable cache_text global sre (thanos-rule@main) [11:54:31] 8151 (RESOLVED) ProbeDown sre (10.2.2.75 ip4 mw-web:4450 probes/service http_mw-web_ip4 eqiad) [11:54:31] 8149 (RESOLVED) TransitPeeringTransportOutSaturation network sre (cr1-codfw:9804 Transport: cr4-ulsfo:xe-0/1/1 (Lumen, 442550294) {#12252_12295-1} xe-1/1/1:0 gnmi codfw) [11:54:31] 8148 (RESOLVED) TransitPeeringTransportOutSaturation network sre (cr1-codfw:9804 Transport: cr4-ulsfo:xe-0/1/1 (Lumen, 442550294) {#12252_12295-1} xe-1/1/1:0 gnmi codfw) [11:55:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 0.831% idle #page - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [11:55:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.89% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [11:55:59] (03CR) 10Brouberol: [C:03+2] Add https://linked.rism.io/api to WDQS allowlist [puppet] - 10https://gerrit.wikimedia.org/r/1308068 (https://phabricator.wikimedia.org/T430264) (owner: 10Brouberol) [11:56:03] !log klausman@cumin2002 START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-eqiad [11:56:05] (03CR) 10Brouberol: [C:03+2] Add https://lgbtdb.wikibase.cloud/query/sparql to WDQS allowlist [puppet] - 10https://gerrit.wikimedia.org/r/1308069 (https://phabricator.wikimedia.org/T430523) (owner: 10Brouberol) [11:56:07] !log klausman@cumin2002 START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1001.eqiad.wmnet [11:56:08] !log klausman@cumin2002 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1001.eqiad.wmnet [11:56:28] (03PS2) 10Brouberol: Add https://linked.rism.io/api to WDQS allowlist [puppet] - 10https://gerrit.wikimedia.org/r/1308068 (https://phabricator.wikimedia.org/T430264) [11:56:28] (03PS2) 10Brouberol: Add https://lgbtdb.wikibase.cloud/query/sparql to WDQS allowlist [puppet] - 10https://gerrit.wikimedia.org/r/1308069 (https://phabricator.wikimedia.org/T430523) [11:57:29] (03PS1) 10Btullis: opensearch: enable the datahub namespace for its OpenSearch cluster [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308088 (https://phabricator.wikimedia.org/T402408) [11:57:56] (03CR) 10Brouberol: [C:03+2] Add https://linked.rism.io/api to WDQS allowlist [puppet] - 10https://gerrit.wikimedia.org/r/1308068 (https://phabricator.wikimedia.org/T430264) (owner: 10Brouberol) [11:58:01] (03CR) 10Brouberol: [C:03+2] Add https://lgbtdb.wikibase.cloud/query/sparql to WDQS allowlist [puppet] - 10https://gerrit.wikimedia.org/r/1308069 (https://phabricator.wikimedia.org/T430523) (owner: 10Brouberol) [11:58:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.76% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [11:58:15] RESOLVED: MediaWikiLatencyExceeded: p75 latency high: eqiad mw-web releases routed via main (k8s) 1.147s - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook#Average_latency_exceeded - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=55&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiLatencyExceeded [11:59:34] !log klausman@cumin2002 START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1001.eqiad.wmnet [11:59:36] !log klausman@cumin2002 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1001.eqiad.wmnet [11:59:42] !log klausman@cumin2002 START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1002.eqiad.wmnet [11:59:44] !log klausman@cumin2002 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1002.eqiad.wmnet [12:00:05] Deploy window Mobileapps/RESTBase/Wikifeeds (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260707T1200) [12:00:06] (03CR) 10Brouberol: [C:03+1] "Nice! 3 less hosts we have to worry about" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308088 (https://phabricator.wikimedia.org/T402408) (owner: 10Btullis) [12:01:54] (03CR) 10Brouberol: [C:03+1] datahub: promote production to DataHub v1.6.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308086 (https://phabricator.wikimedia.org/T402408) (owner: 10Btullis) [12:03:09] !log brouberol@cumin1003 START - Cookbook sre.wdqs.restart [12:03:09] !log brouberol@cumin1003 END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) [12:03:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.76% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [12:03:39] !log brouberol@cumin1003 START - Cookbook sre.wdqs.restart [12:03:39] !log brouberol@cumin1003 END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) [12:03:51] (03Merged) 10jenkins-bot: admin_ng: define the topolvm CSI releases [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306223 (https://phabricator.wikimedia.org/T429331) (owner: 10Btullis) [12:04:43] !log brouberol@cumin1003 START - Cookbook sre.wdqs.restart [12:04:44] !log brouberol@cumin1003 END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) [12:04:53] !log klausman@cumin2002 START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1002.eqiad.wmnet [12:04:55] !log klausman@cumin2002 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1002.eqiad.wmnet [12:04:55] !log klausman@cumin2002 END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-eqiad [12:05:25] !log brouberol@cumin1003 START - Cookbook sre.wdqs.restart [12:05:26] !log brouberol@cumin1003 END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) [12:06:04] !log brouberol@cumin1003 START - Cookbook sre.wdqs.restart [12:06:04] !log brouberol@cumin1003 END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) [12:06:10] !log brouberol@cumin1003 START - Cookbook sre.wdqs.restart [12:07:29] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage [12:07:37] !log brouberol@cumin1003 END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99) [12:08:19] !log brouberol@cumin1003 START - Cookbook sre.wdqs.restart [12:11:16] (03CR) 10Kamila Součková: [C:03+2] mediawiki/php: don't pin libpcre for bookworm (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1307128 (https://phabricator.wikimedia.org/T423714) (owner: 10Kamila Součková) [12:14:17] !log brouberol@cumin1003 END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) [12:14:29] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage [12:14:55] !log brouberol@cumin1003 START - Cookbook sre.wdqs.restart [12:15:17] (03PS1) 10Klausman: hiera/services: Add IPIP LVS config for ML k8s [puppet] - 10https://gerrit.wikimedia.org/r/1308094 [12:15:37] (03CR) 10Marostegui: [C:03+2] tables-catalog.yaml: Add oathauth_user_handles [puppet] - 10https://gerrit.wikimedia.org/r/1308006 (https://phabricator.wikimedia.org/T431124) (owner: 10Marostegui) [12:16:18] (03PS2) 10Klausman: hiera/services: Add IPIP LVS config for ML k8s [puppet] - 10https://gerrit.wikimedia.org/r/1308094 [12:19:38] !log brouberol@cumin1003 END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) [12:20:05] !log brouberol@cumin1003 START - Cookbook sre.wdqs.restart [12:21:07] !log marostegui@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: restarting for replication filter [12:21:15] !log Restart mariadb@s7 on db1155 to pick up new filters - T431124 [12:21:16] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [12:22:09] !log blake@cumin1003 START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1036.eqiad.wmnet [12:22:13] !log blake@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1036.eqiad.wmnet [12:22:20] !log atsuko@cumin1003 START - Cookbook sre.hosts.reimage for host cirrussearch1082.eqiad.wmnet with OS trixie [12:22:49] !log blake@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1036.eqiad.wmnet [12:23:02] !log blake@cumin1003 START - Cookbook sre.hosts.reimage for host wikikube-worker1036.eqiad.wmnet with OS trixie [12:23:30] !log blake@cumin1003 START - Cookbook sre.hosts.move-vlan for host wikikube-worker1036 [12:23:54] (03CR) 10Btullis: [C:03+2] opensearch: enable the datahub namespace for its OpenSearch cluster [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308088 (https://phabricator.wikimedia.org/T402408) (owner: 10Btullis) [12:23:56] !log blake@cumin1003 START - Cookbook sre.dns.netbox [12:26:07] !log brouberol@cumin1003 END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) [12:26:24] (03PS1) 10Muehlenhoff: Depool puppetserver2001/1002 for reboots [dns] - 10https://gerrit.wikimedia.org/r/1308098 [12:26:31] !log brouberol@cumin1003 START - Cookbook sre.wdqs.restart [12:28:18] (03CR) 10Muehlenhoff: [C:03+2] Depool puppetserver2001/1002 for reboots [dns] - 10https://gerrit.wikimedia.org/r/1308098 (owner: 10Muehlenhoff) [12:28:26] !log jmm@dns1004 START - running authdns-update [12:29:50] !log blake@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" [12:29:50] (03CR) 10Brouberol: [C:03+1] Revert "opensearch-ttmserver: increase memory to 1x index" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1302089 (owner: 10Atsuko) [12:29:54] !log blake@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003" [12:29:54] !log blake@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [12:29:55] !log blake@cumin1003 START - Cookbook sre.dns.wipe-cache wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [12:29:58] !log blake@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [12:29:59] !log blake@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1036 [12:30:05] !log jmm@dns1004 END - running authdns-update [12:30:25] (03PS2) 10Atsuko: Revert "opensearch-ttmserver: increase memory to 1x index" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1302089 [12:32:12] !log atsuko@cumin1003 START - Cookbook sre.hosts.reimage for host cirrussearch1087.eqiad.wmnet with OS trixie [12:32:28] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2005.codfw.wmnet with OS trixie [12:32:32] !log blake@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1036 [12:32:32] !log blake@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1036 [12:33:11] !log filippo@cumin1003 START - Cookbook sre.dns.netbox [12:33:30] (03Merged) 10jenkins-bot: opensearch: enable the datahub namespace for its OpenSearch cluster [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308088 (https://phabricator.wikimedia.org/T402408) (owner: 10Btullis) [12:35:59] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. [12:36:40] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. [12:36:51] (03CR) 10Btullis: [C:03+2] datahub: promote production to DataHub v1.6.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308086 (https://phabricator.wikimedia.org/T402408) (owner: 10Btullis) [12:38:08] !log atsuko@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage [12:39:31] (03Merged) 10jenkins-bot: datahub: promote production to DataHub v1.6.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308086 (https://phabricator.wikimedia.org/T402408) (owner: 10Btullis) [12:39:37] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host thanos-fe2006.codfw.wmnet with OS trixie [12:39:49] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12094338 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host thanos-fe2006.codfw.wmnet with OS trixie [12:39:50] !log filippo@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" [12:39:54] !log filippo@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003" [12:39:55] !log filippo@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [12:41:55] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply [12:41:59] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply} [12:42:56] !log jmm@cumin2003 START - Cookbook sre.puppet.disable-merges [12:42:57] !log jmm@cumin2003 END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) [12:43:35] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet [12:44:35] !log atsuko@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage [12:47:29] (03CR) 10Elukey: [C:03+2] profile::benthos: add x-provenance handling for webrequest [puppet] - 10https://gerrit.wikimedia.org/r/1307169 (https://phabricator.wikimedia.org/T427068) (owner: 10Elukey) [12:48:25] !log atsuko@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage [12:48:43] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply [12:49:21] !log blake@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage [12:49:26] (03PS1) 10Kamila Součková: Revert^2 "hiera: add deploy2003 to deployment servers" [puppet] - 10https://gerrit.wikimedia.org/r/1308101 [12:51:31] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet [12:52:11] !log brouberol@cumin1003 END (PASS) - Cookbook sre.wdqs.restart (exit_code=0) [12:52:14] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet [12:52:24] !log atsuko@cumin1003 START - Cookbook sre.hosts.reimage for host cirrussearch1088.eqiad.wmnet with OS trixie [12:53:24] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply} [12:53:36] !log atsuko@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage [12:55:21] 06SRE, 10SRE-Access-Requests, 06Data-Engineering: Personal Hive Database Username - https://phabricator.wikimedia.org/T431412#12094361 (10Urbanecm_WMF) AFAIK, this should be self served – just run `CREATE DATABASE catherinekelsey` to work with it. See https://wikitech.wikimedia.org/wiki/Data_Platform/Systems... [12:55:36] !log jmm@cumin2003 START - Cookbook sre.puppet.disable-merges [12:55:37] !log jmm@cumin2003 END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) [12:56:24] !log cmooney@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin [12:56:33] 06SRE, 06Infrastructure-Foundations, 10netops: cr2-esams rpd failure after enabling bgp 'graceful-shutdown' (June 2026) - https://phabricator.wikimedia.org/T429386#12094364 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=99c0df32-80dc-453a-920f-5d3957387031) set by cmooney@cumin1003 for 1... [12:57:00] (03PS1) 10Muehlenhoff: Revert "Depool puppetserver2001/1002 for reboots" [dns] - 10https://gerrit.wikimedia.org/r/1308103 [12:57:04] !log cmooney@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin [12:57:12] 06SRE, 06Infrastructure-Foundations, 10netops: cr2-esams rpd failure after enabling bgp 'graceful-shutdown' (June 2026) - https://phabricator.wikimedia.org/T429386#12094370 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=8f5cefde-0060-47c2-8575-70bfa00d21f7) set by cmooney@cumin1003 for 1... [12:57:34] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage [12:57:55] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet [12:58:06] !log blake@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage [12:58:25] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.5 point update - https://phabricator.wikimedia.org/T427072#12094379 (10MoritzMuehlenhoff) [12:58:32] !log load updated JunOS on cr3-eqsin T429386 [12:58:34] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [12:58:35] T429386: cr2-esams rpd failure after enabling bgp 'graceful-shutdown' (June 2026) - https://phabricator.wikimedia.org/T429386 [12:59:25] FIRING: SystemdUnitFailed: requestctl-credential-refresh.service on puppetserver2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:00:05] urbanecm and TheresNoTime: #bothumor My software never has bugs. It just develops random features. Rise for UTC afternoon backport window. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260707T1300). [13:00:05] No Gerrit patches in the queue for this window AFAICS. [13:02:17] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage [13:03:25] RESOLVED: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:03:40] (been a quiet run of deployment windows 👀) [13:05:23] (03CR) 10Muehlenhoff: [C:03+2] Revert "Depool puppetserver2001/1002 for reboots" [dns] - 10https://gerrit.wikimedia.org/r/1308103 (owner: 10Muehlenhoff) [13:05:39] !log jmm@dns1004 START - running authdns-update [13:06:11] !log elukey@cumin1003 START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm [13:06:21] (03CR) 10JMeybohm: [C:03+1] admin_ng: Remove obsolete coredns 1.8.7-2 tag, unset everywhere [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305544 (owner: 10RLazarus) [13:07:21] !log jmm@dns1004 END - running authdns-update [13:07:40] !log atsuko@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1082.eqiad.wmnet with OS trixie [13:09:10] !log atsuko@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage [13:10:25] FIRING: CoreRouterInterfaceDown: Core router interface down - cr4-ulsfo:et-0/0/0 (Transport: Arelion (IC-398715) {#temp_IC-398715}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr4-ulsfo:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [13:11:30] (03CR) 10Muehlenhoff: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1308101 (owner: 10Kamila Součková) [13:12:03] !log jmm@dns1004 START - running authdns-update [13:12:21] (03CR) 10Elukey: [C:03+1] docker_registry: migrate nrpe checks to alertmanager [puppet] - 10https://gerrit.wikimedia.org/r/1306351 (https://phabricator.wikimedia.org/T384321) (owner: 10Hnowlan) [13:13:41] !log jmm@dns1004 END - running authdns-update [13:13:56] if nobody is deploying, I'll borrow the window to test the new deploy server [13:14:13] !log atsuko@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage [13:14:45] !log reboot cr3-eqsin to install new JunOS and set PIC 0/0/0 to 100G [13:14:46] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [13:14:55] !log atsuko@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1087.eqiad.wmnet with OS trixie [13:15:34] (03CR) 10Kamila Součková: [C:03+2] Revert^2 "hiera: add deploy2003 to deployment servers" [puppet] - 10https://gerrit.wikimedia.org/r/1308101 (owner: 10Kamila Součková) [13:15:36] !log Istio is being upgraded from 1.24.2 to 1.29.4 on wikikube staging eqiad and codfw - T427401 [13:15:38] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [13:15:39] T427401: Update istio to 1.29 - https://phabricator.wikimedia.org/T427401 [13:16:46] (03PS1) 10Urbanecm: fix(MentorListCleaner): Do not access property before inicialization [extensions/GrowthExperiments] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1308105 (https://phabricator.wikimedia.org/T430689) [13:16:48] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply [13:17:14] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply} [13:17:59] !log elukey@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage [13:18:39] !log blake@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1036.eqiad.wmnet with OS trixie [13:20:35] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2006.codfw.wmnet with OS trixie [13:20:39] FIRING: [2x] CoreBGPDown: Core BGP session down between cr3-ulsfo and cr3-eqsin (103.102.166.131) - group Confed_eqsin - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [13:20:53] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12094435 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host thanos-fe2006.codfw.wmnet with OS trixie compl... [13:21:14] (03PS1) 10Urbanecm: fix(MentorChangeLogFormatter): Remove unused XSS suppression [extensions/GrowthExperiments] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1308106 (https://phabricator.wikimedia.org/T430693) [13:21:25] (03PS2) 10Urbanecm: fix(MentorListCleaner): Do not access property before inicialization [extensions/GrowthExperiments] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1308105 (https://phabricator.wikimedia.org/T430689) [13:21:39] blake@cumin1003 renumber-node (PID 2192371) is awaiting input [13:23:07] 06SRE, 06Infrastructure-Foundations, 10netops: cr2-esams rpd failure after enabling bgp 'graceful-shutdown' (June 2026) - https://phabricator.wikimedia.org/T429386#12094443 (10cmooney) Upgraded cr3-eqsin also as PIC needed bouncing. And created some instructions to do the traffic drain without 'graceful-shu... [13:23:09] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage [13:23:52] (03PS2) 10Elukey: docker_registry: add migration block for ML images in nginx [puppet] - 10https://gerrit.wikimedia.org/r/1304596 (https://phabricator.wikimedia.org/T428022) [13:24:54] (03CR) 10Elukey: docker_registry: add migration block for ML images in nginx (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1304596 (https://phabricator.wikimedia.org/T428022) (owner: 10Elukey) [13:28:33] (03CR) 10Ssingh: [C:03+1] "Assuming VTC is happy, which you confirmed it is, +1" [puppet] - 10https://gerrit.wikimedia.org/r/1308040 (owner: 10Fabfur) [13:29:48] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host thanos-fe2007.codfw.wmnet with OS trixie [13:30:01] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12094469 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host thanos-fe2007.codfw.wmnet with OS trixie [13:30:39] RESOLVED: [2x] CoreBGPDown: Core BGP session down between cr3-ulsfo and cr3-eqsin (103.102.166.131) - group Confed_eqsin - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [13:31:40] !log depooling dns7002 attempting to reimage to trixie [13:32:08] !log cdobbins@cumin1003 conftool action : set/pooled=no; selector: name=dns7002.* [13:32:14] jouncebot: nowandnext [13:32:14] For the next 0 hour(s) and 27 minute(s): UTC afternoon backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260707T1300) [13:32:15] In 0 hour(s) and 27 minute(s): Test Kitchen UI Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260707T1400) [13:32:24] let me steal the window [13:32:28] (03CR) 10Urbanecm: [C:03+2] fix(MentorChangeLogFormatter): Remove unused XSS suppression [extensions/GrowthExperiments] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1308106 (https://phabricator.wikimedia.org/T430693) (owner: 10Urbanecm) [13:32:29] (03CR) 10Urbanecm: [C:03+2] fix(MentorListCleaner): Do not access property before inicialization [extensions/GrowthExperiments] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1308105 (https://phabricator.wikimedia.org/T430689) (owner: 10Urbanecm) [13:32:48] !log cdobbins@cumin1003 START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie [13:32:48] !log atsuko@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1088.eqiad.wmnet with OS trixie [13:33:23] !log reset cr3-eqsin configuration so traffic uses it again after upgrade [13:33:24] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [13:33:44] !log atsuko@cumin1003 START - Cookbook sre.hosts.reimage for host cirrussearch1120.eqiad.wmnet with OS trixie [13:34:13] (03CR) 10TrainBranchBot: [C:03+2] "Approved by urbanecm@deploy1003 using scap backport" [extensions/GrowthExperiments] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1308106 (https://phabricator.wikimedia.org/T430693) (owner: 10Urbanecm) [13:34:14] (03CR) 10TrainBranchBot: [C:03+2] "Approved by urbanecm@deploy1003 using scap backport" [extensions/GrowthExperiments] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1308105 (https://phabricator.wikimedia.org/T430689) (owner: 10Urbanecm) [13:34:40] (03Merged) 10jenkins-bot: fix(MentorChangeLogFormatter): Remove unused XSS suppression [extensions/GrowthExperiments] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1308106 (https://phabricator.wikimedia.org/T430693) (owner: 10Urbanecm) [13:34:43] (03Merged) 10jenkins-bot: fix(MentorListCleaner): Do not access property before inicialization [extensions/GrowthExperiments] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1308105 (https://phabricator.wikimedia.org/T430689) (owner: 10Urbanecm) [13:34:47] (03CR) 10Jgiannelos: [C:03+2] mwparsoid: Mount /srv/mediawiki from mw-experimental [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308029 (https://phabricator.wikimedia.org/T431387) (owner: 10Jgiannelos) [13:35:22] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.5 point update - https://phabricator.wikimedia.org/T427072#12094506 (10MoritzMuehlenhoff) [13:35:32] !log urbanecm@deploy1003 Started scap sync-world: Backport for [[gerrit:1308106|fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105|fix(MentorListCleaner): Do not access property before inicialization (T430689)]] [13:35:36] T430693: UnusedPluginSuppression in GrowthExperiments - https://phabricator.wikimedia.org/T430693 [13:35:37] T430689: Error: Typed property GrowthExperiments\Mentorship\Cleaner\Actions\ActionFactory::$instanceCache must not be accessed before initialization - https://phabricator.wikimedia.org/T430689 [13:36:17] PROBLEM - BFD status on asw1-b4-magru.mgmt is CRITICAL: Down: 2 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [13:37:02] (03Merged) 10jenkins-bot: mwparsoid: Mount /srv/mediawiki from mw-experimental [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308029 (https://phabricator.wikimedia.org/T431387) (owner: 10Jgiannelos) [13:37:40] FIRING: [2x] BFDdown: BFD session down between asw1-b4-magru and 195.200.68.37 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=asw1-b4-magru:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [13:38:09] !log installing e2fsprogs updates from Trixie point release [13:38:10] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [13:39:26] !log jgiannelos@deploy1003 helmfile [codfw] START helmfile.d/services/mw-parsoid: apply [13:40:15] !log jgiannelos@deploy1003 helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply [13:40:25] (03CR) 10Clément Goubert: [C:03+2] redioscope: Add survey for media ratelimit [deployment-charts] - 10https://gerrit.wikimedia.org/r/1295456 (https://phabricator.wikimedia.org/T424051) (owner: 10Clément Goubert) [13:42:31] (03Merged) 10jenkins-bot: redioscope: Add survey for media ratelimit [deployment-charts] - 10https://gerrit.wikimedia.org/r/1295456 (https://phabricator.wikimedia.org/T424051) (owner: 10Clément Goubert) [13:42:45] (03CR) 10Fabfur: ip_reputation_vendors: add basic safeguard to spur downloader (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1307752 (owner: 10Fabfur) [13:44:22] FIRING: [2x] JobUnavailable: Reduced availability for job haproxy in ops@magru - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [13:44:29] !log cgoubert@deploy1003 helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply [13:44:32] !log jgiannelos@deploy1003 helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply [13:45:17] !log jgiannelos@deploy1003 helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply [13:46:24] !log jgiannelos@deploy1003 helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply [13:46:28] !log jgiannelos@deploy1003 helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply [13:46:38] !log atsuko@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage [13:46:51] !log cgoubert@deploy1003 helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply [13:46:53] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage [13:47:07] !log cgoubert@deploy1003 helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply [13:47:17] !log cgoubert@deploy1003 helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply [13:47:28] PROBLEM - Improperly owned -0:0- files in /srv/mediawiki-staging on deploy2003 is CRITICAL: Improperly owned (0:0) files in /srv/mediawiki-staging https://wikitech.wikimedia.org/wiki/Monitoring/bad_directory_owner [13:47:56] (03CR) 10VadymTS1: "@cgoubert@wikimedia.org Can you deploy this change before deployment window, because this is wiki create patch, it’s need fast deploy" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308070 (https://phabricator.wikimedia.org/T429922) (owner: 10VadymTS1) [13:49:42] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.5 point update - https://phabricator.wikimedia.org/T427072#12094559 (10MoritzMuehlenhoff) [13:50:36] !log installing Linux 5.10.259 on Bullseye hosts [13:50:37] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [13:53:45] !log atsuko@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage [13:57:15] PROBLEM - Host titan1002 is DOWN: PING CRITICAL - Packet loss = 66%, RTA = 4547.74 ms [13:57:17] !log cdobbins@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage [13:57:31] RECOVERY - Improperly owned -0:0- files in /srv/mediawiki-staging on deploy2003 is OK: Files ownership is ok. https://wikitech.wikimedia.org/wiki/Monitoring/bad_directory_owner [13:57:58] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage [13:58:07] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm [13:58:25] !log urbanecm@deploy1003 urbanecm: Backport for [[gerrit:1308106|fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105|fix(MentorListCleaner): Do not access property before inicialization (T430689)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [13:58:30] T430693: UnusedPluginSuppression in GrowthExperiments - https://phabricator.wikimedia.org/T430693 [13:58:30] T430689: Error: Typed property GrowthExperiments\Mentorship\Cleaner\Actions\ActionFactory::$instanceCache must not be accessed before initialization - https://phabricator.wikimedia.org/T430689 [13:58:55] !log urbanecm@deploy1003 urbanecm: Continuing with deployment [13:59:03] FIRING: ProbeDown: Service titan1002:443 has failed probes (http_thanos_wikimedia_org_ip6) - https://wikitech.wikimedia.org/wiki/Runbook#titan1002:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [13:59:25] RESOLVED: SystemdUnitFailed: requestctl-credential-refresh.service on puppetserver2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:59:57] RECOVERY - Host titan1002 is UP: PING OK - Packet loss = 0%, RTA = 0.31 ms [14:00:05] Deploy window Test Kitchen UI Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260707T1400) [14:00:09] (03CR) 10JMeybohm: [C:03+1] coredns: Add an internal_only value [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305546 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [14:00:14] (03CR) 10JMeybohm: [C:03+1] coredns: Parameterize `name` and `k8s-app` [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305545 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [14:00:42] !log atsuko@cumin1003 START - Cookbook sre.hosts.reimage for host cirrussearch1083.eqiad.wmnet with OS trixie [14:00:56] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage [14:01:50] 10ops-eqiad, 06cloud-services-team, 10Cloud-VPS, 06DC-Ops: New thermal paste for cloudvirt1071.eqiad.wmnet - https://phabricator.wikimedia.org/T431429 (10Andrew) 03NEW [14:01:57] 10ops-eqiad, 06cloud-services-team, 10Cloud-VPS, 06DC-Ops: New thermal paste for cloudvirt1071.eqiad.wmnet - https://phabricator.wikimedia.org/T431429#12094665 (10Andrew) Oh -- it's out of service now, so y'all can switch it off whenever. [14:03:20] !log urbanecm@deploy1003 Finished scap sync-world: Backport for [[gerrit:1308106|fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105|fix(MentorListCleaner): Do not access property before inicialization (T430689)]] (duration: 27m 48s) [14:04:03] RESOLVED: [2x] ProbeDown: Service titan1002:443 has failed probes (http_thanos_wikimedia_org_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#titan1002:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [14:04:45] !log disable puppet on A:cp-text to selectively apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1308040 [14:04:46] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:04:49] (03CR) 10Ssingh: [C:03+1] hiera/services: Add IPIP LVS config for ML k8s [puppet] - 10https://gerrit.wikimedia.org/r/1308094 (owner: 10Klausman) [14:05:13] !log installing distro-info-data updates from trixie/bookworm point releases [14:05:14] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:07:35] FIRING: CirrusSearchSaneitizerFixRateTooHigh: MediaWiki CirrusSearch Saneitizer is fixing an abnormally high number of documents in eqiad - https://wikitech.wikimedia.org/wiki/Search/CirrusStreamingUpdater#San(e)itizing - https://grafana.wikimedia.org/d/JLK3I_siz/elasticsearch-indexing?viewPanel=59&orgId=1&from=now-6M&to=now&var-search_cluster=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchSaneitizerFixRateTooHigh [14:07:36] PROBLEM - Recursive DNS on 195.200.68.37 is CRITICAL: DNS_QUERY CRITICAL - query timed out https://wikitech.wikimedia.org/wiki/DNS [14:07:49] (03PS1) 10Urbanecm: Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308112 (https://phabricator.wikimedia.org/T427386) [14:07:51] ^ that's fibe [14:07:53] fine even [14:08:22] (03CR) 10Ssingh: [C:03+1] "Will merge in a bit." [puppet] - 10https://gerrit.wikimedia.org/r/1215329 (owner: 10Dzahn) [14:08:25] (03Abandoned) 10Elukey: docker_registry: add migration block for ML images in nginx [puppet] - 10https://gerrit.wikimedia.org/r/1304596 (https://phabricator.wikimedia.org/T428022) (owner: 10Elukey) [14:08:59] (03CR) 10Fabfur: [C:03+2] cache_misc: apply traffic classification [puppet] - 10https://gerrit.wikimedia.org/r/1308040 (owner: 10Fabfur) [14:09:16] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.5 point update - https://phabricator.wikimedia.org/T427072#12094696 (10MoritzMuehlenhoff) [14:10:15] (03CR) 10JMeybohm: [C:03+1] admin_ng: Install coredns-internalonly in wikikube [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305547 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [14:11:13] !log atsuko@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1120.eqiad.wmnet with OS trixie [14:12:36] PROBLEM - Recursive DNS on 2a02:ec80:700:2:195:200:68:37 is CRITICAL: DNS_QUERY CRITICAL - query timed out https://wikitech.wikimedia.org/wiki/DNS [14:12:46] !log klausman@cumin2002 START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@codfw [14:13:21] (03CR) 10Klausman: [C:03+2] hiera/services: Add IPIP LVS config for ML k8s [puppet] - 10https://gerrit.wikimedia.org/r/1308094 (owner: 10Klausman) [14:14:59] !log blake@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1036.eqiad.wmnet [14:15:00] !log blake@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1036.eqiad.wmnet [14:15:02] !log blake@cumin1003 END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1036.eqiad.wmnet [14:16:17] !log atsuko@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage [14:16:56] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2007.codfw.wmnet with OS trixie [14:17:06] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12094786 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host thanos-fe2007.codfw.wmnet with OS trixie compl... [14:18:08] !log blake@cumin1003 START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1037.eqiad.wmnet [14:18:11] !log blake@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1037.eqiad.wmnet [14:18:36] !log klausman@cumin2002 START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs [14:18:47] !log blake@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1037.eqiad.wmnet [14:18:59] !log blake@cumin1003 START - Cookbook sre.hosts.reimage for host wikikube-worker1037.eqiad.wmnet with OS trixie [14:19:36] !log blake@cumin1003 START - Cookbook sre.hosts.move-vlan for host wikikube-worker1037 [14:19:37] !log klausman@cumin2002 END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs [14:19:38] !log klausman@cumin2002 END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@codfw [14:19:52] !log blake@cumin1003 START - Cookbook sre.dns.netbox [14:20:29] !log atsuko@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage [14:21:57] (03PS12) 10Blake: bgp: Remove port from instance and remote_instance names. [alerts] - 10https://gerrit.wikimedia.org/r/1307787 (https://phabricator.wikimedia.org/T430290) [14:23:46] (03PS19) 10Federico Ceratto: mysql: update replication source [cookbooks] - 10https://gerrit.wikimedia.org/r/1238368 (https://phabricator.wikimedia.org/T373436) [14:23:51] (03CR) 10CI reject: [V:04-1] bgp: Remove port from instance and remote_instance names. [alerts] - 10https://gerrit.wikimedia.org/r/1307787 (https://phabricator.wikimedia.org/T430290) (owner: 10Blake) [14:23:56] (03CR) 10JMeybohm: [C:03+1] function-evaluator, function-orchestrator: Allow customizing dnsConfig [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305797 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [14:24:04] !log blake@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" [14:24:08] !log blake@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003" [14:24:08] !log blake@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [14:24:08] !log blake@cumin1003 START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [14:24:12] !log blake@cumin1003 END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [14:24:22] FIRING: [2x] JobUnavailable: Reduced availability for job haproxy in ops@magru - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [14:24:27] !log blake@cumin1003 START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [14:24:30] !log blake@cumin1003 END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [14:24:50] (03CR) 10JMeybohm: [C:03+1] wikifunctions: Enable custom dnsConfig pointing at coredns-internalonly [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305798 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [14:25:02] !log urbanecm@deploy1003 mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # T427386 [14:25:05] T427386: Deploy automated mentor list cleanup to Wikimedia wikis - https://phabricator.wikimedia.org/T427386 [14:25:10] !log blake@cumin1003 START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [14:25:14] !log blake@cumin1003 END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [14:25:17] jouncebot: nowandnext [14:25:17] For the next 0 hour(s) and 4 minute(s): Test Kitchen UI Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260707T1400) [14:25:17] In 0 hour(s) and 4 minute(s): Test Kitchen Experiment Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260707T1430) [14:25:25] (03PS2) 10Urbanecm: Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308112 (https://phabricator.wikimedia.org/T427386) [14:25:30] (03CR) 10Urbanecm: [C:03+2] Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308112 (https://phabricator.wikimedia.org/T427386) (owner: 10Urbanecm) [14:26:06] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [14:26:26] !log fceratto@cumin1003 END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) [14:26:35] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [14:26:41] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host thanos-fe1007.eqiad.wmnet with OS trixie [14:26:52] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [14:26:53] (03Merged) 10jenkins-bot: Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308112 (https://phabricator.wikimedia.org/T427386) (owner: 10Urbanecm) [14:26:58] (03CR) 10CI reject: [V:04-1] mysql: update replication source [cookbooks] - 10https://gerrit.wikimedia.org/r/1238368 (https://phabricator.wikimedia.org/T373436) (owner: 10Federico Ceratto) [14:27:01] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12094848 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host thanos-fe1007.eqiad.wmnet with OS trixie [14:27:41] !log urbanecm@deploy1003 Started scap sync-world: Backport for [[gerrit:1308112|Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] [14:27:51] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [14:28:08] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [14:28:14] blake@cumin1003 renumber-node (PID 2218232) is awaiting input [14:28:20] (03CR) 10JMeybohm: admin_ng: Restrict wikifunctions network access to DNS (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305799 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [14:28:23] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [14:28:38] (03PS1) 10Btullis: datahub: give the production restore-indices job more heap [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308116 (https://phabricator.wikimedia.org/T402408) [14:28:43] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [14:29:10] !log blake@cumin1003 START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [14:29:13] !log blake@cumin1003 END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [14:29:44] PROBLEM - Check unit status of statograph_post on alert1002 is CRITICAL: CRITICAL: Status of the systemd unit statograph_post https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [14:29:48] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [14:29:49] !log klausman@cumin2002 START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@eqiad [14:29:58] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [14:30:04] Deploy window Test Kitchen Experiment Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260707T1430) [14:31:41] (03CR) 10Btullis: [C:03+2] datahub: give the production restore-indices job more heap [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308116 (https://phabricator.wikimedia.org/T402408) (owner: 10Btullis) [14:31:43] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [14:31:58] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [14:32:13] blake@cumin1003 renumber-node (PID 2218232) is awaiting input [14:32:50] !log blake@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1037 [14:33:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 21.12% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [14:33:21] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [14:33:39] (03PS1) 10Jgiannelos: mw-parsoid: Disable default value for mw-experimental [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308119 [14:33:41] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [14:33:47] (03Merged) 10jenkins-bot: datahub: give the production restore-indices job more heap [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308116 (https://phabricator.wikimedia.org/T402408) (owner: 10Btullis) [14:34:04] !log klausman@cumin2002 START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs [14:34:27] !log urbanecm@deploy1003 Finished scap sync-world: Backport for [[gerrit:1308112|Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] (duration: 06m 47s) [14:34:31] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply [14:34:35] !log urbanecm@deploy1003 mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # T427386 [14:34:39] T427386: Deploy automated mentor list cleanup to Wikimedia wikis - https://phabricator.wikimedia.org/T427386 [14:34:39] !log blake@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1037 [14:34:39] !log blake@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1037 [14:34:57] 10ops-ulsfo, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 10netops: ULSFO: OOB IPV6 down - https://phabricator.wikimedia.org/T430599#12094934 (10Papaul) @cmooney see below email from DR thanks ` From our investigation, we confirmed that the source network (2607:fb58:9000:7::/64) is reachable within ou... [14:35:01] (03PS2) 10Jgiannelos: mw-parsoid: Disable default value for mw-experimental after verifying it in codfw [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308119 (https://phabricator.wikimedia.org/T431387) [14:35:05] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply} [14:35:13] (03PS1) 10Scott French: wmnet: Temporarily direct codfw, eqsin, ulsfo etcd clients to eqiad [dns] - 10https://gerrit.wikimedia.org/r/1308114 (https://phabricator.wikimedia.org/T430909) [14:35:19] !log klausman@cumin2002 END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs [14:35:19] !log klausman@cumin2002 END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@eqiad [14:35:23] (03PS2) 10Scott French: hieradata: Temporarily point codfw PyBals at eqiad etcd [puppet] - 10https://gerrit.wikimedia.org/r/1308115 (https://phabricator.wikimedia.org/T430909) [14:35:27] (03CR) 10Scott French: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308115 (https://phabricator.wikimedia.org/T430909) (owner: 10Scott French) [14:35:57] (03PS3) 10Jgiannelos: mw-parsoid: Disable mw-experimental after verifying it in codfw [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308119 (https://phabricator.wikimedia.org/T431387) [14:36:59] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [14:37:11] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [14:37:52] (03CR) 10Clément Goubert: [C:03+1] mw-parsoid: Disable mw-experimental after verifying it in codfw [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308119 (https://phabricator.wikimedia.org/T431387) (owner: 10Jgiannelos) [14:39:44] RECOVERY - Check unit status of statograph_post on alert1002 is OK: OK: Status of the systemd unit statograph_post https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [14:40:00] (03PS20) 10Federico Ceratto: mysql: update replication source [cookbooks] - 10https://gerrit.wikimedia.org/r/1238368 (https://phabricator.wikimedia.org/T373436) [14:40:15] (03CR) 10JMeybohm: admin_ng: Restrict wikifunctions network access to DNS (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305799 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [14:40:36] !log atsuko@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1083.eqiad.wmnet with OS trixie [14:40:41] (03CR) 10Federico Ceratto: "I updated the CR and also added truncation of the heartbeat table." [cookbooks] - 10https://gerrit.wikimedia.org/r/1238368 (https://phabricator.wikimedia.org/T373436) (owner: 10Federico Ceratto) [14:42:53] (03CR) 10CI reject: [V:04-1] mysql: update replication source [cookbooks] - 10https://gerrit.wikimedia.org/r/1238368 (https://phabricator.wikimedia.org/T373436) (owner: 10Federico Ceratto) [14:42:54] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage [14:45:48] (03CR) 10Ssingh: wmnet: Temporarily direct codfw, eqsin, ulsfo etcd clients to eqiad (031 comment) [dns] - 10https://gerrit.wikimedia.org/r/1308114 (https://phabricator.wikimedia.org/T430909) (owner: 10Scott French) [14:47:29] RECOVERY - Host mr1-ulsfo.oob IPv6 is UP: PING OK - Packet loss = 0%, RTA = 68.02 ms [14:49:13] (03CR) 10Ssingh: [C:03+1] hieradata: Temporarily point codfw PyBals at eqiad etcd [puppet] - 10https://gerrit.wikimedia.org/r/1308115 (https://phabricator.wikimedia.org/T430909) (owner: 10Scott French) [14:49:20] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage [14:50:50] !log Deploying Refinery at 7d8dc71f for change 1308087 / T431318 - update filerevision table sqoop and table [14:50:53] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:50:53] T431318: Fix `filerevision` sqoop job - https://phabricator.wikimedia.org/T431318 [14:51:21] (03CR) 10Ladsgroup: [C:04-1] "Wiki creation shouldn't be done by "normal" deployers. It is fragile, complex and has tendency to break easily and break other wikis. Most" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308070 (https://phabricator.wikimedia.org/T429922) (owner: 10VadymTS1) [14:51:25] !log blake@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage [14:51:30] !log arnaudb@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on phab2003.codfw.wmnet,phab[1004-1006].eqiad.wmnet with reason: maintenance [14:51:46] !log javiermonton@deploy1003 Started deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] [14:51:54] (03CR) 10Scott French: "Thanks, Sukhbir!" [dns] - 10https://gerrit.wikimedia.org/r/1308114 (https://phabricator.wikimedia.org/T430909) (owner: 10Scott French) [14:53:00] (03CR) 10Ssingh: [C:03+1] wmnet: Temporarily direct codfw, eqsin, ulsfo etcd clients to eqiad (031 comment) [dns] - 10https://gerrit.wikimedia.org/r/1308114 (https://phabricator.wikimedia.org/T430909) (owner: 10Scott French) [14:53:46] !log javiermonton@deploy1003 Finished deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] (duration: 02m 00s) [14:53:52] 10ops-ulsfo, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 10netops: ULSFO: OOB IPV6 down - https://phabricator.wikimedia.org/T430599#12095011 (10cmooney) @papaul hmm it looks like it's working different now: ` cmooney@mr1-ulsfo> traceroute 2620:0:861:3:208:80:154:78 routing-instance oob-transit wait 2... [14:53:57] !log javiermonton@deploy1003 Started deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] [14:54:06] (03CR) 10VadymTS1: "Sorry, I wasn't aware of that. Do I need to abandon this change? It was created according to the instructions in the task." [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308070 (https://phabricator.wikimedia.org/T429922) (owner: 10VadymTS1) [14:54:11] (03CR) 10Ladsgroup: [C:04-1] "And please don't just add people asking them to deploy changes. Specially when it's not their job. It's impolite." [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308070 (https://phabricator.wikimedia.org/T429922) (owner: 10VadymTS1) [14:54:59] !log blake@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage [14:55:18] (03CR) 10VadymTS1: "Sorry, okay." [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308070 (https://phabricator.wikimedia.org/T429922) (owner: 10VadymTS1) [14:56:47] (03CR) 10Dpogorzelski: [C:03+1] hiera/manifests/install: Remove ml-cache machines [puppet] - 10https://gerrit.wikimedia.org/r/1306682 (https://phabricator.wikimedia.org/T430654) (owner: 10Klausman) [14:58:12] !log javiermonton@deploy1003 Finished deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] (duration: 04m 14s) [14:58:29] !log javiermonton@deploy1003 Started deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] [14:58:54] PROBLEM - Host mr1-ulsfo.oob IPv6 is DOWN: PING CRITICAL - Packet loss = 100% [14:59:52] (03PS1) 10Cathal Mooney: Eqiad AVOIDED-PATHS: add path to TELX/DR through HE [homer/public] - 10https://gerrit.wikimedia.org/r/1308123 (https://phabricator.wikimedia.org/T430599) [15:00:05] jelto, arnoldokoth, mutante, and arnaudb: Time to do the SRE Collaboration Services office hours deploy. Don't look at me like that. You signed up for it. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260707T1500). [15:00:19] (03PS13) 10Blake: bgp: Remove port from instance and remote_instance names. [alerts] - 10https://gerrit.wikimedia.org/r/1307787 (https://phabricator.wikimedia.org/T430290) [15:00:40] !log javiermonton@deploy1003 Finished deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] (duration: 02m 10s) [15:03:05] !log brennen@deploy1003 Started deploy [phabricator/deployment@7e02037]: deploy phab2003 for T431440 [15:03:08] T431440: Deploy Phab/Phorge 2026-07-07 - https://phabricator.wikimedia.org/T431440 [15:03:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.49% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [15:03:55] !log brennen@deploy1003 Finished deploy [phabricator/deployment@7e02037]: deploy phab2003 for T431440 (duration: 00m 51s) [15:04:11] !log brennen@deploy1003 Started deploy [phabricator/deployment@7e02037]: deploy phab1004 for T431440 [15:04:46] FIRING: [4x] GerritHAProxyBackendUnavailable: Gerrit backend is unavilable for tcp-proxy (HAProxy) gerrit_ssh - https://wikitech.wikimedia.org/wiki/Gerrit/Operations#GerritHAProxyBackendUnavailable - grafana.wikimedia.org/d/459365f6-df37-48d6-8142-82b22c1875e7/gerrit-tcp-proxy?viewPanel=panel-15 - https://alerts.wikimedia.org/?q=alertname%3DGerritHAProxyBackendUnavailable [15:04:51] FIRING: [2x] GerritHAProxyServiceUnavailable: Gerrit tcp-proxy (HAProxy) service gerrit_ssh is DOWN in codfw - https://wikitech.wikimedia.org/wiki/Gerrit/Operations#GerritHAProxyServiceUnavailable - grafana.wikimedia.org/d/459365f6-df37-48d6-8142-82b22c1875e7/gerrit-tcp-proxy?viewPanel=panel-15 - https://alerts.wikimedia.org/?q=alertname%3DGerritHAProxyServiceUnavailable [15:04:58] !log brennen@deploy1003 Finished deploy [phabricator/deployment@7e02037]: deploy phab1004 for T431440 (duration: 00m 47s) [15:05:31] FIRING: [4x] ProbeDown: Service gerrit2003:29418 has failed probes (tcp_gerrit_ssh_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [15:05:43] atsuko@cumin1003 reimage (PID 2246350) is awaiting input [15:05:45] I was about to ask - is gerrit down? :D [15:06:21] was it under maintenance or did it just restart? [15:06:36] elukey: scrapers, most likely. being investigated. [15:06:57] (03PS1) 10Filippo Giunchedi: dumps: add nfs_client for dumps-nfs.w.o [puppet] - 10https://gerrit.wikimedia.org/r/1308126 (https://phabricator.wikimedia.org/T411248) [15:06:58] (03PS1) 10Filippo Giunchedi: statistics: optional support for dumps-nfs.w.o [puppet] - 10https://gerrit.wikimedia.org/r/1308127 (https://phabricator.wikimedia.org/T411248) [15:06:58] (03PS1) 10Filippo Giunchedi: wmcs: add new dumps-nfs.w.o support [puppet] - 10https://gerrit.wikimedia.org/r/1308128 (https://phabricator.wikimedia.org/T411248) [15:06:58] (03PS1) 10Filippo Giunchedi: pontoon: better pre-receive check for unclean destination repository [puppet] - 10https://gerrit.wikimedia.org/r/1308129 [15:07:35] brennen: o/ the gerrit systemd unit shows 1 minute of uptime, so it may have crashed [15:08:14] !log andrew@cumin2002 START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org [15:08:31] !log andrew@cumin2002 END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org [15:08:34] brennen: Terminating due to java.lang.OutOfMemoryError: Java heap space [15:08:41] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1007.eqiad.wmnet with OS trixie [15:08:52] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12095173 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host thanos-fe1007.eqiad.wmnet with OS trixie compl... [15:08:55] cc: arnaudb, jelto --^ [15:09:00] (03CR) 10CI reject: [V:04-1] dumps: add nfs_client for dumps-nfs.w.o [puppet] - 10https://gerrit.wikimedia.org/r/1308126 (https://phabricator.wikimedia.org/T411248) (owner: 10Filippo Giunchedi) [15:09:24] !log atsuko@cumin1003 START - Cookbook sre.hosts.reimage for host cirrussearch1121.eqiad.wmnet with OS trixie [15:09:38] !log andrew@cumin2002 START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org [15:09:43] !log andrew@cumin2002 END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org [15:09:46] RESOLVED: [4x] GerritHAProxyBackendUnavailable: Gerrit backend is unavilable for tcp-proxy (HAProxy) gerrit_ssh - https://wikitech.wikimedia.org/wiki/Gerrit/Operations#GerritHAProxyBackendUnavailable - grafana.wikimedia.org/d/459365f6-df37-48d6-8142-82b22c1875e7/gerrit-tcp-proxy?viewPanel=panel-15 - https://alerts.wikimedia.org/?q=alertname%3DGerritHAProxyBackendUnavailable [15:09:51] RESOLVED: [2x] GerritHAProxyServiceUnavailable: Gerrit tcp-proxy (HAProxy) service gerrit_ssh is DOWN in codfw - https://wikitech.wikimedia.org/wiki/Gerrit/Operations#GerritHAProxyServiceUnavailable - grafana.wikimedia.org/d/459365f6-df37-48d6-8142-82b22c1875e7/gerrit-tcp-proxy?viewPanel=panel-15 - https://alerts.wikimedia.org/?q=alertname%3DGerritHAProxyServiceUnavailable [15:09:52] !log atsuko@cumin1003 START - Cookbook sre.hosts.move-vlan for host cirrussearch1121 [15:09:57] !log andrew@cumin2002 START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org [15:10:31] RESOLVED: [4x] ProbeDown: Service gerrit2003:29418 has failed probes (tcp_gerrit_ssh_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [15:11:14] !log atsuko@cumin1003 START - Cookbook sre.dns.netbox [15:11:30] (03CR) 10Jgiannelos: [C:03+2] mw-parsoid: Disable mw-experimental after verifying it in codfw [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308119 (https://phabricator.wikimedia.org/T431387) (owner: 10Jgiannelos) [15:12:10] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - dumps-lb6_2049: Servers clouddumps1002.wikimedia.org are marked down but pooled: dumps-lb_2049: Servers clouddumps1002.wikimedia.org are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [15:12:30] (03Abandoned) 10VadymTS1: Init Wikiquote Minangkabau [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308070 (https://phabricator.wikimedia.org/T429922) (owner: 10VadymTS1) [15:12:54] PROBLEM - PyBal backends health check on lvs1018 is CRITICAL: PYBAL CRITICAL - CRITICAL - dumps-lb6_2049: Servers clouddumps1002.wikimedia.org are marked down but pooled: dumps-lb_2049: Servers clouddumps1002.wikimedia.org are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [15:13:46] (03CR) 10Ladsgroup: "nah, don't worry. I get the new wiki created later today." [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308070 (https://phabricator.wikimedia.org/T429922) (owner: 10VadymTS1) [15:14:06] !log blake@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1037.eqiad.wmnet with OS trixie [15:14:10] (03Merged) 10jenkins-bot: mw-parsoid: Disable mw-experimental after verifying it in codfw [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308119 (https://phabricator.wikimedia.org/T431387) (owner: 10Jgiannelos) [15:14:15] 07sre-alert-triage, 06Data-Platform-SRE (2026-07-03 - 2026-07-31): Alert in need of triage: KubernetesAPIErrorRate - https://phabricator.wikimedia.org/T414413#12095199 (10BTullis) 05Open→03Resolved a:03BTullis We're not too concerned about 404s at the moment, but we can try to look again if this cont... [15:14:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.93% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [15:14:43] (03PS2) 10Aklapper: phabricator: drop diffusion.ssh-host config [puppet] - 10https://gerrit.wikimedia.org/r/1305044 (https://phabricator.wikimedia.org/T429367) [15:14:56] (03PS1) 10Filippo Giunchedi: fixup! dumps: add nfs_client for dumps-nfs.w.o [puppet] - 10https://gerrit.wikimedia.org/r/1308131 [15:15:48] (03CR) 10Aklapper: "Should be fine to review/merge now for next time" [puppet] - 10https://gerrit.wikimedia.org/r/1305044 (https://phabricator.wikimedia.org/T429367) (owner: 10Aklapper) [15:15:57] (03PS2) 10Filippo Giunchedi: dumps: add nfs_client for dumps-nfs.w.o [puppet] - 10https://gerrit.wikimedia.org/r/1308126 (https://phabricator.wikimedia.org/T411248) [15:15:57] (03PS2) 10Filippo Giunchedi: statistics: optional support for dumps-nfs.w.o [puppet] - 10https://gerrit.wikimedia.org/r/1308127 (https://phabricator.wikimedia.org/T411248) [15:15:57] (03PS2) 10Filippo Giunchedi: wmcs: add new dumps-nfs.w.o support [puppet] - 10https://gerrit.wikimedia.org/r/1308128 (https://phabricator.wikimedia.org/T411248) [15:15:58] (03PS2) 10Filippo Giunchedi: pontoon: better pre-receive check for unclean destination repository [puppet] - 10https://gerrit.wikimedia.org/r/1308129 [15:16:02] (03Abandoned) 10Filippo Giunchedi: fixup! dumps: add nfs_client for dumps-nfs.w.o [puppet] - 10https://gerrit.wikimedia.org/r/1308131 (owner: 10Filippo Giunchedi) [15:16:21] !log atsuko@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" [15:16:25] !log atsuko@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003" [15:16:26] !log atsuko@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [15:16:26] !log atsuko@cumin1003 START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [15:16:29] !log atsuko@cumin1003 END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [15:16:48] !log cgoubert@deploy1003 helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply [15:16:48] !log atsuko@cumin1003 START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [15:16:51] !log atsuko@cumin1003 END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [15:16:56] !log cgoubert@deploy1003 helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply [15:17:06] blake@cumin1003 renumber-node (PID 2218232) is awaiting input [15:17:40] !log atsuko@cumin1003 START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [15:17:44] !log atsuko@cumin1003 END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [15:17:49] !log atsuko@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1121 [15:18:10] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [15:18:20] !log andrew@cumin2002 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1002.wikimedia.org [15:18:21] !log atsuko@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1121 [15:18:21] !log atsuko@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1121 [15:18:48] (03CR) 10VadymTS1: "I respect criticism, so for me, this is fine. Thank you." [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308070 (https://phabricator.wikimedia.org/T429922) (owner: 10VadymTS1) [15:18:50] (03CR) 10Filippo Giunchedi: [C:03+1] openstack: wmcs-webproxy: Use singular 'backend' for API requests [puppet] - 10https://gerrit.wikimedia.org/r/1308062 (https://phabricator.wikimedia.org/T429960) (owner: 10Majavah) [15:18:54] RECOVERY - PyBal backends health check on lvs1018 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [15:19:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.93% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [15:19:18] (03CR) 10CI reject: [V:04-1] dumps: add nfs_client for dumps-nfs.w.o [puppet] - 10https://gerrit.wikimedia.org/r/1308126 (https://phabricator.wikimedia.org/T411248) (owner: 10Filippo Giunchedi) [15:19:26] (03CR) 10Filippo Giunchedi: [C:03+1] dynamicproxy: Allow using singular 'backend' for writing data [puppet] - 10https://gerrit.wikimedia.org/r/1308061 (https://phabricator.wikimedia.org/T429960) (owner: 10Majavah) [15:19:42] (03CR) 10Filippo Giunchedi: [C:03+1] dynamicproxy: Include singular 'backend' key with API responses [puppet] - 10https://gerrit.wikimedia.org/r/1308060 (https://phabricator.wikimedia.org/T429960) (owner: 10Majavah) [15:20:28] (03CR) 10Filippo Giunchedi: [C:03+1] haproxy: Make logging work with default config [puppet] - 10https://gerrit.wikimedia.org/r/1308053 (owner: 10Majavah) [15:20:48] !log andrew@cumin2002 START - Cookbook sre.hosts.reboot-single for host clouddumps1001.wikimedia.org [15:21:16] (03PS3) 10Filippo Giunchedi: pontoon: better pre-receive check for unclean destination repository [puppet] - 10https://gerrit.wikimedia.org/r/1308129 [15:21:49] (03PS1) 10Kamila Součková: wmnet: update deployment CNAME record to deploy2003 [dns] - 10https://gerrit.wikimedia.org/r/1308134 (https://phabricator.wikimedia.org/T423714) [15:22:09] (03CR) 10Filippo Giunchedi: [C:03+2] pontoon: better pre-receive check for unclean destination repository [puppet] - 10https://gerrit.wikimedia.org/r/1308129 (owner: 10Filippo Giunchedi) [15:22:24] (03CR) 10Majavah: [C:03+2] dynamicproxy: Include singular 'backend' key with API responses [puppet] - 10https://gerrit.wikimedia.org/r/1308060 (https://phabricator.wikimedia.org/T429960) (owner: 10Majavah) [15:22:41] (03CR) 10Majavah: [C:03+2] dynamicproxy: Allow using singular 'backend' for writing data [puppet] - 10https://gerrit.wikimedia.org/r/1308061 (https://phabricator.wikimedia.org/T429960) (owner: 10Majavah) [15:22:45] (03CR) 10Kamila Součková: "To be deployed during Wednesday UTC late infra window." [dns] - 10https://gerrit.wikimedia.org/r/1308134 (https://phabricator.wikimedia.org/T423714) (owner: 10Kamila Součková) [15:22:59] (03CR) 10Majavah: [C:03+2] openstack: wmcs-webproxy: Use singular 'backend' for API requests [puppet] - 10https://gerrit.wikimedia.org/r/1308062 (https://phabricator.wikimedia.org/T429960) (owner: 10Majavah) [15:24:10] (03CR) 10Majavah: [V:03+1 C:03+2] haproxy: Make logging work with default config [puppet] - 10https://gerrit.wikimedia.org/r/1308053 (owner: 10Majavah) [15:25:03] (03PS1) 10Andrew Bogott: scan-and-shrink: raise deletion threshold score from 2 to 3 [puppet] - 10https://gerrit.wikimedia.org/r/1308135 (https://phabricator.wikimedia.org/T431340) [15:26:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.25% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [15:26:17] 10ops-eqiad, 06SRE, 06cloud-services-team, 10Cloud-VPS, 06DC-Ops: New thermal paste for cloudvirt1071.eqiad.wmnet - https://phabricator.wikimedia.org/T431429#12095333 (10Andrew) Let's do a memory scan on this server too, while it's out of service. Thanks! [15:27:05] (03CR) 10Scott French: [C:03+1] wmnet: update deployment CNAME record to deploy2003 [dns] - 10https://gerrit.wikimedia.org/r/1308134 (https://phabricator.wikimedia.org/T423714) (owner: 10Kamila Součková) [15:28:46] (03CR) 10Scott French: [C:03+2] hieradata: Mark configcluster hosts as manual depool [puppet] - 10https://gerrit.wikimedia.org/r/1307852 (https://phabricator.wikimedia.org/T430930) (owner: 10Scott French) [15:29:21] !log andrew@cumin2002 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1001.wikimedia.org [15:30:05] jelto, arnoldokoth, mutante, and arnaudb: I seem to be stuck in Groundhog week. Sigh. Time for (yet another) SRE Collaboration Services office hours deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260707T1500). [15:30:05] mutante and hashar: Deploy window CI jenkins maintenance downtime, CI may not work during this time (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260707T1530) [15:30:42] !log atsuko@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage [15:32:31] FIRING: [2x] ProbeDown: Service gerrit2003:443 has failed probes (http_gerrit_tls_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#gerrit2003:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [15:33:16] !log cdobbins@cumin1003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host dns7002.wikimedia.org with OS trixie [15:33:52] !log cdobbins@cumin1003 START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS bookworm [15:34:46] FIRING: [4x] GerritHAProxyBackendUnavailable: Gerrit backend is unavilable for tcp-proxy (HAProxy) gerrit_ssh - https://wikitech.wikimedia.org/wiki/Gerrit/Operations#GerritHAProxyBackendUnavailable - grafana.wikimedia.org/d/459365f6-df37-48d6-8142-82b22c1875e7/gerrit-tcp-proxy?viewPanel=panel-15 - https://alerts.wikimedia.org/?q=alertname%3DGerritHAProxyBackendUnavailable [15:34:56] !log atsuko@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage [15:34:57] FIRING: [2x] GerritHAProxyServiceUnavailable: Gerrit tcp-proxy (HAProxy) service gerrit_ssh is DOWN in codfw - https://wikitech.wikimedia.org/wiki/Gerrit/Operations#GerritHAProxyServiceUnavailable - grafana.wikimedia.org/d/459365f6-df37-48d6-8142-82b22c1875e7/gerrit-tcp-proxy?viewPanel=panel-15 - https://alerts.wikimedia.org/?q=alertname%3DGerritHAProxyServiceUnavailable [15:36:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.45% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [15:37:31] FIRING: [4x] ProbeDown: Service gerrit2003:29418 has failed probes (tcp_gerrit_ssh_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [15:39:22] FIRING: [4x] JobUnavailable: Reduced availability for job gerrit in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [15:39:46] RESOLVED: [4x] GerritHAProxyBackendUnavailable: Gerrit backend is unavilable for tcp-proxy (HAProxy) gerrit_ssh - https://wikitech.wikimedia.org/wiki/Gerrit/Operations#GerritHAProxyBackendUnavailable - grafana.wikimedia.org/d/459365f6-df37-48d6-8142-82b22c1875e7/gerrit-tcp-proxy?viewPanel=panel-15 - https://alerts.wikimedia.org/?q=alertname%3DGerritHAProxyBackendUnavailable [15:39:57] RESOLVED: [2x] GerritHAProxyServiceUnavailable: Gerrit tcp-proxy (HAProxy) service gerrit_ssh is DOWN in codfw - https://wikitech.wikimedia.org/wiki/Gerrit/Operations#GerritHAProxyServiceUnavailable - grafana.wikimedia.org/d/459365f6-df37-48d6-8142-82b22c1875e7/gerrit-tcp-proxy?viewPanel=panel-15 - https://alerts.wikimedia.org/?q=alertname%3DGerritHAProxyServiceUnavailable [15:40:20] blake@cumin1003 renumber-node (PID 2218232) is awaiting input [15:41:08] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host thanos-fe1006.eqiad.wmnet with OS trixie [15:42:31] RESOLVED: [4x] ProbeDown: Service gerrit2003:29418 has failed probes (tcp_gerrit_ssh_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [15:42:46] (03Abandoned) 10BCornwall: cache_misc: apply traffic classification [puppet] - 10https://gerrit.wikimedia.org/r/1307824 (owner: 10BCornwall) [15:42:50] !log blake@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1037.eqiad.wmnet [15:42:52] !log blake@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1037.eqiad.wmnet [15:42:54] !log blake@cumin1003 END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1037.eqiad.wmnet [15:44:22] FIRING: [5x] JobUnavailable: Reduced availability for job gerrit in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [15:44:33] (03CR) 10Dzahn: [C:03+2] ci: disable jenkins on role::ci machines [puppet] - 10https://gerrit.wikimedia.org/r/1307850 (https://phabricator.wikimedia.org/T418109) (owner: 10Dzahn) [15:44:41] (03CR) 10Cathal Mooney: [C:03+2] Eqiad AVOIDED-PATHS: add path to TELX/DR through HE [homer/public] - 10https://gerrit.wikimedia.org/r/1308123 (https://phabricator.wikimedia.org/T430599) (owner: 10Cathal Mooney) [15:45:22] (03CR) 10Clément Goubert: "There's no reason for an emergency SRE deployment of this patch, it is scheduled for tomorrow's backport window and in my opinion can occu" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308070 (https://phabricator.wikimedia.org/T429922) (owner: 10VadymTS1) [15:46:03] (03Merged) 10jenkins-bot: Eqiad AVOIDED-PATHS: add path to TELX/DR through HE [homer/public] - 10https://gerrit.wikimedia.org/r/1308123 (https://phabricator.wikimedia.org/T430599) (owner: 10Cathal Mooney) [15:47:13] (03PS1) 10Kamila Součková: hiera: update deployment server to deploy2003 [puppet] - 10https://gerrit.wikimedia.org/r/1308147 (https://phabricator.wikimedia.org/T423714) [15:49:05] (03CR) 10Blake: "After discussion on IRC: 'hostname' isn't what's used in the matcher, it is, in fact, 'instance'. And the matcher makes allowances for the" [alerts] - 10https://gerrit.wikimedia.org/r/1307787 (https://phabricator.wikimedia.org/T430290) (owner: 10Blake) [15:49:36] (03CR) 10SBassett: Revert^2 "varnish: Add CSP report-only header value" (033 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1307826 (owner: 10BCornwall) [15:50:23] (03PS1) 10Arnaudb: gerrit: filter paths on httpd [puppet] - 10https://gerrit.wikimedia.org/r/1308148 (https://phabricator.wikimedia.org/T431398) [15:51:17] (03CR) 10Blake: "recheck" [alerts] - 10https://gerrit.wikimedia.org/r/1307787 (https://phabricator.wikimedia.org/T430290) (owner: 10Blake) [15:51:30] PROBLEM - jenkins_service_running on contint1002 is CRITICAL: PROCS CRITICAL: 0 processes with regex args .*/bin/java .*-jar /usr/share/java/jenkins.war https://wikitech.wikimedia.org/wiki/Jenkins [15:54:16] ^ maintenance period [15:54:27] !log jenkins down in planned maintenance window [15:54:28] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [15:56:29] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage [15:57:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.17% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [15:58:42] !log cdobbins@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage [15:58:49] !log atsuko@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1121.eqiad.wmnet with OS trixie [15:59:42] PROBLEM - Host cloudrabbit2001-dev is DOWN: PING CRITICAL - Packet loss = 100% [16:00:02] (03CR) 10Dzahn: [C:03+2] Revert^2 "integration: switch integration-agent-docker VMs to Java 21" [puppet] - 10https://gerrit.wikimedia.org/r/1307851 (owner: 10Dzahn) [16:00:05] mutante and hashar: Deploy window CI jenkins maintenance downtime, CI may not work during this time (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260707T1530) [16:00:05] jhathaway and rzl: #bothumor My software never has bugs. It just develops random features. Rise for Puppet request window. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260707T1600). [16:00:05] Dreamy_Jazz, kostajh, and dancy: A patch you scheduled for Puppet request window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [16:00:11] o/ [16:00:15] \o [16:00:30] rzl: there is currently no jenkins, it's in maintenance. if it matters. [16:01:01] (03PS1) 10Btullis: Switch new an-test-master and an-test-coord nodes into their role [puppet] - 10https://gerrit.wikimedia.org/r/1308149 (https://phabricator.wikimedia.org/T406746) [16:01:11] mutante: it might :) when is it expected back? [16:01:16] (03CR) 10Scott French: [C:03+1] "Looks good!" [puppet] - 10https://gerrit.wikimedia.org/r/1308147 (https://phabricator.wikimedia.org/T423714) (owner: 10Kamila Součková) [16:01:19] RECOVERY - Host cloudrabbit2001-dev is UP: PING OK - Packet loss = 0%, RTA = 31.61 ms [16:01:19] RECOVERY - Host mr1-ulsfo.oob IPv6 is UP: PING OK - Packet loss = 0%, RTA = 68.85 ms [16:01:20] rzl: half an hour [16:01:25] hi [16:01:25] oof. okay [16:01:33] (03PS1) 10CDobbins: hieradata: set `use_new_pdns_cfg` flag to true [puppet] - 10https://gerrit.wikimedia.org/r/1308150 (https://phabricator.wikimedia.org/T401832) [16:02:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.66% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [16:02:21] (03CR) 10Ssingh: "Looks good, will +1 when we have verified the config offline." [puppet] - 10https://gerrit.wikimedia.org/r/1308150 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [16:02:33] kostajh: hi! sorry but I can't take https://gerrit.wikimedia.org/r/1307100 in the puppet window -- please ask Data Persistence to review it, do you need me to find someone for you? [16:03:02] I already asked in their slack [16:03:03] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage [16:03:08] I think they said they'd deploy it [16:03:11] On monday [16:03:15] But probably didn't happen [16:04:29] got it - I'd still like to get their review. we can follow up afterward, it doesn't have to wait for the next window, it can be merged whenever, but I don't know enough about that to be confident reviewing it for you [16:04:34] rzl: if it already got a V+2 I think it should be ok, depending on the type of patch [16:04:40] to just V+2 I mean [16:05:21] jhathaway: if you're around, are you happy with https://gerrit.wikimedia.org/r/1306745? I see you took a look already [16:05:36] o/, looking... [16:05:39] (03PS1) 10Clément Goubert: redioscope: fix metrics prefix [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308151 [16:05:45] (03CR) 10Andrew Bogott: [C:03+2] scan-and-shrink: raise deletion threshold score from 2 to 3 [puppet] - 10https://gerrit.wikimedia.org/r/1308135 (https://phabricator.wikimedia.org/T431340) (owner: 10Andrew Bogott) [16:05:55] Yeah, fair re the wikireplicas patch [16:06:25] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage [16:06:29] I think Kosta added it just to see if it could be moved along :D [16:06:39] I'll talk to them on slack about coordinating that [16:06:49] PROBLEM - Host cloudrabbit2002-dev is DOWN: PING CRITICAL - Packet loss = 100% [16:07:13] rzl: yes I'm okay with it [16:07:13] yeah, understand :) thanks, let me know if you need me to help coordinate [16:07:18] yes, just trying to move it along [16:07:25] jhathaway: thanks! [16:07:27] (03CR) 10JHathaway: [C:03+1] profile::puppet::agent: write pinned CA cert with a trailing newline [puppet] - 10https://gerrit.wikimedia.org/r/1306745 (https://phabricator.wikimedia.org/T429413) (owner: 10Ahmon Dancy) [16:07:43] (03CR) 10RLazarus: [C:03+2] profile::puppet::agent: write pinned CA cert with a trailing newline [puppet] - 10https://gerrit.wikimedia.org/r/1306745 (https://phabricator.wikimedia.org/T429413) (owner: 10Ahmon Dancy) [16:07:45] o/ [16:08:08] rzl: who should we tag to look at https://gerrit.wikimedia.org/r/c/operations/puppet/+/1307100 ? [16:08:19] RECOVERY - Host cloudrabbit2002-dev is UP: PING OK - Packet loss = 0%, RTA = 31.62 ms [16:08:25] kostajh: data persistence, Dreamy_Jazz and I were just talking about that [16:08:42] sure, was looking for a suggestion about a person [16:08:51] We are in the data-engineering channel on Slack, is that the same people? [16:08:53] I've pinged them again at the slack thread I made about this last week [16:09:02] no, that's a different team :) data persistence sre [16:09:22] FIRING: [4x] JobUnavailable: Reduced availability for job haproxy in ops@magru - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:09:24] Last time we were modifying this I went to the data engineering channel [16:09:30] (They pointed me there) [16:09:34] oh, okay [16:09:39] PROBLEM - Recursive DNS on 195.200.68.37 is CRITICAL: DNS_QUERY CRITICAL - query timed out https://wikitech.wikimedia.org/wiki/DNS [16:10:08] sounds like you know more than me, in that case :) [16:10:21] :D [16:10:24] (03PS4) 10Dzahn: contint: stop including local jenkins proxy, switch ext proxy to /ci [puppet] - 10https://gerrit.wikimedia.org/r/1306446 (https://phabricator.wikimedia.org/T418521) [16:10:39] I've had to make several security fixes to that file before :D [16:10:46] (03CR) 10Hashar: [C:03+1] contint: stop including local jenkins proxy, switch ext proxy to /ci [puppet] - 10https://gerrit.wikimedia.org/r/1306446 (https://phabricator.wikimedia.org/T418521) (owner: 10Dzahn) [16:11:18] (03CR) 10Dzahn: [V:03+2 C:03+2] contint: stop including local jenkins proxy, switch ext proxy to /ci [puppet] - 10https://gerrit.wikimedia.org/r/1306446 (https://phabricator.wikimedia.org/T418521) (owner: 10Dzahn) [16:11:33] in the meantime I'm not sure if I want to start that LVS removal without Jenkins available -- I know there's deployment server work scheduled for 17:00, so we have a hard stop, and 30 minutes at the end of the window might not be enough [16:11:57] I'm sorry, I know we already pushed this off, I wish I'd known that maintenance was scheduled [16:12:31] PROBLEM - Host cloudrabbit2003-dev is DOWN: PING CRITICAL - Packet loss = 100% [16:14:06] Sure, it's as it goes. Not a priority from PSI to get this undeployed. Should we pick the next free puppet window again 😀 [16:14:19] RECOVERY - Host cloudrabbit2003-dev is UP: PING OK - Packet loss = 0%, RTA = 31.63 ms [16:14:22] FIRING: [4x] JobUnavailable: Reduced availability for job haproxy in ops@magru - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:14:39] PROBLEM - Recursive DNS on 2a02:ec80:700:2:195:200:68:37 is CRITICAL: DNS_QUERY CRITICAL - query timed out https://wikitech.wikimedia.org/wiki/DNS [16:15:01] I think that's probably best -- it conflicts with the end of the WMF staff meeting, but that's okay with me if it is for you [16:15:25] if the CI maintenance ends on time, maybe we go ahead with the LVS part of the work today, and leave the rest for Thursday [16:16:15] Yeah Thursday would be fine for me [16:16:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.93% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [16:16:32] I usually idly write PHP through the staff meeting anyway 😀 [16:16:39] :o [16:16:41] RECOVERY - jenkins_service_running on contint1003 is OK: PROCS OK: 1 process with regex args .*/bin/java .*-jar /usr/share/java/jenkins.war https://wikitech.wikimedia.org/wiki/Jenkins [16:17:07] I'll add it to the calendar then [16:19:37] (03PS1) 10Cathal Mooney: HE Transport esams<->eqiad, lower OSPF cost to make live [homer/public] - 10https://gerrit.wikimedia.org/r/1308152 (https://phabricator.wikimedia.org/T424839) [16:20:50] lgtm, thanks [16:21:14] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1006.eqiad.wmnet with OS trixie [16:21:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.93% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [16:22:11] (03PS1) 10Dzahn: jenkins: switch jenkins prefix to /ci on new hosts [puppet] - 10https://gerrit.wikimedia.org/r/1308153 (https://phabricator.wikimedia.org/T418521) [16:22:45] (03CR) 10CDobbins: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308150 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [16:23:21] (03CR) 10Dzahn: [V:03+2 C:03+2] jenkins: switch jenkins prefix to /ci on new hosts [puppet] - 10https://gerrit.wikimedia.org/r/1308153 (https://phabricator.wikimedia.org/T418521) (owner: 10Dzahn) [16:25:19] (03PS1) 10Dzahn: jenkins: enable jenkins service on new host [puppet] - 10https://gerrit.wikimedia.org/r/1308154 (https://phabricator.wikimedia.org/T418521) [16:26:27] (03CR) 10Dzahn: [V:03+2 C:03+2] jenkins: enable jenkins service on new host [puppet] - 10https://gerrit.wikimedia.org/r/1308154 (https://phabricator.wikimedia.org/T418521) (owner: 10Dzahn) [16:26:37] (03PS2) 10Dzahn: jenkins: enable jenkins service on new host [puppet] - 10https://gerrit.wikimedia.org/r/1308154 (https://phabricator.wikimedia.org/T418521) [16:26:42] (03CR) 10Dzahn: [V:03+2 C:03+2] jenkins: enable jenkins service on new host [puppet] - 10https://gerrit.wikimedia.org/r/1308154 (https://phabricator.wikimedia.org/T418521) (owner: 10Dzahn) [16:27:12] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host cirrussearch1111.eqiad.wmnet with OS trixie [16:29:22] FIRING: [4x] JobUnavailable: Reduced availability for job haproxy in ops@magru - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:30:19] (03CR) 10Zabe: "Yes, Amir and I usually do the new wiki creations. But I would do it at a time where I have time for it and not during tomorrow's window." [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308070 (https://phabricator.wikimedia.org/T429922) (owner: 10VadymTS1) [16:33:32] (03CR) 10Dzahn: [V:03+1 C:03+1] "recheck" [puppet] - 10https://gerrit.wikimedia.org/r/1215329 (owner: 10Dzahn) [16:35:04] (03PS1) 10Ladsgroup: swift: Load the file async [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1308156 (https://phabricator.wikimedia.org/T333445) [16:35:31] RECOVERY - jenkins_service_running on contint1002 is OK: PROCS OK: 1 process with regex args .*/bin/java .*-jar /usr/share/java/jenkins.war https://wikitech.wikimedia.org/wiki/Jenkins [16:35:44] !log cmooney@cumin1003 START - Cookbook sre.network.peering with action 'configure' for AS: 47794 [16:36:16] Dreamy_Jazz: okay, let's see how far we get :) [16:36:31] Yeah :) [16:36:33] Thanks [16:36:39] (03CR) 10RLazarus: [C:03+2] "Done" [puppet] - 10https://gerrit.wikimedia.org/r/1306899 (https://phabricator.wikimedia.org/T416623) (owner: 10Clément Goubert) [16:38:14] (03PS1) 10Andrew Bogott: daily_file_cleanup: remove --dry-run [puppet] - 10https://gerrit.wikimedia.org/r/1308157 (https://phabricator.wikimedia.org/T431340) [16:38:31] PROBLEM - jenkins_service_running on contint1002 is CRITICAL: PROCS CRITICAL: 0 processes with regex args .*/bin/java .*-jar /usr/share/java/jenkins.war https://wikitech.wikimedia.org/wiki/Jenkins [16:38:33] (03CR) 10Ladsgroup: "I can't test it. Is there a way to test this change somewhere? 😞" [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1308156 (https://phabricator.wikimedia.org/T333445) (owner: 10Ladsgroup) [16:38:34] !log cmooney@cumin1003 END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'configure' for AS: 47794 [16:39:14] running puppet on O:lvs::balancer now [16:39:37] RECOVERY - Recursive DNS on 195.200.68.37 is OK: DNS_QUERY OK - Success https://wikitech.wikimedia.org/wiki/DNS [16:39:37] RECOVERY - Recursive DNS on 2a02:ec80:700:2:195:200:68:37 is OK: DNS_QUERY OK - Success https://wikitech.wikimedia.org/wiki/DNS [16:40:15] !log bking@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage [16:40:51] (03CR) 10Andrew Bogott: [C:03+2] daily_file_cleanup: remove --dry-run [puppet] - 10https://gerrit.wikimedia.org/r/1308157 (https://phabricator.wikimedia.org/T431340) (owner: 10Andrew Bogott) [16:40:57] disabled at lvs2012 for maintenance, expected, and already in progress at lvs2014, but done now [16:43:33] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage [16:44:22] RESOLVED: JobUnavailable: Reduced availability for job pdnsrec in ops@magru - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:47:34] (03CR) 10Hnowlan: prometheus: refactor prometheus-es-exporter to use config file (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1305718 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [16:48:20] (03CR) 10Hnowlan: prometheus: refactor prometheus-es-exporter to use config file (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1305718 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [16:48:50] (03PS1) 10Dzahn: jenkins: add firewall rule for controller to agent [puppet] - 10https://gerrit.wikimedia.org/r/1308158 (https://phabricator.wikimedia.org/T418521) [16:49:40] (03PS2) 10Dzahn: jenkins: add firewall rule for controller to agent [puppet] - 10https://gerrit.wikimedia.org/r/1308158 (https://phabricator.wikimedia.org/T418521) [16:50:02] (03PS3) 10Dzahn: jenkins: add firewall rule for controller to agent [puppet] - 10https://gerrit.wikimedia.org/r/1308158 (https://phabricator.wikimedia.org/T418521) [16:50:23] (03CR) 10Dzahn: [V:03+2 C:03+2] jenkins: add firewall rule for controller to agent [puppet] - 10https://gerrit.wikimedia.org/r/1308158 (https://phabricator.wikimedia.org/T418521) (owner: 10Dzahn) [16:51:31] RECOVERY - BFD status on asw1-b4-magru.mgmt is OK: UP: 2 AdminDown: 0 Down: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [16:51:32] (03PS2) 10Hnowlan: cassandra: remove disabled CQL check [puppet] - 10https://gerrit.wikimedia.org/r/1307167 (https://phabricator.wikimedia.org/T407120) [16:52:40] RESOLVED: [2x] BFDdown: BFD session down between asw1-b4-magru and 195.200.68.37 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=asw1-b4-magru:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [16:52:55] restarted pybal on A:lvs-secondary-eqiad [16:53:35] oh I should log that properly I guess [16:53:57] !log rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-eqiad' 'systemctl restart pybal.service' # lvs1020, T416623 [16:53:59] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [16:54:00] T416623: Decommission NodeJS IPoid service - https://phabricator.wikimedia.org/T416623 [16:54:36] (03PS2) 10Btullis: Switch new an-test-master and an-test-coord nodes into their role [puppet] - 10https://gerrit.wikimedia.org/r/1308149 (https://phabricator.wikimedia.org/T406746) [16:55:12] !log rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-eqiad' 'systemctl restart pybal.service' # lvs1019, T416623 [16:55:14] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [16:55:26] (03PS1) 10Dzahn: Revert "jenkins: add firewall rule for controller to agent" [puppet] - 10https://gerrit.wikimedia.org/r/1308162 [16:55:39] waiting 300 seconds, which means I'll be a few minutes late for the end of the window :( sorry Raine [16:57:59] rzl: I don't need today's window, I'll be needing the one tomorrow :-) [16:58:20] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS bookworm [16:58:31] !log andrew@cumin2002 START - Cookbook sre.hosts.reboot-single for host cloudrabbit1001.eqiad.wmnet [16:58:48] OH it does say July 8. that's tomorrow! [16:59:08] great news, thanks [16:59:40] in that case, swfrench-wmf as the other usual suspect: I'll be ending this window a little late, creeping into the infra window -- sorry about that [17:00:05] Deploy window MediaWiki infrastructure (UTC late) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260707T1700) [17:00:56] !log rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal.service' # lvs2014, T416623 [17:00:58] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [17:00:59] T416623: Decommission NodeJS IPoid service - https://phabricator.wikimedia.org/T416623 [17:01:38] (03PS1) 10Hashar: Add ci::agent profile to the new Jenkins controller [puppet] - 10https://gerrit.wikimedia.org/r/1308163 [17:02:08] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1111.eqiad.wmnet with OS trixie [17:04:34] !log andrew@cumin2002 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1001.eqiad.wmnet [17:05:06] (03PS1) 10Dzahn: jenkins: include profile for jenkins agent and user setup [puppet] - 10https://gerrit.wikimedia.org/r/1308164 (https://phabricator.wikimedia.org/T418521) [17:05:45] (03CR) 10Dzahn: [V:03+2 C:03+2] jenkins: include profile for jenkins agent and user setup [puppet] - 10https://gerrit.wikimedia.org/r/1308164 (https://phabricator.wikimedia.org/T418521) (owner: 10Dzahn) [17:06:13] !log rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal.service' # lvs2013, T416623 [17:06:16] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [17:06:16] T416623: Decommission NodeJS IPoid service - https://phabricator.wikimedia.org/T416623 [17:06:17] (03Abandoned) 10Hashar: Add ci::agent profile to the new Jenkins controller [puppet] - 10https://gerrit.wikimedia.org/r/1308163 (owner: 10Hashar) [17:06:42] !log andrew@cumin2002 START - Cookbook sre.hosts.reboot-single for host cloudrabbit1002.eqiad.wmnet [17:07:16] (03PS1) 10Blake: mw-pretrain: Remove default quota and ranges. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308161 (https://phabricator.wikimedia.org/T427668) [17:07:46] rzl: apologies, missed your mention before. I don't plan to start etcd-related prep work for T430909 until the second half of the window. all yours! [17:07:47] T430909: codfw: rack B3 maintenance - https://phabricator.wikimedia.org/T430909 [17:07:55] sweet, thanks [17:09:56] (03PS1) 10Dzahn: jenkins: add Hiera keys for jenkins agent [puppet] - 10https://gerrit.wikimedia.org/r/1308165 (https://phabricator.wikimedia.org/T418521) [17:10:27] (03CR) 10Dzahn: [V:03+2 C:03+2] jenkins: add Hiera keys for jenkins agent [puppet] - 10https://gerrit.wikimedia.org/r/1308165 (https://phabricator.wikimedia.org/T418521) (owner: 10Dzahn) [17:11:01] (03CR) 10RLazarus: [C:03+2] service: Remove ipoid service [puppet] - 10https://gerrit.wikimedia.org/r/1306900 (https://phabricator.wikimedia.org/T416623) (owner: 10Clément Goubert) [17:12:04] (03CR) 10Ssingh: [C:03+1] "@bcornwall@wikimedia.org: can you check this once please? It looks good to me but needs your input." [puppet] - 10https://gerrit.wikimedia.org/r/1215329 (owner: 10Dzahn) [17:12:53] (03CR) 10CI reject: [V:04-1] swift: Load the file async [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1308156 (https://phabricator.wikimedia.org/T333445) (owner: 10Ladsgroup) [17:12:53] (03CR) 10CI reject: [V:04-1] bgp: Remove port from instance and remote_instance names. [alerts] - 10https://gerrit.wikimedia.org/r/1307787 (https://phabricator.wikimedia.org/T430290) (owner: 10Blake) [17:14:48] !log andrew@cumin2002 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1002.eqiad.wmnet [17:14:59] hm, gerrit down? [17:15:11] looks it [17:15:14] same [17:15:24] !log andrew@cumin2002 START - Cookbook sre.hosts.reboot-single for host cloudrabbit1003.eqiad.wmnet [17:15:27] mutante: just in case related to your work earlier [17:15:46] FIRING: [4x] GerritHAProxyBackendUnavailable: Gerrit backend is unavilable for tcp-proxy (HAProxy) gerrit_ssh - https://wikitech.wikimedia.org/wiki/Gerrit/Operations#GerritHAProxyBackendUnavailable - grafana.wikimedia.org/d/459365f6-df37-48d6-8142-82b22c1875e7/gerrit-tcp-proxy?viewPanel=panel-15 - https://alerts.wikimedia.org/?q=alertname%3DGerritHAProxyBackendUnavailable [17:15:57] FIRING: [2x] GerritHAProxyServiceUnavailable: Gerrit tcp-proxy (HAProxy) service gerrit_ssh is DOWN in codfw - https://wikitech.wikimedia.org/wiki/Gerrit/Operations#GerritHAProxyServiceUnavailable - grafana.wikimedia.org/d/459365f6-df37-48d6-8142-82b22c1875e7/gerrit-tcp-proxy?viewPanel=panel-15 - https://alerts.wikimedia.org/?q=alertname%3DGerritHAProxyServiceUnavailable [17:16:31] FIRING: [4x] ProbeDown: Service gerrit2003:29418 has failed probes (tcp_gerrit_ssh_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [17:16:56] checking to make sure I didn't just cause it via touching pybal, but I'd expect gerrit to be on high-traffic [17:17:21] rzl: yeah unrelated [17:17:34] I mean gerrit is ht1 and we didn't do anything with those bits [17:17:44] 👍 [17:18:25] (03PS1) 10Dzahn: jenkins: add contint1003 to jenkins_controller_hosts [puppet] - 10https://gerrit.wikimedia.org/r/1308166 (https://phabricator.wikimedia.org/T418521) [17:18:47] works for mutante, clearly :D [17:18:58] ah yep, back up for me too [17:19:13] (03CR) 10Dzahn: [V:03+2 C:03+2] jenkins: add contint1003 to jenkins_controller_hosts [puppet] - 10https://gerrit.wikimedia.org/r/1308166 (https://phabricator.wikimedia.org/T418521) (owner: 10Dzahn) [17:19:38] (03CR) 10Hnowlan: "Unfortunately not easily when it comes to swift logic :( Shipping the change to staging only via values-staging.yaml is the only real prac" [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1308156 (https://phabricator.wikimedia.org/T333445) (owner: 10Ladsgroup) [17:19:51] rzl: yea, it works. it's not related to earlier work, still fixing something but it's not this [17:20:06] cool :) sorry for the false alarm [17:20:46] RESOLVED: [4x] GerritHAProxyBackendUnavailable: Gerrit backend is unavilable for tcp-proxy (HAProxy) gerrit_ssh - https://wikitech.wikimedia.org/wiki/Gerrit/Operations#GerritHAProxyBackendUnavailable - grafana.wikimedia.org/d/459365f6-df37-48d6-8142-82b22c1875e7/gerrit-tcp-proxy?viewPanel=panel-15 - https://alerts.wikimedia.org/?q=alertname%3DGerritHAProxyBackendUnavailable [17:20:51] RESOLVED: [2x] GerritHAProxyServiceUnavailable: Gerrit tcp-proxy (HAProxy) service gerrit_ssh is DOWN in codfw - https://wikitech.wikimedia.org/wiki/Gerrit/Operations#GerritHAProxyServiceUnavailable - grafana.wikimedia.org/d/459365f6-df37-48d6-8142-82b22c1875e7/gerrit-tcp-proxy?viewPanel=panel-15 - https://alerts.wikimedia.org/?q=alertname%3DGerritHAProxyServiceUnavailable [17:21:31] RESOLVED: [4x] ProbeDown: Service gerrit2003:29418 has failed probes (tcp_gerrit_ssh_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [17:21:51] !log andrew@cumin2002 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1003.eqiad.wmnet [17:21:57] (03PS5) 10Clément Goubert: service: Remove ipoid service [puppet] - 10https://gerrit.wikimedia.org/r/1306900 (https://phabricator.wikimedia.org/T416623) [17:23:06] (03CR) 10RLazarus: [C:03+2] service: Remove ipoid service [puppet] - 10https://gerrit.wikimedia.org/r/1306900 (https://phabricator.wikimedia.org/T416623) (owner: 10Clément Goubert) [17:23:18] (03CR) 10Blake: "recheck" [alerts] - 10https://gerrit.wikimedia.org/r/1307787 (https://phabricator.wikimedia.org/T430290) (owner: 10Blake) [17:23:35] (03PS2) 10Ladsgroup: swift: Load the file async [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1308156 (https://phabricator.wikimedia.org/T333445) [17:23:43] (03CR) 10CI reject: [V:04-1] swift: Load the file async [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1308156 (https://phabricator.wikimedia.org/T333445) (owner: 10Ladsgroup) [17:23:43] (03CR) 10Ladsgroup: swift: Load the file async (031 comment) [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1308156 (https://phabricator.wikimedia.org/T333445) (owner: 10Ladsgroup) [17:26:10] (03CR) 10Ladsgroup: "If there are tests on staging we can run. That's still good enough. The change isn't that big and it either works or explodes so it should" [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1308156 (https://phabricator.wikimedia.org/T333445) (owner: 10Ladsgroup) [17:27:08] (03CR) 10Ladsgroup: "recheck" [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1308156 (https://phabricator.wikimedia.org/T333445) (owner: 10Ladsgroup) [17:28:53] LVS part of the ipoid removal work is done, I'll follow up later with the Kubernetes part [17:29:22] might just go wild and do it in an empty slot on the deployment calendar, ahead of Thursday :) just an ordinary service deploy at this point [17:30:20] Thanks for fitting it in beyond the window [17:31:12] (03PS7) 10Clément Goubert: deployment_server: absent ipoid kubernetes service [puppet] - 10https://gerrit.wikimedia.org/r/1306779 (https://phabricator.wikimedia.org/T416623) (owner: 10Dreamy Jazz) [17:31:12] (03PS8) 10Clément Goubert: deployment_server: remove ipoid users [puppet] - 10https://gerrit.wikimedia.org/r/1306782 (https://phabricator.wikimedia.org/T416623) (owner: 10Dreamy Jazz) [17:32:11] I presume it's still safe for the DBAs to drop the ipoid DB on m5? [17:32:28] I guess it was safe from when I removed the service from prod [17:32:34] rzl: I'm back to get started on etcd prep for the rack maintenance. how are things going? [17:33:06] (If you'd prefer for that to wait, I can update the task I filed for the DBAs) [17:34:51] ah, I see the LVS work is done, which is good, since we'll also be messing with LVS [17:37:07] (03CR) 10Ssingh: [C:03+2] hieradata: Temporarily point codfw PyBals at eqiad etcd [puppet] - 10https://gerrit.wikimedia.org/r/1308115 (https://phabricator.wikimedia.org/T430909) (owner: 10Scott French) [17:40:18] (03CR) 10Scott French: [C:03+2] wmnet: Temporarily direct codfw, eqsin, ulsfo etcd clients to eqiad [dns] - 10https://gerrit.wikimedia.org/r/1308114 (https://phabricator.wikimedia.org/T430909) (owner: 10Scott French) [17:40:48] !log restart pybal on lvs2014 to switch from conf2004 to conf1008: T430909 [17:40:50] !log swfrench@dns1004 START - running authdns-update [17:40:50] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [17:40:51] T430909: codfw: rack B3 maintenance - https://phabricator.wikimedia.org/T430909 [17:43:03] !log swfrench@dns1004 END - running authdns-update [17:44:03] (03PS5) 10Cwhite: prometheus: refactor prometheus-es-exporter to use config file [puppet] - 10https://gerrit.wikimedia.org/r/1305718 (https://phabricator.wikimedia.org/T350516) [17:44:05] !log switched codfw, eqsin, ulsfo etcd client SRV records to eqiad - T430909 [17:44:08] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [17:44:55] (03PS1) 10Dzahn: jenkins: ensure workspaces_dir exists [puppet] - 10https://gerrit.wikimedia.org/r/1308169 (https://phabricator.wikimedia.org/T418521) [17:45:06] (03CR) 10Dzahn: [C:03+2] Revert "jenkins: add firewall rule for controller to agent" [puppet] - 10https://gerrit.wikimedia.org/r/1308162 (owner: 10Dzahn) [17:45:59] (03CR) 10CI reject: [V:04-1] jenkins: ensure workspaces_dir exists [puppet] - 10https://gerrit.wikimedia.org/r/1308169 (https://phabricator.wikimedia.org/T418521) (owner: 10Dzahn) [17:46:15] (03CR) 10Dzahn: [V:03+2 C:03+2] jenkins: ensure workspaces_dir exists [puppet] - 10https://gerrit.wikimedia.org/r/1308169 (https://phabricator.wikimedia.org/T418521) (owner: 10Dzahn) [17:46:37] !log restart pybal on lvs2013 to switch from conf2004 to conf1008: T430909 [17:46:42] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [17:46:43] T430909: codfw: rack B3 maintenance - https://phabricator.wikimedia.org/T430909 [17:46:48] (03CR) 10Ladsgroup: "recheck" [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1308156 (https://phabricator.wikimedia.org/T333445) (owner: 10Ladsgroup) [17:51:00] !log restart pybal on lvs2012 to switch from conf2004 to conf1008 [puppet re-enabled there]: T430909 [17:51:02] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [17:51:07] (03PS1) 10Dzahn: jenkins: use literal qualified path for workspace dir [puppet] - 10https://gerrit.wikimedia.org/r/1308170 (https://phabricator.wikimedia.org/T418521) [17:51:42] (03CR) 10Dzahn: [V:03+2 C:03+2] jenkins: use literal qualified path for workspace dir [puppet] - 10https://gerrit.wikimedia.org/r/1308170 (https://phabricator.wikimedia.org/T418521) (owner: 10Dzahn) [17:52:40] !log restart pybal on lvs2011 to switch from conf2004 to conf1008: T430909 [17:52:43] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [17:52:44] T430909: codfw: rack B3 maintenance - https://phabricator.wikimedia.org/T430909 [17:53:44] (03PS2) 10Dzahn: jenkins: use literal qualified path for workspace dir [puppet] - 10https://gerrit.wikimedia.org/r/1308170 (https://phabricator.wikimedia.org/T418521) [17:54:20] (03CR) 10Dzahn: [V:03+2 C:03+2] jenkins: use literal qualified path for workspace dir [puppet] - 10https://gerrit.wikimedia.org/r/1308170 (https://phabricator.wikimedia.org/T418521) (owner: 10Dzahn) [17:54:56] (03Abandoned) 10Dzahn: jenkins: ensure workspaces_dir exists [puppet] - 10https://gerrit.wikimedia.org/r/1308169 (https://phabricator.wikimedia.org/T418521) (owner: 10Dzahn) [17:59:14] !log restarted ulsfo confds, confirmed now connected to eqiad backends - T430909 [17:59:17] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [17:59:19] T430909: codfw: rack B3 maintenance - https://phabricator.wikimedia.org/T430909 [18:00:05] thcipriani and thcipriani: #bothumor Q:Why did functions stop calling each other? A:They had arguments. Rise for MediaWiki train - Utc-7 Version . (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260707T1800). [18:01:13] !log restarted navtiming on webperf2003 - T430909 [18:01:16] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [18:01:24] (03CR) 10Cwhite: "PCC OK: https://puppet-compiler.wmflabs.org/output/1305718/8943/" [puppet] - 10https://gerrit.wikimedia.org/r/1305718 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [18:03:15] swfrench-wmf: ah sorry, I popped afk briefly! the answer was go for it, I'm glad you did :) [18:05:17] no worries at all [18:06:56] (03PS1) 10Bking: cirrussearch: move cloudelastic to its own file [alerts] - 10https://gerrit.wikimedia.org/r/1308176 (https://phabricator.wikimedia.org/T430869) [18:11:35] (03CR) 10CI reject: [V:04-1] cirrussearch: move cloudelastic to its own file [alerts] - 10https://gerrit.wikimedia.org/r/1308176 (https://phabricator.wikimedia.org/T430869) (owner: 10Bking) [18:11:53] !log restarted eqsin, codfw confds - T430909 [18:11:56] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [18:11:56] T430909: codfw: rack B3 maintenance - https://phabricator.wikimedia.org/T430909 [18:12:45] (03PS1) 10JHathaway: nftables: add support for the VRRP protocol [puppet] - 10https://gerrit.wikimedia.org/r/1308177 (https://phabricator.wikimedia.org/T427799) [18:12:55] PROBLEM - NTP peers and stratum check on dns7002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/NTP [18:12:59] (03CR) 10Ladsgroup: "recheck" [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1308156 (https://phabricator.wikimedia.org/T333445) (owner: 10Ladsgroup) [18:13:30] (03PS2) 10JHathaway: nftables: add support for the VRRP protocol [puppet] - 10https://gerrit.wikimedia.org/r/1308177 (https://phabricator.wikimedia.org/T427799) [18:13:41] (03CR) 10Bking: "recheck" [alerts] - 10https://gerrit.wikimedia.org/r/1308176 (https://phabricator.wikimedia.org/T430869) (owner: 10Bking) [18:14:03] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308177 (https://phabricator.wikimedia.org/T427799) (owner: 10JHathaway) [18:15:00] (03CR) 10JHathaway: "This should work and I agree it is cumbersome, perhaps we should add support? First cut attempt here, 1308177" [puppet] - 10https://gerrit.wikimedia.org/r/1297093 (https://phabricator.wikimedia.org/T427799) (owner: 10Majavah) [18:19:13] (03PS3) 10Btullis: Switch new an-test-master and an-test-coord nodes into their role [puppet] - 10https://gerrit.wikimedia.org/r/1308149 (https://phabricator.wikimedia.org/T406746) [18:19:19] (03CR) 10Btullis: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308149 (https://phabricator.wikimedia.org/T406746) (owner: 10Btullis) [18:21:40] (03CR) 10CDobbins: [V:03+1] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/8944/co" [puppet] - 10https://gerrit.wikimedia.org/r/1308150 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [18:27:40] (03CR) 10Btullis: [C:03+1] cirrussearch: move cloudelastic to its own file [alerts] - 10https://gerrit.wikimedia.org/r/1308176 (https://phabricator.wikimedia.org/T430869) (owner: 10Bking) [18:33:38] (03CR) 10CDanis: [C:03+1] "Thanks for flagging! We're good, the data pipeline isn't active at the moment." [puppet] - 10https://gerrit.wikimedia.org/r/1308147 (https://phabricator.wikimedia.org/T423714) (owner: 10Kamila Součková) [18:38:49] (03PS1) 10Btullis: Add dummy keytabs for new an-test nodes. [labs/private] - 10https://gerrit.wikimedia.org/r/1308182 (https://phabricator.wikimedia.org/T406746) [18:38:50] !log swfrench@cumin2002 START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo (T430909) [18:38:53] T430909: codfw: rack B3 maintenance - https://phabricator.wikimedia.org/T430909 [18:39:08] (03CR) 10Btullis: [V:03+2 C:03+2] Add dummy keytabs for new an-test nodes. [labs/private] - 10https://gerrit.wikimedia.org/r/1308182 (https://phabricator.wikimedia.org/T406746) (owner: 10Btullis) [18:39:53] (03CR) 10Btullis: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308149 (https://phabricator.wikimedia.org/T406746) (owner: 10Btullis) [18:40:41] !log swfrench@cumin2002 END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo (T430909) [18:41:41] (03CR) 10Ladsgroup: "recheck" [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1308156 (https://phabricator.wikimedia.org/T333445) (owner: 10Ladsgroup) [18:44:27] (03PS1) 10Clare Ming: Move Test Kitchen config from CommonSettings.php [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308183 (https://phabricator.wikimedia.org/T431257) [18:45:29] (03CR) 10CI reject: [V:04-1] Move Test Kitchen config from CommonSettings.php [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308183 (https://phabricator.wikimedia.org/T431257) (owner: 10Clare Ming) [18:46:39] FIRING: CirrusSearchNodeIndexingNotIncreasing: Elasticsearch instance cirrussearch1092-production-search-eqiad is not indexing - https://wikitech.wikimedia.org/wiki/Search/Elasticsearch_Administration#Indexing_hung_and_not_making_progress - https://grafana.wikimedia.org/d/JLK3I_siz/elasticsearch-indexing?orgId=1&from=now-3d&to=now&viewPanel=57 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchNodeIndexingNotIncreasing [18:47:16] (03PS2) 10Clare Ming: Move Test Kitchen config from CommonSettings.php [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308183 (https://phabricator.wikimedia.org/T431257) [18:49:09] !log swfrench@cumin2002 START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin (T430909) [18:49:12] T430909: codfw: rack B3 maintenance - https://phabricator.wikimedia.org/T430909 [18:49:43] (03PS3) 10Clare Ming: Move Test Kitchen config from CommonSettings.php [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308183 (https://phabricator.wikimedia.org/T431257) [18:50:42] (03CR) 10Bking: [C:03+2] cirrussearch: move cloudelastic to its own file [alerts] - 10https://gerrit.wikimedia.org/r/1308176 (https://phabricator.wikimedia.org/T430869) (owner: 10Bking) [18:52:08] !log swfrench@cumin2002 END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin (T430909) [18:52:45] PROBLEM - Check if ntpsec.service has been restarted after /etc/ntpsec/ntp.conf was changed on dns7002 is CRITICAL: CRITICAL: Service ntpsec.service has not been restarted after /etc/ntpsec/ntp.conf was changed (gt 2h). https://wikitech.wikimedia.org/wiki/NTP%23Monitoring [18:53:07] cjd91: ^ let's restart ntpsec.service on dns7002 [18:54:04] (03CR) 10Clare Ming: [C:03+2] test-kitchen Chart: Set phabricator discovery URL as the default one [deployment-charts] - 10https://gerrit.wikimedia.org/r/1307797 (https://phabricator.wikimedia.org/T428984) (owner: 10Santiago Faci) [18:54:22] FIRING: CoreRouterInterfaceDown: Core router interface down - cr2-eqiad:et-1/1/0 () - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [18:55:10] (03CR) 10Ayounsi: "Circuit has a 79 ms latency so if we keep the logic it should be 790, which should be enough to bring it ahead of the colt link (800)" [homer/public] - 10https://gerrit.wikimedia.org/r/1308152 (https://phabricator.wikimedia.org/T424839) (owner: 10Cathal Mooney) [18:56:11] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host cirrussearch1092.eqiad.wmnet with OS trixie [18:56:41] (03Merged) 10jenkins-bot: test-kitchen Chart: Set phabricator discovery URL as the default one [deployment-charts] - 10https://gerrit.wikimedia.org/r/1307797 (https://phabricator.wikimedia.org/T428984) (owner: 10Santiago Faci) [18:58:38] (03CR) 10Ayounsi: [C:03+1] "just saw your task comment, lgtm !" [homer/public] - 10https://gerrit.wikimedia.org/r/1308152 (https://phabricator.wikimedia.org/T424839) (owner: 10Cathal Mooney) [18:58:43] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host cirrussearch1109.eqiad.wmnet with OS trixie [18:58:45] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host cirrussearch1123.eqiad.wmnet with OS trixie [18:59:05] RECOVERY - Check if ntpsec.service has been restarted after /etc/ntpsec/ntp.conf was changed on dns7002 is OK: OK: ntpsec.service was restarted after /etc/ntpsec/ntp.conf was changed. https://wikitech.wikimedia.org/wiki/NTP%23Monitoring [19:06:06] (03CR) 10Ladsgroup: "recheck" [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1308156 (https://phabricator.wikimedia.org/T333445) (owner: 10Ladsgroup) [19:06:12] RECOVERY - NTP peers and stratum check on dns7002 is OK: NTP OK: Offset 0.000314299 secs, stratum=2 https://wikitech.wikimedia.org/wiki/NTP [19:06:46] (03CR) 10Eevans: [C:03+1] Add depool policy for restbase/sessionstore/cassandra_dev [puppet] - 10https://gerrit.wikimedia.org/r/1308030 (https://phabricator.wikimedia.org/T327300) (owner: 10Ayounsi) [19:07:25] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [19:10:41] (03PS1) 10Ebenezer Rao: Fixed the typo of the word return in move-vlan.py [cookbooks] - 10https://gerrit.wikimedia.org/r/1308185 (https://phabricator.wikimedia.org/T201491) [19:10:58] !log bking@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage [19:11:20] !log bking@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage [19:12:05] 10ops-eqiad, 06SRE, 10Data-Persistence-Backup, 06DC-Ops: db1245 crashed - https://phabricator.wikimedia.org/T431115#12096694 (10VRiley-WMF) The TSR report showed the following and Dell has provided some troubleshooting tips which I will be walking through Post TSR analysis we can see multiple errors repo... [19:12:29] !log cdobbins@cumin2002 conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update [19:13:23] !log cdobbins@dns1004 START - running authdns-update [19:14:10] (03CR) 10Jasmine: [C:03+2] Add Kubernetes POD IP reverse range delegations for wikikube-ctrl1006 [dns] - 10https://gerrit.wikimedia.org/r/1307215 (https://phabricator.wikimedia.org/T418920) (owner: 10Jasmine) [19:14:25] (03CR) 10Jasmine: [C:03+2] templates/wmnet: Add new control plane wikikube-ctrl1006 to etcd-server SRV record [dns] - 10https://gerrit.wikimedia.org/r/1307216 (https://phabricator.wikimedia.org/T418920) (owner: 10Jasmine) [19:14:34] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage [19:15:00] !log cdobbins@dns1004 END - running authdns-update [19:15:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.1% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [19:15:25] !log jasmine@dns1004 START - running authdns-update [19:15:42] !log bking@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage [19:16:18] (03PS1) 10Dzahn: ci: add contint1003 to list of firewall_hosts allowed to git over 9418 [puppet] - 10https://gerrit.wikimedia.org/r/1308186 (https://phabricator.wikimedia.org/T418521) [19:17:06] (03CR) 10Dzahn: [C:03+2] ci: add contint1003 to list of firewall_hosts allowed to git over 9418 [puppet] - 10https://gerrit.wikimedia.org/r/1308186 (https://phabricator.wikimedia.org/T418521) (owner: 10Dzahn) [19:17:15] !log jasmine@dns1004 END - running authdns-update [19:17:57] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage [19:18:21] (03CR) 10Dzahn: [C:03+2] "18:47 < Amir1> now it tries to connect to contint2002.wikimedia.org and failing https://integration.wikimedia.org/ci/job/thumbor-plugins-p" [puppet] - 10https://gerrit.wikimedia.org/r/1308186 (https://phabricator.wikimedia.org/T418521) (owner: 10Dzahn) [19:19:17] !log cdobbins@cumin2002 conftool action : set/pooled=yes; selector: name=dns7002.* [19:20:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.1% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [19:22:14] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage [19:23:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.76% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [19:28:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.97% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [19:29:12] (03PS1) 10Eevans: preseed.yaml: Use new (EFI) Cassandra JBOD preseed [puppet] - 10https://gerrit.wikimedia.org/r/1308188 (https://phabricator.wikimedia.org/T416538) [19:32:41] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1123.eqiad.wmnet with OS trixie [19:38:47] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1092.eqiad.wmnet with OS trixie [19:44:33] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1109.eqiad.wmnet with OS trixie [19:45:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.83% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [19:48:18] (03CR) 10Ayounsi: [C:03+2] Add depool policy for restbase/sessionstore/cassandra_dev [puppet] - 10https://gerrit.wikimedia.org/r/1308030 (https://phabricator.wikimedia.org/T327300) (owner: 10Ayounsi) [19:50:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.83% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [19:52:02] (03PS1) 10Dzahn: ci::docker: ensure jenkins agent user is in docker group also in prod [puppet] - 10https://gerrit.wikimedia.org/r/1308194 (https://phabricator.wikimedia.org/T418521) [19:52:18] (03CR) 10JHathaway: [C:03+1] "looks good, minor suggestions" [cookbooks] - 10https://gerrit.wikimedia.org/r/1302859 (https://phabricator.wikimedia.org/T426180) (owner: 10Elukey) [19:52:43] (03CR) 10CI reject: [V:04-1] ci::docker: ensure jenkins agent user is in docker group also in prod [puppet] - 10https://gerrit.wikimedia.org/r/1308194 (https://phabricator.wikimedia.org/T418521) (owner: 10Dzahn) [19:53:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.76% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [19:54:12] (03PS2) 10Dzahn: ci::docker: ensure jenkins agent user is in docker group also in prod [puppet] - 10https://gerrit.wikimedia.org/r/1308194 (https://phabricator.wikimedia.org/T418521) [19:57:02] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie [19:58:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.69% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [19:58:43] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie [19:59:09] (03CR) 10Dzahn: "recheck" [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1308156 (https://phabricator.wikimedia.org/T333445) (owner: 10Ladsgroup) [20:00:05] RoanKattouw, urbanecm, TheresNoTime, kindrobot, and cjming: #bothumor I � Unicode. All rise for UTC late backport window deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260707T2000). [20:00:05] No Gerrit patches in the queue for this window AFAICS. [20:01:23] !log bking@cumin2003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=93) for host cirrussearch1108.eqiad.wmnet with OS trixie [20:03:53] (03PS1) 10Arlolra: Bump Parsoid image limit to 1250 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308196 (https://phabricator.wikimedia.org/T430854) [20:04:22] (03CR) 10Scott French: [C:03+1] mw-pretrain: Remove default quota and ranges. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308161 (https://phabricator.wikimedia.org/T427668) (owner: 10Blake) [20:07:25] 10ops-eqiad, 06SRE, 06DC-Ops: Degraded RAID on backup1019 - https://phabricator.wikimedia.org/T431367#12096933 (10VRiley-WMF) Created Dell ticket 228753175 for this hard drive issue. They are shipping one out. [20:09:21] !log remove 2026-04 swift log archives from centrallog2002 to free some space [20:09:22] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [20:10:32] (03CR) 10Subramanya Sastry: [C:03+1] Bump Parsoid image limit to 1250 (031 comment) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308196 (https://phabricator.wikimedia.org/T430854) (owner: 10Arlolra) [20:11:15] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, July 07 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308196 (https://phabricator.wikimedia.org/T430854) (owner: 10Arlolra) [20:11:38] If noone else is deploying, I will backport a config change [20:12:45] (03CR) 10Cwhite: [C:03+2] logstash: configure provisioning of admin certificate [puppet] - 10https://gerrit.wikimedia.org/r/1306001 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [20:13:20] (03CR) 10TrainBranchBot: [C:03+2] "Approved by arlolra@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308196 (https://phabricator.wikimedia.org/T430854) (owner: 10Arlolra) [20:13:36] (03PS4) 10Cwhite: opensearch: add config required to authenticate curator requests [puppet] - 10https://gerrit.wikimedia.org/r/1306969 (https://phabricator.wikimedia.org/T350516) [20:13:49] (03CR) 10Scott French: [C:03+1] "Awesome - thank you, Chris!" [puppet] - 10https://gerrit.wikimedia.org/r/1308147 (https://phabricator.wikimedia.org/T423714) (owner: 10Kamila Součková) [20:14:19] (03Merged) 10jenkins-bot: Bump Parsoid image limit to 1250 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308196 (https://phabricator.wikimedia.org/T430854) (owner: 10Arlolra) [20:14:45] !log arlolra@deploy1003 Started scap sync-world: Backport for [[gerrit:1308196|Bump Parsoid image limit to 1250 (T430854)]] [20:14:48] T430854: Improve handling of image limits by fixing the UX on pages where these limits are hit - https://phabricator.wikimedia.org/T430854 [20:14:55] (03CR) 10Cwhite: [C:03+2] opensearch: add config required to authenticate curator requests [puppet] - 10https://gerrit.wikimedia.org/r/1306969 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [20:15:10] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie [20:16:02] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie [20:16:02] (03PS8) 10Cwhite: logstash: add parameters needed for security plugin [puppet] - 10https://gerrit.wikimedia.org/r/1305769 (https://phabricator.wikimedia.org/T350516) [20:16:49] !log arlolra@deploy1003 arlolra: Backport for [[gerrit:1308196|Bump Parsoid image limit to 1250 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [20:17:53] !log arlolra@deploy1003 arlolra: Continuing with deployment [20:18:02] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host cirrussearch1091.eqiad.wmnet with OS trixie [20:19:07] (03PS2) 10CDobbins: hieradata: issue fixes [puppet] - 10https://gerrit.wikimedia.org/r/1308150 (https://phabricator.wikimedia.org/T401832) [20:20:41] !log "homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1006"" [20:20:43] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [20:20:43] (03CR) 10CDobbins: [V:03+1] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/8945/co" [puppet] - 10https://gerrit.wikimedia.org/r/1308150 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [20:22:14] !log arlolra@deploy1003 Finished scap sync-world: Backport for [[gerrit:1308196|Bump Parsoid image limit to 1250 (T430854)]] (duration: 07m 29s) [20:22:17] T430854: Improve handling of image limits by fixing the UX on pages where these limits are hit - https://phabricator.wikimedia.org/T430854 [20:26:12] 06SRE, 06ServiceOps new, 07Service-deployment-requests: New Service Request: Headless VE - https://phabricator.wikimedia.org/T431497 (10Ottomata) 03NEW [20:26:14] (03CR) 10Scott French: [C:03+1] "Thanks for the detailed write-up!" [puppet] - 10https://gerrit.wikimedia.org/r/1306905 (https://phabricator.wikimedia.org/T428909) (owner: 10Clément Goubert) [20:26:42] 06SRE, 06ServiceOps new, 07Service-deployment-requests: New Service Request: Headless VE - https://phabricator.wikimedia.org/T431497#12097085 (10Ottomata) [20:27:02] !log "homer lsw1-c2-eqiad* commit "Added new stacked control plane wikikube-ctrl1006"" [20:27:03] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [20:27:24] 06SRE, 06ServiceOps new, 07Service-deployment-requests: New Service Request: Headless VE - https://phabricator.wikimedia.org/T431497#12097090 (10Ottomata) [20:29:12] 10ops-eqiad, 06SRE, 06DC-Ops: Degraded RAID on backup1019 - https://phabricator.wikimedia.org/T431367#12097102 (10VRiley-WMF) a:03VRiley-WMF [20:30:18] !log jasmine@cumin2002 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1006.eqiad.wmnet [20:30:18] !log bking@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage [20:30:21] !log jasmine@cumin2002 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1006.eqiad.wmnet [20:30:32] 06SRE, 06ServiceOps new, 10ServiceOps-Upgrades-Hardware: wikikube-ctrl100[56] implementation tracking - https://phabricator.wikimedia.org/T418920#12097103 (10ops-monitoring-bot) Cookbook cookbooks.sre.k8s.pool-depool-node started by jasmine@cumin2002 pool for host wikikube-ctrl1006.eqiad.wmnet completed: - w... [20:30:50] !log bking@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage [20:32:31] FIRING: [2x] ProbeDown: Service gerrit2003:443 has failed probes (http_gerrit_tls_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#gerrit2003:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [20:32:46] FIRING: [2x] GerritHAProxyBackendUnavailable: Gerrit backend is unavilable for tcp-proxy (HAProxy) gerrit_ssh - https://wikitech.wikimedia.org/wiki/Gerrit/Operations#GerritHAProxyBackendUnavailable - grafana.wikimedia.org/d/459365f6-df37-48d6-8142-82b22c1875e7/gerrit-tcp-proxy?viewPanel=panel-15 - https://alerts.wikimedia.org/?q=alertname%3DGerritHAProxyBackendUnavailable [20:33:17] !log bking@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage [20:33:35] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage [20:33:46] FIRING: [2x] GerritHAProxyServiceUnavailable: Gerrit tcp-proxy (HAProxy) service gerrit_ssh is DOWN in codfw - https://wikitech.wikimedia.org/wiki/Gerrit/Operations#GerritHAProxyServiceUnavailable - grafana.wikimedia.org/d/459365f6-df37-48d6-8142-82b22c1875e7/gerrit-tcp-proxy?viewPanel=panel-15 - https://alerts.wikimedia.org/?q=alertname%3DGerritHAProxyServiceUnavailable [20:34:26] Gerrit appears down [20:36:39] Now back for me [20:36:52] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage [20:37:31] RESOLVED: [4x] ProbeDown: Service gerrit2003:29418 has failed probes (tcp_gerrit_ssh_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [20:37:46] RESOLVED: [4x] GerritHAProxyBackendUnavailable: Gerrit backend is unavilable for tcp-proxy (HAProxy) gerrit_ssh - https://wikitech.wikimedia.org/wiki/Gerrit/Operations#GerritHAProxyBackendUnavailable - grafana.wikimedia.org/d/459365f6-df37-48d6-8142-82b22c1875e7/gerrit-tcp-proxy?viewPanel=panel-15 - https://alerts.wikimedia.org/?q=alertname%3DGerritHAProxyBackendUnavailable [20:38:46] RESOLVED: [2x] GerritHAProxyServiceUnavailable: Gerrit tcp-proxy (HAProxy) service gerrit_ssh is DOWN in codfw - https://wikitech.wikimedia.org/wiki/Gerrit/Operations#GerritHAProxyServiceUnavailable - grafana.wikimedia.org/d/459365f6-df37-48d6-8142-82b22c1875e7/gerrit-tcp-proxy?viewPanel=panel-15 - https://alerts.wikimedia.org/?q=alertname%3DGerritHAProxyServiceUnavailable [20:40:45] (03CR) 10CDobbins: [V:03+1] "The reason for putting quotes around all IPs is that `/` and `:` may cause problems otherwise: https://doc.powerdns.com/recursor/yamlsett" [puppet] - 10https://gerrit.wikimedia.org/r/1308150 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [20:43:56] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage [20:47:14] 06SRE, 06ServiceOps new, 10ServiceOps-Upgrades-Hardware: wikikube-ctrl100[56] implementation tracking - https://phabricator.wikimedia.org/T418920#12097176 (10jasmine_) 05Open→03Resolved Alright, sweet `wikikube-ctrl1006` has been added to the cluster now too🎉 ` wikikube-ctrl1002.eqiad.wmnet... [20:48:46] FIRING: GerritDiskSpaceExhaustionIncoming: Gerrit disk space runway on gerrit2003:/srv is too low - https://wikitech.wikimedia.org/wiki/Gerrit/Operations#TODO - https://grafana.wikimedia.org/d/rYdddlPWk/node-exporter-collaboration-services?var-job=node&var-nodename=gerrit2003&var-node=gerrit2003%3A9100&refresh=1m - https://alerts.wikimedia.org/?q=alertname%3DGerritDiskSpaceExhaustionIncoming [20:48:47] 10ops-eqiad, 06SRE, 06cloud-services-team, 10Cloud-VPS, 06DC-Ops: New thermal paste for cloudvirt1071.eqiad.wmnet - https://phabricator.wikimedia.org/T431429#12097198 (10VRiley-WMF) a:03VRiley-WMF [20:54:47] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1090.eqiad.wmnet with OS trixie [20:58:13] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1091.eqiad.wmnet with OS trixie [21:00:05] Deploy window Readers deployment window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260707T2100) [21:04:04] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, 06cloud-services-team (FY2025/2026-Q3-Q4): Ensure cloudvirt capacity is more evenly spread out among racks - https://phabricator.wikimedia.org/T424658#12097230 (10VRiley-WMF) a:03VRiley-WMF [21:05:13] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1108.eqiad.wmnet with OS trixie [21:11:33] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply [21:12:02] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply [21:12:46] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply [21:13:57] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply [21:21:09] 06SRE, 06ServiceOps new, 10ServiceOps-Upgrades-Hardware, 07Kubernetes, 13Patch-For-Review: wikikube-ctrl2006 implementation tracking - https://phabricator.wikimedia.org/T406596#12097262 (10jasmine_) To recap the IRC discussion from earlier, it appears the PXE was enable on both 10g nic's and disabling o... [21:33:46] RESOLVED: GerritDiskSpaceExhaustionIncoming: Gerrit disk space runway on gerrit2003:/srv is too low - https://wikitech.wikimedia.org/wiki/Gerrit/Operations#TODO - https://grafana.wikimedia.org/d/rYdddlPWk/node-exporter-collaboration-services?var-job=node&var-nodename=gerrit2003&var-node=gerrit2003%3A9100&refresh=1m - https://alerts.wikimedia.org/?q=alertname%3DGerritDiskSpaceExhaustionIncoming [21:50:18] (03PS1) 10Hashar: gerrit: comment out HeapDumpOnOutOfMemoryError [puppet] - 10https://gerrit.wikimedia.org/r/1308213 [21:57:58] (03PS3) 10RLazarus: admin_ng: Restrict wikifunctions network access to DNS [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305799 (https://phabricator.wikimedia.org/T427864) [22:02:43] (03CR) 10Cwhite: [C:03+2] gerrit: comment out HeapDumpOnOutOfMemoryError [puppet] - 10https://gerrit.wikimedia.org/r/1308213 (owner: 10Hashar) [22:03:34] jouncebot: now [22:03:34] No deployments scheduled for the next 7 hour(s) and 56 minute(s) [22:07:50] !log Restarting Gerrit on gerrit2003 [22:07:51] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [22:08:12] Answers my question about gerrit :D [22:08:23] sorry :-\ [22:08:50] Np, needed to get up from the PC anyway :D [22:09:13] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host cirrussearch1097.eqiad.wmnet with OS trixie [22:11:01] the good thing is Gerrit comes back reasonably fast [22:12:18] Yeah definitely :D [22:14:36] !log Restarting Gerrit on gerrit2002 and gerrit1003 (replicas) [22:14:37] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [22:18:13] (03PS1) 10Andrew Bogott: Horizon: update codfw1dev build version [puppet] - 10https://gerrit.wikimedia.org/r/1308226 (https://phabricator.wikimedia.org/T430164) [22:20:18] (03CR) 10Andrew Bogott: [C:03+2] Horizon: update codfw1dev build version [puppet] - 10https://gerrit.wikimedia.org/r/1308226 (https://phabricator.wikimedia.org/T430164) (owner: 10Andrew Bogott) [22:21:44] (03PS1) 10Jdlrobson: Improve responsiveness of toolbelt [skins/Vector] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1308227 (https://phabricator.wikimedia.org/T429518) [22:22:06] (03CR) 10RLazarus: "recheck" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305799 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [22:24:16] !log bking@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage [22:28:59] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage [22:38:19] (03CR) 10Santiago Faci: [C:03+1] Move Test Kitchen config from CommonSettings.php [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308183 (https://phabricator.wikimedia.org/T431257) (owner: 10Clare Ming) [22:41:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 17.18% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [22:42:20] (03CR) 10RLazarus: admin_ng: Restrict wikifunctions network access to DNS (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305799 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [22:43:08] (03CR) 10BCornwall: [V:03+1 C:03+1] "I can second what Daniel said: No hits in the last 30 days for that referer regex. I'll deploy this tomorrow." [puppet] - 10https://gerrit.wikimedia.org/r/1215329 (owner: 10Dzahn) [22:46:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.45% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [22:49:47] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1097.eqiad.wmnet with OS trixie [22:54:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 18.87% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [22:55:09] (03PS9) 10Cwhite: opensearch_dashboards: add settings required to enable the security plugin [puppet] - 10https://gerrit.wikimedia.org/r/1305774 (https://phabricator.wikimedia.org/T350516) [22:55:09] (03CR) 10Cwhite: "PCC OK: https://puppet-compiler.wmflabs.org/output/1305774/8947/" [puppet] - 10https://gerrit.wikimedia.org/r/1305774 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [22:55:25] FIRING: CoreRouterInterfaceDown: Core router interface down - cr2-eqiad:et-1/1/0 () - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [22:55:55] (03CR) 10Cwhite: [C:03+1] cassandra: remove disabled CQL check [puppet] - 10https://gerrit.wikimedia.org/r/1307167 (https://phabricator.wikimedia.org/T407120) (owner: 10Hnowlan) [23:04:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.72% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [23:07:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.42% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [23:07:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [23:12:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.24% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [23:30:44] (03PS3) 10Ladsgroup: swift: Load the file async [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1308156 (https://phabricator.wikimedia.org/T333445) [23:35:52] 14SRE-Sprint-Week-Sustainability-March2023, 06Traffic, 13Patch-For-Review, 07Sustainability (Incident Followup): Experiment with single backend CDN nodes - https://phabricator.wikimedia.org/T288106#12097579 (10BCornwall) It's been a week since running and the results can be viewed: [[ https://grafana-rw.w... [23:42:32] (03PS1) 10TrainBranchBot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1308234 [23:42:32] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1308234 (owner: 10TrainBranchBot) [23:50:30] (03Merged) 10jenkins-bot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1308234 (owner: 10TrainBranchBot)