[00:05:38] FIRING: [12x] CertAlmostExpired: gNMI TLS certificate for lsw1-c2-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [00:28:36] FIRING: [5x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-3/1/2 (Transport: Hurricane Electric (dc4841.dal5) {#122601}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [00:46:46] (03PS2) 10Ladsgroup: thumbor: Allow 2 connection per worker [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309832 (https://phabricator.wikimedia.org/T431767) [01:12:04] (03PS1) 10TrainBranchBot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1309889 [01:12:04] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1309889 (owner: 10TrainBranchBot) [01:18:41] (03Merged) 10jenkins-bot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1309889 (owner: 10TrainBranchBot) [01:26:10] 06SRE, 10SRE-swift-storage, 06Commons: Some files had disappeared from Commons after renaming - https://phabricator.wikimedia.org/T111838#12113145 (10Pppery) [01:42:33] 14SRE-Sprint-Week-Sustainability-March2023, 06DBA, 07Sustainability (Incident Followup): Alert when auto-increment fields on any MW-related databases reach a threshold - https://phabricator.wikimedia.org/T291332#12113176 (10Krinkle) [01:50:57] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:00:14] !log mwpresync@deploy2003 Started scap build-images: Publishing wmf/next image [02:02:55] PROBLEM - SSH on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/SSH/monitoring [02:02:55] PROBLEM - librenms.wikimedia.org requires authentication on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [02:02:55] PROBLEM - librenms.wikimedia.org tls expiry on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [02:04:35] FIRING: HelmReleaseBadStatus: Helm release wdqs/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [02:05:23] FIRING: [13x] CertAlmostExpired: gNMI TLS certificate for lsw1-c2-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [02:05:42] FIRING: [2x] JobUnavailable: Reduced availability for job atlas_exporter in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:06:52] !log mwpresync@deploy2003 Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) [02:07:19] FIRING: KafkaUnderReplicatedPartitions: Under replicated partitions for Kafka cluster test-eqiad in eqiad - https://wikitech.wikimedia.org/wiki/Kafka/Administration - https://grafana.wikimedia.org/d/000000027/kafka?orgId=1&var-datasource=eqiad%20prometheus/ops&var-kafka_cluster=test-eqiad - https://alerts.wikimedia.org/?q=alertname%3DKafkaUnderReplicatedPartitions [02:10:42] FIRING: [4x] JobUnavailable: Reduced availability for job atlas_exporter in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:15:42] FIRING: [4x] JobUnavailable: Reduced availability for job atlas_exporter in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:16:45] RECOVERY - librenms.wikimedia.org tls expiry on netmon2002 is OK: OK - Certificate librenms.wikimedia.org will expire on Thu 10 Sep 2026 02:24:49 AM GMT +0000. https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [02:16:45] RECOVERY - SSH on netmon2002 is OK: SSH OK - OpenSSH_9.2p1 Debian-2+deb12u10 (protocol 2.0) https://wikitech.wikimedia.org/wiki/SSH/monitoring [02:16:45] RECOVERY - librenms.wikimedia.org requires authentication on netmon2002 is OK: HTTP OK: Status line output matched HTTP/1.1 302 - 701 bytes in 0.138 second response time https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [02:20:42] FIRING: [4x] JobUnavailable: Reduced availability for job atlas_exporter in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:38:55] RECOVERY - Kafka Broker Server on kafka-test1007 is OK: PROCS OK: 1 process with command name java, args Kafka /etc/kafka/server.properties https://wikitech.wikimedia.org/wiki/Kafka/Administration [02:41:55] PROBLEM - Kafka Broker Server on kafka-test1007 is CRITICAL: PROCS CRITICAL: 0 processes with command name java, args Kafka /etc/kafka/server.properties https://wikitech.wikimedia.org/wiki/Kafka/Administration [02:46:47] (03CR) 10Anzx: [C:04-1] "recheck" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309865 (https://phabricator.wikimedia.org/T431094) (owner: 10Priyankar22) [02:47:34] (03CR) 10CI reject: [V:04-1] thwiki: Change to Wikipedia 25 logo [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309865 (https://phabricator.wikimedia.org/T431094) (owner: 10Priyankar22) [03:08:55] RECOVERY - Kafka Broker Server on kafka-test1007 is OK: PROCS OK: 1 process with command name java, args Kafka /etc/kafka/server.properties https://wikitech.wikimedia.org/wiki/Kafka/Administration [03:11:55] PROBLEM - Kafka Broker Server on kafka-test1007 is CRITICAL: PROCS CRITICAL: 0 processes with command name java, args Kafka /etc/kafka/server.properties https://wikitech.wikimedia.org/wiki/Kafka/Administration [03:32:31] FIRING: [2x] Traffic bill over quota: Alert for device cr2-eqiad.wikimedia.org - Traffic bill over quota - https://alerts.wikimedia.org/?q=alertname%3DTraffic+bill+over+quota [03:37:31] FIRING: [3x] Traffic bill over quota: Alert for device cr1-codfw.wikimedia.org - Traffic bill over quota - https://alerts.wikimedia.org/?q=alertname%3DTraffic+bill+over+quota [03:52:31] FIRING: [3x] Traffic bill over quota: Alert for device cr1-codfw.wikimedia.org - Traffic bill over quota - https://alerts.wikimedia.org/?q=alertname%3DTraffic+bill+over+quota [03:57:31] RESOLVED: Traffic bill over quota: Alert for device cr1-codfw.wikimedia.org - Traffic bill over quota - https://alerts.wikimedia.org/?q=alertname%3DTraffic+bill+over+quota [04:07:55] RECOVERY - Kafka Broker Server on kafka-test1009 is OK: PROCS OK: 1 process with command name java, args Kafka /etc/kafka/server.properties https://wikitech.wikimedia.org/wiki/Kafka/Administration [04:10:55] PROBLEM - Kafka Broker Server on kafka-test1009 is CRITICAL: PROCS CRITICAL: 0 processes with command name java, args Kafka /etc/kafka/server.properties https://wikitech.wikimedia.org/wiki/Kafka/Administration [04:28:36] FIRING: [5x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-3/1/2 (Transport: Hurricane Electric (dc4841.dal5) {#122601}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [04:38:55] RECOVERY - Kafka Broker Server on kafka-test1007 is OK: PROCS OK: 1 process with command name java, args Kafka /etc/kafka/server.properties https://wikitech.wikimedia.org/wiki/Kafka/Administration [04:41:55] PROBLEM - Kafka Broker Server on kafka-test1007 is CRITICAL: PROCS CRITICAL: 0 processes with command name java, args Kafka /etc/kafka/server.properties https://wikitech.wikimedia.org/wiki/Kafka/Administration [05:08:55] RECOVERY - Kafka Broker Server on kafka-test1007 is OK: PROCS OK: 1 process with command name java, args Kafka /etc/kafka/server.properties https://wikitech.wikimedia.org/wiki/Kafka/Administration [05:11:00] !log Drop users_to_rename table T431842 [05:11:02] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [05:11:03] T431842: Drop users_to_rename table from production - https://phabricator.wikimedia.org/T431842 [05:11:55] PROBLEM - Kafka Broker Server on kafka-test1007 is CRITICAL: PROCS CRITICAL: 0 processes with command name java, args Kafka /etc/kafka/server.properties https://wikitech.wikimedia.org/wiki/Kafka/Administration [05:12:04] (03PS1) 10Marostegui: tables-catalog.yaml: Remove users_to_rename [puppet] - 10https://gerrit.wikimedia.org/r/1309893 (https://phabricator.wikimedia.org/T431842) [05:12:36] (03PS1) 10Gkm563: Remove nonexistent autopatrolled group from Outreach Wiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309894 [05:15:17] (03PS2) 10Gkm563: Remove nonexistent autopatrolled group from Outreach Wiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309894 (https://phabricator.wikimedia.org/T431959) [05:22:20] FIRING: CirrusSearchBackendMemoryIssue: CirrusSearch backend failed 55.470309444444446 times in the last 10 minutes due to memory usage (mw@codfw to dnsdisc) - https://wikitech.wikimedia.org/wiki/Search#Health/Activity_Monitoring - https://grafana.wikimedia.org/d/dc04b9f2-b8d5-4ab6-9482-5d9a75728951/elasticsearch-percentiles?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchBackendMemoryIssue [05:26:34] 10ops-eqiad, 06Data-Persistence, 06DBA, 06DC-Ops, and 3 others: es7 primary (es1039.eqiad.wmnet) crashed - https://phabricator.wikimedia.org/T430764#12113283 (10Marostegui) Thanks Valerie! Starting some more tests now. [05:27:21] RESOLVED: CirrusSearchBackendMemoryIssue: CirrusSearch backend failed 22.22222222222222 times in the last 10 minutes due to memory usage (mw@codfw to dnsdisc) - https://wikitech.wikimedia.org/wiki/Search#Health/Activity_Monitoring - https://grafana.wikimedia.org/d/dc04b9f2-b8d5-4ab6-9482-5d9a75728951/elasticsearch-percentiles?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchBackendMemoryIssue [05:33:20] FIRING: CirrusSearchBackendMemoryIssue: CirrusSearch backend failed 38.88888888888889 times in the last 10 minutes due to memory usage (mw@codfw to dnsdisc) - https://wikitech.wikimedia.org/wiki/Search#Health/Activity_Monitoring - https://grafana.wikimedia.org/d/dc04b9f2-b8d5-4ab6-9482-5d9a75728951/elasticsearch-percentiles?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchBackendMemoryIssue [05:37:59] !log marostegui@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1217,1228].eqiad.wmnet with reason: cloning [05:38:20] RESOLVED: CirrusSearchBackendMemoryIssue: CirrusSearch backend failed 20.000006172873796 times in the last 10 minutes due to memory usage (mw@codfw to dnsdisc) - https://wikitech.wikimedia.org/wiki/Search#Health/Activity_Monitoring - https://grafana.wikimedia.org/d/dc04b9f2-b8d5-4ab6-9482-5d9a75728951/elasticsearch-percentiles?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchBackendMemoryIssue [05:43:05] (03PS1) 10Marostegui: mariadb: Move db1228 to m5 [puppet] - 10https://gerrit.wikimedia.org/r/1309895 (https://phabricator.wikimedia.org/T430903) [05:45:21] FIRING: CirrusSearchBackendMemoryIssue: CirrusSearch backend failed 80.00000000000001 times in the last 10 minutes due to memory usage (mw@codfw to dnsdisc) - https://wikitech.wikimedia.org/wiki/Search#Health/Activity_Monitoring - https://grafana.wikimedia.org/d/dc04b9f2-b8d5-4ab6-9482-5d9a75728951/elasticsearch-percentiles?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchBackendMemoryIssue [05:46:38] (03CR) 10Marostegui: [C:03+2] mariadb: Move db1228 to m5 [puppet] - 10https://gerrit.wikimedia.org/r/1309895 (https://phabricator.wikimedia.org/T430903) (owner: 10Marostegui) [05:57:20] PROBLEM - haproxy failover on dbproxy1027 is CRITICAL: CRITICAL check_failover servers up 1 down 1: https://wikitech.wikimedia.org/wiki/HAProxy [05:57:20] PROBLEM - haproxy failover on dbproxy1029 is CRITICAL: CRITICAL check_failover servers up 1 down 1: https://wikitech.wikimedia.org/wiki/HAProxy [06:00:20] RESOLVED: CirrusSearchBackendMemoryIssue: CirrusSearch backend failed 38.888888888888886 times in the last 10 minutes due to memory usage (mw@codfw to dnsdisc) - https://wikitech.wikimedia.org/wiki/Search#Health/Activity_Monitoring - https://grafana.wikimedia.org/d/dc04b9f2-b8d5-4ab6-9482-5d9a75728951/elasticsearch-percentiles?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchBackendMemoryIssue [06:00:52] dbproxy alerts are expeted [06:03:22] !log marostegui@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on dbproxy[2005-2008].codfw.wmnet with reason: reboot [06:04:35] FIRING: HelmReleaseBadStatus: Helm release wdqs/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [06:05:38] FIRING: [13x] CertAlmostExpired: gNMI TLS certificate for lsw1-c2-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [06:07:19] FIRING: KafkaUnderReplicatedPartitions: Under replicated partitions for Kafka cluster test-eqiad in eqiad - https://wikitech.wikimedia.org/wiki/Kafka/Administration - https://grafana.wikimedia.org/d/000000027/kafka?orgId=1&var-datasource=eqiad%20prometheus/ops&var-kafka_cluster=test-eqiad - https://alerts.wikimedia.org/?q=alertname%3DKafkaUnderReplicatedPartitions [06:07:20] RECOVERY - haproxy failover on dbproxy1029 is OK: OK check_failover servers up 2 down 0: https://wikitech.wikimedia.org/wiki/HAProxy [06:07:20] RECOVERY - haproxy failover on dbproxy1027 is OK: OK check_failover servers up 2 down 0: https://wikitech.wikimedia.org/wiki/HAProxy [06:08:34] 06SRE, 06DBA: db1228 crashed - https://phabricator.wikimedia.org/T430934#12113312 (10Marostegui) 05Open→03Resolved So far so good, closing for now. [06:15:24] PROBLEM - Check unit status of httpbb_kubernetes_mw-web-next_hourly on cumin2003 is CRITICAL: CRITICAL: Status of the systemd unit httpbb_kubernetes_mw-web-next_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [06:17:24] !log marostegui@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1022.eqiad.wmnet with reason: reboot [06:20:42] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [06:21:05] !log marostegui@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1025.eqiad.wmnet with reason: reboot [06:24:18] (03PS1) 10Marostegui: clouddb1028: Enable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1309897 (https://phabricator.wikimedia.org/T409557) [06:24:55] !log marostegui@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1028.eqiad.wmnet with reason: reboot [06:25:31] (03CR) 10Marostegui: [C:03+2] clouddb1028: Enable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1309897 (https://phabricator.wikimedia.org/T409557) (owner: 10Marostegui) [06:29:21] !log marostegui@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot [06:31:28] (03PS1) 10Marostegui: check_private_data_report: Add clouddb1028 [puppet] - 10https://gerrit.wikimedia.org/r/1309898 (https://phabricator.wikimedia.org/T409557) [06:32:29] (03CR) 10Marostegui: [C:03+2] check_private_data_report: Add clouddb1028 [puppet] - 10https://gerrit.wikimedia.org/r/1309898 (https://phabricator.wikimedia.org/T409557) (owner: 10Marostegui) [06:34:11] !log Drop m5 ipoid database T431007 [06:34:14] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [06:34:15] T431007: Drop the ipoid database from MariaDB m5 - https://phabricator.wikimedia.org/T431007 [06:37:54] RECOVERY - Kafka Broker Server on kafka-test1009 is OK: PROCS OK: 1 process with command name java, args Kafka /etc/kafka/server.properties https://wikitech.wikimedia.org/wiki/Kafka/Administration [06:38:00] (03PS1) 10Marostegui: dumps-eqiad-m5.sql.erb: Remove grants [puppet] - 10https://gerrit.wikimedia.org/r/1309900 (https://phabricator.wikimedia.org/T431007) [06:40:54] PROBLEM - Kafka Broker Server on kafka-test1009 is CRITICAL: PROCS CRITICAL: 0 processes with command name java, args Kafka /etc/kafka/server.properties https://wikitech.wikimedia.org/wiki/Kafka/Administration [06:42:38] (03CR) 10Marostegui: "@jcrespo@wikimedia.org I've already dropped the grants in db1217 and db2160." [puppet] - 10https://gerrit.wikimedia.org/r/1309900 (https://phabricator.wikimedia.org/T431007) (owner: 10Marostegui) [06:44:04] (03PS2) 10Marostegui: dumps-eqiad-m5.sql.erb: Remove grants [puppet] - 10https://gerrit.wikimedia.org/r/1309900 (https://phabricator.wikimedia.org/T431007) [06:47:51] (03CR) 10Neriah: [C:03+1] "recheck" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309894 (https://phabricator.wikimedia.org/T431959) (owner: 10Gkm563) [06:49:13] !log jelto@cumin1003 START - Cookbook sre.hosts.reboot-single for host gitlab1003.wikimedia.org [06:50:12] (03CR) 10Neriah: [C:03+1] "You will need to schedule this patch for deployment following the process at https://wikitech.wikimedia.org/wiki/Backport_windows#How_to_s" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309894 (https://phabricator.wikimedia.org/T431959) (owner: 10Gkm563) [06:50:59] PROBLEM - Host gitlab-replica-a.wikimedia.org is DOWN: PING CRITICAL - Packet loss = 100% [06:51:03] (03CR) 10Neriah: "You will need to schedule this patch for deployment following the process at https://wikitech.wikimedia.org/wiki/Backport_windows#How_to_s" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309727 (https://phabricator.wikimedia.org/T431858) (owner: 10Esanders) [06:52:01] PROBLEM - Docker registry HTTPS interface on registry1005 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/Docker [06:52:51] RECOVERY - Docker registry HTTPS interface on registry1005 is OK: HTTP OK: HTTP/1.1 200 OK - 3746 bytes in 0.196 second response time https://wikitech.wikimedia.org/wiki/Docker [06:55:33] !log jelto@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1003.wikimedia.org [06:55:42] FIRING: [2x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [06:55:52] !log jelto@cumin1003 START - Cookbook sre.hosts.reboot-single for host gitlab2002.wikimedia.org [06:56:01] RECOVERY - Host gitlab-replica-a.wikimedia.org is UP: PING OK - Packet loss = 0%, RTA = 0.35 ms [07:00:05] Amir1, urbanecm, and awight: Time to do the UTC morning backport window deploy. Don't look at me like that. You signed up for it. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260713T0700). [07:00:05] Msz2001 and dyepezg: A patch you scheduled for UTC morning backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [07:00:15] 06SRE, 06cloud-services-team, 10Infrastructure Security, 06Infrastructure-Foundations, and 5 others: July 2026 Bookworm/Trixie reboots: Data Platform SRE - https://phabricator.wikimedia.org/T431826#12113401 (10Jelto) [07:00:42] FIRING: [2x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [07:02:17] !log jelto@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2002.wikimedia.org [07:02:29] !log jelto@cumin1003 START - Cookbook sre.hosts.reboot-single for host gitlab2003.wikimedia.org [07:03:51] (03CR) 10TrainBranchBot: [C:03+2] "Approved by mszwarc@deploy2003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309559 (https://phabricator.wikimedia.org/T430512) (owner: 10Mszwarc) [07:04:45] (03Merged) 10jenkins-bot: plwiki: Switch back to normal tagline [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309559 (https://phabricator.wikimedia.org/T430512) (owner: 10Mszwarc) [07:05:11] !log mszwarc@deploy2003 Started scap sync-world: Backport for [[gerrit:1309559|plwiki: Switch back to normal tagline (T430512)]] [07:05:14] T430512: Temporarily change plwiki tagline to reflect 1.7M articles - https://phabricator.wikimedia.org/T430512 [07:05:42] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [07:05:53] (03PS1) 10Anzx: minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308652 (https://phabricator.wikimedia.org/T429943) [07:06:55] RECOVERY - Kafka Broker Server on kafka-test1009 is OK: PROCS OK: 1 process with command name java, args Kafka /etc/kafka/server.properties https://wikitech.wikimedia.org/wiki/Kafka/Administration [07:07:35] Msz2001: i change that needed to be deployed https://gerrit.wikimedia.org/r/c/operations/mediawiki-config/+/1308652 can you deploy it [07:07:54] Sure, once my deployment is done [07:07:55] RECOVERY - Kafka broker TLS certificate validity on kafka-test1009 is OK: SSL OK - Certificate kafka-test1009.eqiad.wmnet valid until 2027-04-03 10:04:00 +0000 (expires in 264 days) https://wikitech.wikimedia.org/wiki/Kafka/Administration%23Renew_TLS_certificate [07:08:14] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, July 13 UTC morning backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-it" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308652 (https://phabricator.wikimedia.org/T429943) (owner: 10Anzx) [07:08:49] !log jelto@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2003.wikimedia.org [07:08:51] i have added it to calendar, thanks [07:08:55] RECOVERY - Kafka Broker Server on kafka-test1007 is OK: PROCS OK: 1 process with command name java, args Kafka /etc/kafka/server.properties https://wikitech.wikimedia.org/wiki/Kafka/Administration [07:08:55] RECOVERY - Kafka broker TLS certificate validity on kafka-test1007 is OK: SSL OK - Certificate kafka-test1007.eqiad.wmnet valid until 2027-04-03 08:35:00 +0000 (expires in 264 days) https://wikitech.wikimedia.org/wiki/Kafka/Administration%23Renew_TLS_certificate [07:11:06] !log jelto@cumin1003 START - Cookbook sre.hosts.reboot-single for host aphlict1002.eqiad.wmnet [07:11:55] PROBLEM - Kafka Broker Server on kafka-test1007 is CRITICAL: PROCS CRITICAL: 0 processes with command name java, args Kafka /etc/kafka/server.properties https://wikitech.wikimedia.org/wiki/Kafka/Administration [07:11:55] PROBLEM - Kafka broker TLS certificate validity on kafka-test1007 is CRITICAL: SSL CRITICAL - failed to connect or SSL handshake:Connection refused https://wikitech.wikimedia.org/wiki/Kafka/Administration%23Renew_TLS_certificate [07:13:14] (03PS1) 10Danielyepezgarces: cowikimedia: Update localtimezone to America/Bogota [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309901 (https://phabricator.wikimedia.org/T431965) [07:13:51] (03CR) 10Brouberol: "`" [puppet] - 10https://gerrit.wikimedia.org/r/1309679 (https://phabricator.wikimedia.org/T429150) (owner: 10Brouberol) [07:14:55] RECOVERY - Kafka Broker Server on kafka-test1007 is OK: PROCS OK: 1 process with command name java, args Kafka /etc/kafka/server.properties https://wikitech.wikimedia.org/wiki/Kafka/Administration [07:14:55] RECOVERY - Kafka broker TLS certificate validity on kafka-test1007 is OK: SSL OK - Certificate kafka-test1007.eqiad.wmnet valid until 2027-04-03 08:35:00 +0000 (expires in 264 days) https://wikitech.wikimedia.org/wiki/Kafka/Administration%23Renew_TLS_certificate [07:15:04] !log jelto@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict1002.eqiad.wmnet [07:15:25] RECOVERY - Check unit status of httpbb_kubernetes_mw-web-next_hourly on cumin2003 is OK: OK: Status of the systemd unit httpbb_kubernetes_mw-web-next_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [07:15:44] (03CR) 10Brouberol: [C:03+1] Add a component/topolvm to bookworm-wikimedia in reprepro [puppet] - 10https://gerrit.wikimedia.org/r/1309701 (https://phabricator.wikimedia.org/T428895) (owner: 10Btullis) [07:17:04] RESOLVED: KafkaUnderReplicatedPartitions: Under replicated partitions for Kafka cluster test-eqiad in eqiad - https://wikitech.wikimedia.org/wiki/Kafka/Administration - https://grafana.wikimedia.org/d/000000027/kafka?orgId=1&var-datasource=eqiad%20prometheus/ops&var-kafka_cluster=test-eqiad - https://alerts.wikimedia.org/?q=alertname%3DKafkaUnderReplicatedPartitions [07:19:03] !log jelto@cumin1003 START - Cookbook sre.hosts.reboot-single for host gerrit2002.wikimedia.org [07:21:14] !log mszwarc@deploy2003 mszwarc: Backport for [[gerrit:1309559|plwiki: Switch back to normal tagline (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [07:21:18] T430512: Temporarily change plwiki tagline to reflect 1.7M articles - https://phabricator.wikimedia.org/T430512 [07:22:41] !log mszwarc@deploy2003 mszwarc: Continuing with deployment [07:23:55] (03CR) 10Brouberol: [C:03+1] admin_ng: configure topolvm storage classes and node placement (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309261 (https://phabricator.wikimedia.org/T428235) (owner: 10Btullis) [07:25:12] dyepezg: Are you around? I will proceed to deploying your patch in a moment [07:25:14] !log jelto@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2002.wikimedia.org [07:25:31] yeap, i am here [07:25:42] FIRING: [2x] JobUnavailable: Reduced availability for job atlas_exporter in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [07:26:05] Okay, the current deployment will finish in a few minutes, after which I'll deploy both your patches [07:26:24] (03CR) 10Brouberol: [C:03+2] trafficserver: redirect query-next.w.o/sparql to qlever in dse-k8s [puppet] - 10https://gerrit.wikimedia.org/r/1309679 (https://phabricator.wikimedia.org/T429150) (owner: 10Brouberol) [07:26:58] Could you please review change 1309901, which I made at the last minute to add it to the queue? [07:27:12] (03CR) 10Mszwarc: [C:03+2] "Ahead of deployment" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309302 (https://phabricator.wikimedia.org/T431765) (owner: 10Danielyepezgarces) [07:27:21] (03CR) 10Mszwarc: [C:03+2] "Ahead of deployment" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1307836 (https://phabricator.wikimedia.org/T431334) (owner: 10Danielyepezgarces) [07:27:37] Looking [07:28:10] Okay, will deploy that one as well [07:28:21] (03CR) 10Mszwarc: [C:03+2] "Ahead of deployment" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309901 (https://phabricator.wikimedia.org/T431965) (owner: 10Danielyepezgarces) [07:28:26] (03Merged) 10jenkins-bot: Enable campaignEvents on cowikimedia [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309302 (https://phabricator.wikimedia.org/T431765) (owner: 10Danielyepezgarces) [07:28:34] (03Merged) 10jenkins-bot: extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1307836 (https://phabricator.wikimedia.org/T431334) (owner: 10Danielyepezgarces) [07:29:26] (03Merged) 10jenkins-bot: cowikimedia: Update localtimezone to America/Bogota [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309901 (https://phabricator.wikimedia.org/T431965) (owner: 10Danielyepezgarces) [07:30:42] FIRING: [2x] JobUnavailable: Reduced availability for job atlas_exporter in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [07:32:13] (03CR) 10Mszwarc: [C:03+2] "Ahead of deployment" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308652 (https://phabricator.wikimedia.org/T429943) (owner: 10Anzx) [07:33:06] (03Merged) 10jenkins-bot: minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308652 (https://phabricator.wikimedia.org/T429943) (owner: 10Anzx) [07:35:14] !log mszwarc@deploy2003 Finished scap sync-world: Backport for [[gerrit:1309559|plwiki: Switch back to normal tagline (T430512)]] (duration: 30m 03s) [07:35:18] T430512: Temporarily change plwiki tagline to reflect 1.7M articles - https://phabricator.wikimedia.org/T430512 [07:36:01] !log mszwarc@deploy2003 Started scap sync-world: Backport for [[gerrit:1309302|Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836|extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901|cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652|minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943) [07:36:01] ]] [07:36:11] T431765: Enable CampaignEvents extension on cowikimedia - https://phabricator.wikimedia.org/T431765 [07:36:11] T431334: extwiki: Update wgSitename due wrong title - https://phabricator.wikimedia.org/T431334 [07:36:12] T431965: cowikimedia: Update timezone to America/Bogota - https://phabricator.wikimedia.org/T431965 [07:36:12] T429943: Post-creation work for minwikiquote - https://phabricator.wikimedia.org/T429943 [07:37:24] (03CR) 10Marostegui: es1039: Enable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1307678 (https://phabricator.wikimedia.org/T430764) (owner: 10Marostegui) [07:37:33] (03CR) 10Marostegui: [C:03+2] es1039: Enable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1307678 (https://phabricator.wikimedia.org/T430764) (owner: 10Marostegui) [07:39:27] !log marostegui@cumin1003 START - Cookbook sre.mysql.pool pool es1039: Repooling after testing [07:39:32] !log mszwarc@deploy2003 mszwarc, danielyepezgarces, anzx: Backport for [[gerrit:1309302|Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836|extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901|cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652|minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark [07:39:33] (T429943)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [07:39:37] dyepezg, anzx: Please verify the changes [07:39:43] 10ops-eqiad, 06Data-Persistence, 06DBA, 06DC-Ops, and 3 others: es7 primary (es1039.eqiad.wmnet) crashed - https://phabricator.wikimedia.org/T430764#12113488 (10ops-monitoring-bot) Starting pool of es1039 by marostegui@cumin1003: Repooling after testing [07:39:57] looking [07:40:34] Msz2001: looks good, ok to sync [07:40:44] (03PS1) 10Marostegui: eqiad.yaml: Add clouddb1028 to the LB [puppet] - 10https://gerrit.wikimedia.org/r/1309903 (https://phabricator.wikimedia.org/T409557) [07:41:27] (03PS2) 10Brouberol: cache/text: enable caching for query-next.w.o [puppet] - 10https://gerrit.wikimedia.org/r/1309680 (https://phabricator.wikimedia.org/T429150) [07:41:34] (03CR) 10Marostegui: [C:03+2] "@fnegri@wikimedia.org FYI" [puppet] - 10https://gerrit.wikimedia.org/r/1309903 (https://phabricator.wikimedia.org/T409557) (owner: 10Marostegui) [07:41:40] (03CR) 10Brouberol: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1309680 (https://phabricator.wikimedia.org/T429150) (owner: 10Brouberol) [07:43:32] I have verified the other ones as well, proceeding [07:43:38] !log mszwarc@deploy2003 mszwarc, danielyepezgarces, anzx: Continuing with deployment [07:44:31] Sorry for the delay. That works for me too. [07:44:54] !log marostegui@cumin1003 conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 [07:44:58] !log marostegui@cumin1003 conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s6 [07:45:11] !log marostegui@cumin1003 conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s4 [07:45:16] !log marostegui@cumin1003 conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s6 [07:46:18] !log marostegui@cumin1003 conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 [07:46:33] !log marostegui@cumin1003 conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4 [07:49:30] !log filippo@cumin1003 START - Cookbook sre.hosts.reboot-single for host cloudbackup1001-dev.eqiad.wmnet [07:50:13] !log mszwarc@deploy2003 Finished scap sync-world: Backport for [[gerrit:1309302|Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836|extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901|cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652|minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943 [07:50:13] )]] (duration: 14m 12s) [07:50:17] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host thanos-be2005.codfw.wmnet with OS trixie [07:50:21] T431765: Enable CampaignEvents extension on cowikimedia - https://phabricator.wikimedia.org/T431765 [07:50:22] T431334: extwiki: Update wgSitename due wrong title - https://phabricator.wikimedia.org/T431334 [07:50:22] T431965: cowikimedia: Update timezone to America/Bogota - https://phabricator.wikimedia.org/T431965 [07:50:23] T429943: Post-creation work for minwikiquote - https://phabricator.wikimedia.org/T429943 [07:50:31] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12113512 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host thanos-be2005.codfw.wmnet with OS trixie [07:51:16] 10ops-eqiad, 06Data-Persistence, 06DBA, 06DC-Ops, and 3 others: es7 primary (es1039.eqiad.wmnet) crashed - https://phabricator.wikimedia.org/T430764#12113514 (10Marostegui) I am slowly repooling this host after a few successful tests. [07:51:17] Deployed the changes [07:51:42] !Log UTC morning backport+config window done [07:52:37] !log UTC morning backport+config window done [07:52:39] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [07:52:46] Seems it must be lowercase [07:53:49] !log marostegui@cumin1003 conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s4 [07:53:53] !log marostegui@cumin1003 conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s6 [07:54:15] !log jelto@cumin1003 START - Cookbook sre.hosts.reboot-single for host gerrit1003.wikimedia.org [07:54:49] !log filippo@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1001-dev.eqiad.wmnet [07:54:53] !log filippo@cumin1003 START - Cookbook sre.hosts.reboot-single for host cloudbackup1002-dev.eqiad.wmnet [07:56:29] 06SRE, 10SRE-SLO, 06Traffic: Page on ATS backend errors relative to traffic - https://phabricator.wikimedia.org/T400675#12113544 (10fgiunchedi) This is still an issue: SRE got paged for 10 5xx req/s on swift which was doing 2k req/s at the time (esams). I very much doubt it is worth paging engineers on absol... [07:58:12] !log trueg@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply [07:58:20] !log trueg@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply [07:58:41] !log filippo@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1002-dev.eqiad.wmnet [07:58:45] !log filippo@cumin1003 START - Cookbook sre.hosts.reboot-single for host cloudbackup1003.eqiad.wmnet [08:00:18] !log jelto@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit1003.wikimedia.org [08:02:25] (03PS1) 10Marostegui: wmnet: Failover m1-master [dns] - 10https://gerrit.wikimedia.org/r/1309906 (https://phabricator.wikimedia.org/T431660) [08:04:40] (03CR) 10Marostegui: [C:03+2] wmnet: Failover m1-master [dns] - 10https://gerrit.wikimedia.org/r/1309906 (https://phabricator.wikimedia.org/T431660) (owner: 10Marostegui) [08:04:56] !log marostegui@dns1004 START - running authdns-update [08:05:04] !log trueg@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply [08:05:05] !log marostegui@dns1004 START - running authdns-update [08:05:09] !log marostegui@dns1004 START - running authdns-update [08:05:10] !log jelto@cumin1003 START - Cookbook sre.hosts.reboot-single for host phab1005.eqiad.wmnet [08:05:17] !log trueg@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply [08:05:28] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage [08:05:42] !log marostegui@dns1004 START - running authdns-update [08:06:37] !log trueg@deploy2003 helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply [08:06:42] !log trueg@deploy2003 helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply [08:06:55] !log trueg@deploy2003 helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply [08:07:09] !log trueg@deploy2003 helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply [08:07:22] !log filippo@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1003.eqiad.wmnet [08:07:25] !log filippo@cumin1003 START - Cookbook sre.hosts.reboot-single for host cloudbackup1004.eqiad.wmnet [08:07:31] !log marostegui@dns1004 END - running authdns-update [08:08:26] (03PS1) 10Trueg: WDQS: fixed Qlever service hostname in proxy config [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309907 [08:08:47] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage [08:09:26] (03PS2) 10Trueg: WDQS: fixed Qlever service hostname in proxy config [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309907 (https://phabricator.wikimedia.org/T429150) [08:10:50] (03CR) 10Brouberol: [C:03+1] WDQS: fixed Qlever service hostname in proxy config [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309907 (https://phabricator.wikimedia.org/T429150) (owner: 10Trueg) [08:11:15] (03PS1) 10JavierMonton: stream: pageview-trending-relative [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309908 (https://phabricator.wikimedia.org/T430675) [08:11:40] !log jelto@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host phab1005.eqiad.wmnet [08:14:31] !log btullis@cumin1003 START - Cookbook sre.hosts.reboot-single for host db1208.eqiad.wmnet [08:14:38] (03CR) 10Gmodena: [C:03+2] WDQS: fixed Qlever service hostname in proxy config [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309907 (https://phabricator.wikimedia.org/T429150) (owner: 10Trueg) [08:16:36] (03CR) 10Btullis: [C:03+2] Add a component/topolvm to bookworm-wikimedia in reprepro [puppet] - 10https://gerrit.wikimedia.org/r/1309701 (https://phabricator.wikimedia.org/T428895) (owner: 10Btullis) [08:17:12] (03Merged) 10jenkins-bot: WDQS: fixed Qlever service hostname in proxy config [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309907 (https://phabricator.wikimedia.org/T429150) (owner: 10Trueg) [08:17:53] !log filippo@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1004.eqiad.wmnet [08:17:56] (03CR) 10Filippo Giunchedi: [C:03+2] openstack: remove openstack::monitor::neutron::l3_agent_conntrack [puppet] - 10https://gerrit.wikimedia.org/r/1308626 (https://phabricator.wikimedia.org/T328502) (owner: 10Filippo Giunchedi) [08:17:57] !log filippo@cumin1003 START - Cookbook sre.hosts.reboot-single for host cloudbackup2003.codfw.wmnet [08:20:52] (03CR) 10Ayounsi: [C:03+2] depool-rack: various fix [cookbooks] - 10https://gerrit.wikimedia.org/r/1309030 (https://phabricator.wikimedia.org/T327300) (owner: 10Ayounsi) [08:21:34] !log marostegui@cumin1003 conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s6 [08:21:37] !log marostegui@cumin1003 conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4 [08:23:21] !log marostegui@cumin1003 conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 [08:23:24] !log marostegui@cumin1003 conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 [08:23:49] !log marostegui@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb1016.eqiad.wmnet with reason: cloning [08:23:50] (03Merged) 10jenkins-bot: depool-rack: various fix [cookbooks] - 10https://gerrit.wikimedia.org/r/1309030 (https://phabricator.wikimedia.org/T327300) (owner: 10Ayounsi) [08:24:56] !log marostegui@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1039: Repooling after testing [08:25:10] 10ops-eqiad, 06Data-Persistence, 06DBA, 06DC-Ops, and 3 others: es7 primary (es1039.eqiad.wmnet) crashed - https://phabricator.wikimedia.org/T430764#12113645 (10ops-monitoring-bot) Completed pooling of es1039 by marostegui@cumin1003: Repooling after testing [08:25:42] RECOVERY - MariaDB disk space on db1208 is OK: DISK OK https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting [08:25:42] RECOVERY - MariaDB Replica SQL: analytics_meta on db1208 is OK: OK slave_sql_state Slave_SQL_Running: Yes https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [08:25:42] RECOVERY - MariaDB Replica IO: analytics_meta on db1208 is OK: OK slave_io_state Slave_IO_Running: Yes https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [08:25:44] RECOVERY - MariaDB read only analytics_meta on db1208 is OK: Version 10.6.18-MariaDB-log, Uptime 30s, read_only: True, event_scheduler: True, 6610.62 QPS, connection latency: 0.038250s https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Master_comes_back_in_read_only [08:26:42] RECOVERY - mysqld processes on db1208 is OK: PROCS OK: 2 processes with command name mysqld https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting [08:26:42] RECOVERY - MariaDB Replica SQL: matomo on db1208 is OK: OK slave_sql_state Slave_SQL_Running: Yes https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [08:26:42] RECOVERY - MariaDB read only matomo on db1208 is OK: Version 10.6.18-MariaDB-log, Uptime 48s, read_only: True, event_scheduler: True, 11.20 QPS, connection latency: 0.035233s https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Master_comes_back_in_read_only [08:27:42] RECOVERY - MariaDB Replica IO: matomo on db1208 is OK: OK slave_io_state Slave_IO_Running: Yes https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [08:28:20] !log filippo@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2003.codfw.wmnet [08:28:21] (03PS1) 10Marostegui: mariadb: Productionize clouddb1029 [puppet] - 10https://gerrit.wikimedia.org/r/1310003 (https://phabricator.wikimedia.org/T409557) [08:28:24] !log filippo@cumin1003 START - Cookbook sre.hosts.reboot-single for host cloudbackup2004.codfw.wmnet [08:28:35] (03CR) 10Brouberol: [C:03+2] cache/text: enable caching for query-next.w.o [puppet] - 10https://gerrit.wikimedia.org/r/1309680 (https://phabricator.wikimedia.org/T429150) (owner: 10Brouberol) [08:28:36] FIRING: [5x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-3/1/2 (Transport: Hurricane Electric (dc4841.dal5) {#122601}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [08:28:54] PROBLEM - Postgres Replication Lag on puppetdb2003 is CRITICAL: POSTGRES_HOT_STANDBY_DELAY CRITICAL: DB puppetdb (host:localhost) 224155392 and 8 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [08:29:26] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2005.codfw.wmnet with OS trixie [08:29:40] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12113659 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host thanos-be2005.codfw.wmnet with OS trixie compl... [08:30:25] (03CR) 10Marostegui: [C:03+2] mariadb: Productionize clouddb1029 [puppet] - 10https://gerrit.wikimedia.org/r/1310003 (https://phabricator.wikimedia.org/T409557) (owner: 10Marostegui) [08:30:50] !log marostegui@cumin1003 conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=x3 [08:30:54] RECOVERY - Postgres Replication Lag on puppetdb2003 is OK: POSTGRES_HOT_STANDBY_DELAY OK: DB puppetdb (host:localhost) 3039744 and 1 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [08:32:02] (03PS1) 10Marostegui: clouddb1028: Remove note [puppet] - 10https://gerrit.wikimedia.org/r/1310010 [08:32:58] (03CR) 10Marostegui: [C:03+2] clouddb1028: Remove note [puppet] - 10https://gerrit.wikimedia.org/r/1310010 (owner: 10Marostegui) [08:33:45] FIRING: SwiftLowContainerAvailability: Swift eqiad container availability low - https://wikitech.wikimedia.org/wiki/Swift/How_To - https://grafana.wikimedia.org/d/OPgmB1Eiz/swift?panelId=8&fullscreen&orgId=1&var-DC=eqiad - https://alerts.wikimedia.org/?q=alertname%3DSwiftLowContainerAvailability [08:33:55] !log btullis@cumin1003 END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host db1208.eqiad.wmnet [08:34:09] jouncebot: now [08:34:09] No deployments scheduled for the next 1 hour(s) and 25 minute(s) [08:34:17] (03CR) 10Arthur taylor: [C:03+2] "merging so that it can be deployed." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309103 (https://phabricator.wikimedia.org/T431080) (owner: 10Arthur taylor) [08:34:26] PROBLEM - MariaDB Replica Lag: analytics_meta on db1208 is CRITICAL: CRITICAL slave_sql_lag Replication lag: 265454.00 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [08:34:28] PROBLEM - MariaDB Replica Lag: matomo on db1208 is CRITICAL: CRITICAL slave_sql_lag Replication lag: 282761.00 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [08:36:53] (03Merged) 10jenkins-bot: Upgrade to the newly-released version of the Queryservice GUI [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309103 (https://phabricator.wikimedia.org/T431080) (owner: 10Arthur taylor) [08:37:59] !log arthurtaylor@deploy2003 helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply [08:39:56] (03PS1) 10Brouberol: trafficserver: redirect query-scholarly next.w.o/sparql to wdqs-scholarly in dse-k8s [puppet] - 10https://gerrit.wikimedia.org/r/1310012 (https://phabricator.wikimedia.org/T429150) [08:39:59] (03PS1) 10Brouberol: cache/text: enable caching for query-scholarly next.w.o [puppet] - 10https://gerrit.wikimedia.org/r/1310013 (https://phabricator.wikimedia.org/T429150) [08:40:30] (03CR) 10CI reject: [V:04-1] trafficserver: redirect query-scholarly next.w.o/sparql to wdqs-scholarly in dse-k8s [puppet] - 10https://gerrit.wikimedia.org/r/1310012 (https://phabricator.wikimedia.org/T429150) (owner: 10Brouberol) [08:40:39] (03CR) 10Trueg: [C:03+1] trafficserver: redirect query-scholarly next.w.o/sparql to wdqs-scholarly in dse-k8s [puppet] - 10https://gerrit.wikimedia.org/r/1310012 (https://phabricator.wikimedia.org/T429150) (owner: 10Brouberol) [08:40:57] (03CR) 10Trueg: [C:03+1] cache/text: enable caching for query-scholarly next.w.o [puppet] - 10https://gerrit.wikimedia.org/r/1310013 (https://phabricator.wikimedia.org/T429150) (owner: 10Brouberol) [08:40:59] !log filippo@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2004.codfw.wmnet [08:41:32] !log arthurtaylor@deploy2003 helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply [08:41:38] (03PS2) 10Brouberol: trafficserver: redirect query-scholarly next.w.o/sparql to dse/wdqs-scholarly [puppet] - 10https://gerrit.wikimedia.org/r/1310012 (https://phabricator.wikimedia.org/T429150) [08:41:38] (03PS2) 10Brouberol: cache/text: enable caching for query-scholarly next.w.o [puppet] - 10https://gerrit.wikimedia.org/r/1310013 (https://phabricator.wikimedia.org/T429150) [08:42:09] !log trueg@deploy2003 helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply [08:42:12] !log trueg@deploy2003 helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply [08:42:21] !log arthurtaylor@deploy2003 helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply [08:42:33] (03CR) 10Gmodena: [C:03+1] trafficserver: redirect query-scholarly next.w.o/sparql to dse/wdqs-scholarly [puppet] - 10https://gerrit.wikimedia.org/r/1310012 (https://phabricator.wikimedia.org/T429150) (owner: 10Brouberol) [08:42:57] (03CR) 10Gmodena: [C:03+1] cache/text: enable caching for query-scholarly next.w.o [puppet] - 10https://gerrit.wikimedia.org/r/1310013 (https://phabricator.wikimedia.org/T429150) (owner: 10Brouberol) [08:43:20] !log arthurtaylor@deploy2003 helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply [08:43:39] !log arthurtaylor@deploy2003 helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply [08:44:06] !log arthurtaylor@deploy2003 helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply [08:44:50] (03CR) 10Hnowlan: [C:03+1] thumbor: Allow 2 connection per worker [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309832 (https://phabricator.wikimedia.org/T431767) (owner: 10Ladsgroup) [08:45:27] RECOVERY - MariaDB Replica Lag: matomo on db1208 is OK: OK slave_sql_lag Replication lag: 0.00 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [08:46:45] (03CR) 10Btullis: [C:03+1] trafficserver: redirect query-scholarly next.w.o/sparql to dse/wdqs-scholarly [puppet] - 10https://gerrit.wikimedia.org/r/1310012 (https://phabricator.wikimedia.org/T429150) (owner: 10Brouberol) [08:47:08] (03CR) 10Btullis: [C:03+1] cache/text: enable caching for query-scholarly next.w.o [puppet] - 10https://gerrit.wikimedia.org/r/1310013 (https://phabricator.wikimedia.org/T429150) (owner: 10Brouberol) [08:48:27] RECOVERY - MariaDB Replica Lag: analytics_meta on db1208 is OK: OK slave_sql_lag Replication lag: 0.00 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [08:48:45] RESOLVED: SwiftLowContainerAvailability: Swift eqiad container availability low - https://wikitech.wikimedia.org/wiki/Swift/How_To - https://grafana.wikimedia.org/d/OPgmB1Eiz/swift?panelId=8&fullscreen&orgId=1&var-DC=eqiad - https://alerts.wikimedia.org/?q=alertname%3DSwiftLowContainerAvailability [08:52:09] (03PS2) 10JavierMonton: stream: pageview-trending-relative [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309908 (https://phabricator.wikimedia.org/T430675) [08:52:44] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, July 13 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309894 (https://phabricator.wikimedia.org/T431959) (owner: 10Gkm563) [08:55:48] (03PS2) 10Priyankar22: thwiki: Change to Wikipedia 25 logo [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309865 (https://phabricator.wikimedia.org/T431094) [08:57:25] (03CR) 10Anzx: "recheck" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309865 (https://phabricator.wikimedia.org/T431094) (owner: 10Priyankar22) [08:57:45] (03CR) 10Priyankar22: "recheck" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309865 (https://phabricator.wikimedia.org/T431094) (owner: 10Priyankar22) [08:58:13] (03CR) 10Brouberol: [C:03+2] trafficserver: redirect query-scholarly next.w.o/sparql to dse/wdqs-scholarly [puppet] - 10https://gerrit.wikimedia.org/r/1310012 (https://phabricator.wikimedia.org/T429150) (owner: 10Brouberol) [08:59:07] (03CR) 10Filippo Giunchedi: [C:03+2] statistics: optional support for dumps-nfs.w.o [puppet] - 10https://gerrit.wikimedia.org/r/1308127 (https://phabricator.wikimedia.org/T411248) (owner: 10Filippo Giunchedi) [08:59:13] (03CR) 10Filippo Giunchedi: [C:03+2] wmcs: add dumps-nfs.w.o support [puppet] - 10https://gerrit.wikimedia.org/r/1308128 (https://phabricator.wikimedia.org/T411248) (owner: 10Filippo Giunchedi) [08:59:16] (03CR) 10Filippo Giunchedi: [C:03+2] labstore: don't check after umount [puppet] - 10https://gerrit.wikimedia.org/r/1309200 (https://phabricator.wikimedia.org/T411248) (owner: 10Filippo Giunchedi) [09:01:00] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host thanos-be2006.codfw.wmnet with OS trixie [09:01:16] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12113785 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host thanos-be2006.codfw.wmnet with OS trixie [09:01:47] 06SRE, 06MediaWiki-Media-Platform-Team, 06ServiceOps new, 10ServiceOps-Services-Oids, 10Thumbor: Thumbor-k8s performance improvements - https://phabricator.wikimedia.org/T333445#12113786 (10hnowlan) >>! In T333445#12107169, @Ladsgroup wrote: > Cormac suggested we close this ticket since it has been a whi... [09:02:03] (03CR) 10Hnowlan: [C:04-1] "LGTM, but needs a Chart bump." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309832 (https://phabricator.wikimedia.org/T431767) (owner: 10Ladsgroup) [09:03:12] (03CR) 10Anzx: "you need to edit `logos/config.yaml` file again, some already merged chages are being reverted by this change" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309865 (https://phabricator.wikimedia.org/T431094) (owner: 10Priyankar22) [09:06:22] !log trueg@deploy2003 helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply [09:06:25] !log trueg@deploy2003 helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply [09:06:42] !log marostegui@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1024.eqiad.wmnet with reason: reboot [09:06:48] (03PS3) 10Ladsgroup: thumbor: Allow 2 connection per worker [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309832 (https://phabricator.wikimedia.org/T431767) [09:07:22] (03CR) 10Ladsgroup: "Done!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309832 (https://phabricator.wikimedia.org/T431767) (owner: 10Ladsgroup) [09:11:00] PROBLEM - Postgres Replication Lag on puppetdb2003 is CRITICAL: POSTGRES_HOT_STANDBY_DELAY CRITICAL: DB puppetdb (host:localhost) 187319136 and 3 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [09:12:00] RECOVERY - Postgres Replication Lag on puppetdb2003 is OK: POSTGRES_HOT_STANDBY_DELAY OK: DB puppetdb (host:localhost) 1288712 and 0 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [09:13:28] (03CR) 10Ayounsi: "lgtm!" [homer/public] - 10https://gerrit.wikimedia.org/r/1309707 (https://phabricator.wikimedia.org/T431849) (owner: 10Cathal Mooney) [09:17:33] (03PS2) 10Atsuko: Revert "cirrussearch: alert on tripped circuit breakers" [alerts] - 10https://gerrit.wikimedia.org/r/1309691 [09:17:38] (03PS1) 10JavierMonton: topic: webrequest-pageview [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310018 (https://phabricator.wikimedia.org/T430675) [09:18:31] (03PS3) 10Atsuko: Revert "cirrussearch: alert on tripped circuit breakers" [alerts] - 10https://gerrit.wikimedia.org/r/1309691 (https://phabricator.wikimedia.org/T431827) [09:20:03] (03CR) 10Hnowlan: "Sorry - a few more notes on things I already commented upon 🫣" [puppet] - 10https://gerrit.wikimedia.org/r/1305774 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [09:20:37] (03CR) 10Majavah: [C:03+1] tables-catalog.yaml: Remove users_to_rename [puppet] - 10https://gerrit.wikimedia.org/r/1309893 (https://phabricator.wikimedia.org/T431842) (owner: 10Marostegui) [09:20:47] (03CR) 10Marostegui: [C:03+2] tables-catalog.yaml: Remove users_to_rename [puppet] - 10https://gerrit.wikimedia.org/r/1309893 (https://phabricator.wikimedia.org/T431842) (owner: 10Marostegui) [09:22:31] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage [09:24:55] (03CR) 10Ayounsi: Nokia SR-Linux: add export policy for unicast IBGP clusters (031 comment) [homer/public] - 10https://gerrit.wikimedia.org/r/1309643 (https://phabricator.wikimedia.org/T423430) (owner: 10Cathal Mooney) [09:29:09] PROBLEM - Docker registry HTTPS interface on registry1004 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/Docker [09:29:41] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage [09:30:09] RECOVERY - Docker registry HTTPS interface on registry1004 is OK: HTTP OK: HTTP/1.1 200 OK - 3746 bytes in 9.467 second response time https://wikitech.wikimedia.org/wiki/Docker [09:32:15] (03PS3) 10Priyankar22: thwiki: Change to Wikipedia 25 logo [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309865 (https://phabricator.wikimedia.org/T431094) [09:32:31] (03PS1) 10Marostegui: Revert "wmnet: Failover m1-master" [dns] - 10https://gerrit.wikimedia.org/r/1310023 [09:33:41] (03PS1) 10David Caro: toolforge::k8s::haproxy: tweak timeouts, limits and open files [puppet] - 10https://gerrit.wikimedia.org/r/1310024 (https://phabricator.wikimedia.org/T431946) [09:33:42] (03CR) 10Priyankar22: "recheck" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309865 (https://phabricator.wikimedia.org/T431094) (owner: 10Priyankar22) [09:36:45] (03PS1) 10JavierMonton: topic: webrequest-page-view-next [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310025 (https://phabricator.wikimedia.org/T430675) [09:40:46] (03CR) 10Marostegui: [C:03+2] Revert "wmnet: Failover m1-master" [dns] - 10https://gerrit.wikimedia.org/r/1310023 (owner: 10Marostegui) [09:40:50] !log marostegui@dns1004 START - running authdns-update [09:42:39] !log marostegui@dns1004 END - running authdns-update [09:43:21] (03CR) 10Anzx: [C:03+1] "looks good, please schedule it for deployment by following instructions on https://wikitech.wikimedia.org/wiki/Deployments#Getting_started" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309865 (https://phabricator.wikimedia.org/T431094) (owner: 10Priyankar22) [09:43:34] (03CR) 10Anzx: [C:03+1] "recheck" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309865 (https://phabricator.wikimedia.org/T431094) (owner: 10Priyankar22) [09:43:40] (03CR) 10David Caro: [V:03+1] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/8980/co" [puppet] - 10https://gerrit.wikimedia.org/r/1310024 (https://phabricator.wikimedia.org/T431946) (owner: 10David Caro) [09:44:47] (03CR) 10Hnowlan: [C:03+2] hadoop: emit HDFS HA status to prometheus [puppet] - 10https://gerrit.wikimedia.org/r/1304766 (https://phabricator.wikimedia.org/T407138) (owner: 10Hnowlan) [09:44:56] (03CR) 10Atsuko: [C:03+2] Revert "cirrussearch: alert on tripped circuit breakers" [alerts] - 10https://gerrit.wikimedia.org/r/1309691 (https://phabricator.wikimedia.org/T431827) (owner: 10Atsuko) [09:45:44] (03CR) 10Kevin Bazira: [C:03+1] ml-services: enable revertrisk-wikidata event stream predictions [deployment-charts] - 10https://gerrit.wikimedia.org/r/1307441 (https://phabricator.wikimedia.org/T420883) (owner: 10AikoChou) [09:46:51] (03Merged) 10jenkins-bot: Revert "cirrussearch: alert on tripped circuit breakers" [alerts] - 10https://gerrit.wikimedia.org/r/1309691 (https://phabricator.wikimedia.org/T431827) (owner: 10Atsuko) [09:48:00] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2006.codfw.wmnet with OS trixie [09:48:12] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12113882 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host thanos-be2006.codfw.wmnet with OS trixie compl... [09:48:39] (03CR) 10Kevin Bazira: [C:03+1] EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1307438 (https://phabricator.wikimedia.org/T420883) (owner: 10AikoChou) [09:50:47] (03CR) 10Kevin Bazira: [C:03+1] changeprop: add liftwing revertrisk-wikidata to production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1307434 (https://phabricator.wikimedia.org/T420883) (owner: 10AikoChou) [09:53:32] (03PS2) 10AikoChou: EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1307438 (https://phabricator.wikimedia.org/T420883) [09:53:45] FIRING: SwiftLowContainerAvailability: Swift eqiad container availability low - https://wikitech.wikimedia.org/wiki/Swift/How_To - https://grafana.wikimedia.org/d/OPgmB1Eiz/swift?panelId=8&fullscreen&orgId=1&var-DC=eqiad - https://alerts.wikimedia.org/?q=alertname%3DSwiftLowContainerAvailability [09:54:53] (03PS1) 10Marostegui: sections.yaml: Add x4 in codfw [puppet] - 10https://gerrit.wikimedia.org/r/1310026 (https://phabricator.wikimedia.org/T404715) [09:55:45] (03CR) 10Marostegui: [C:03+2] sections.yaml: Add x4 in codfw [puppet] - 10https://gerrit.wikimedia.org/r/1310026 (https://phabricator.wikimedia.org/T404715) (owner: 10Marostegui) [09:56:29] (03PS3) 10Blake: mw-pretrain: Basic Ingress configuration. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309641 (https://phabricator.wikimedia.org/T427668) [09:58:32] (03CR) 10Jcrespo: [C:03+1] dumps-eqiad-m5.sql.erb: Remove grants [puppet] - 10https://gerrit.wikimedia.org/r/1309900 (https://phabricator.wikimedia.org/T431007) (owner: 10Marostegui) [10:00:04] Deploy window MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260713T1000) [10:03:45] RESOLVED: SwiftLowContainerAvailability: Swift eqiad container availability low - https://wikitech.wikimedia.org/wiki/Swift/How_To - https://grafana.wikimedia.org/d/OPgmB1Eiz/swift?panelId=8&fullscreen&orgId=1&var-DC=eqiad - https://alerts.wikimedia.org/?q=alertname%3DSwiftLowContainerAvailability [10:04:35] FIRING: HelmReleaseBadStatus: Helm release wdqs/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [10:04:54] (03PS1) 10Mszwarc: SI: Fix client side instrumentation [extensions/CheckUser] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1310027 (https://phabricator.wikimedia.org/T431977) [10:05:23] FIRING: [13x] CertAlmostExpired: gNMI TLS certificate for lsw1-c2-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [10:05:55] (03PS1) 10Majavah: P:toolforge::k8s::haproxy: Increase open file descriptor limit [puppet] - 10https://gerrit.wikimedia.org/r/1310028 (https://phabricator.wikimedia.org/T431946) [10:05:58] (03PS1) 10Majavah: P:toolforge::k8s::haproxy: Set a global connection limit [puppet] - 10https://gerrit.wikimedia.org/r/1310029 (https://phabricator.wikimedia.org/T431946) [10:06:01] (03PS1) 10Majavah: P:toolforge::k8s::haproxy: Set aggressive timeouts on HTTP redirects [puppet] - 10https://gerrit.wikimedia.org/r/1310030 (https://phabricator.wikimedia.org/T431946) [10:06:03] (03PS1) 10Majavah: P:toolforge::k8s::haproxy: Decrease client timeouts [puppet] - 10https://gerrit.wikimedia.org/r/1310031 (https://phabricator.wikimedia.org/T431946) [10:07:22] (03CR) 10Marostegui: [C:03+2] dumps-eqiad-m5.sql.erb: Remove grants [puppet] - 10https://gerrit.wikimedia.org/r/1309900 (https://phabricator.wikimedia.org/T431007) (owner: 10Marostegui) [10:09:35] (03CR) 10David Caro: [C:03+1] P:toolforge::k8s::haproxy: Increase open file descriptor limit [puppet] - 10https://gerrit.wikimedia.org/r/1310028 (https://phabricator.wikimedia.org/T431946) (owner: 10Majavah) [10:10:49] (03CR) 10Majavah: [C:03+2] P:toolforge::k8s::haproxy: Increase open file descriptor limit [puppet] - 10https://gerrit.wikimedia.org/r/1310028 (https://phabricator.wikimedia.org/T431946) (owner: 10Majavah) [10:11:01] (03CR) 10David Caro: [C:03+1] "During the outage I was seeing issues when passing ~60k connections, though might be coupled with the leftover CLOSE-WAIT sockets also pil" [puppet] - 10https://gerrit.wikimedia.org/r/1310029 (https://phabricator.wikimedia.org/T431946) (owner: 10Majavah) [10:11:56] (03CR) 10David Caro: [C:03+1] P:toolforge::k8s::haproxy: Set aggressive timeouts on HTTP redirects [puppet] - 10https://gerrit.wikimedia.org/r/1310030 (https://phabricator.wikimedia.org/T431946) (owner: 10Majavah) [10:12:06] (03CR) 10Majavah: [C:03+2] P:toolforge::k8s::haproxy: Set a global connection limit [puppet] - 10https://gerrit.wikimedia.org/r/1310029 (https://phabricator.wikimedia.org/T431946) (owner: 10Majavah) [10:12:56] (03CR) 10David Caro: [C:03+1] "Just fyi. this is not related to the timeout when waiting for logs with toolforge jobs logs -f (that one still waits for 10min even with t" [puppet] - 10https://gerrit.wikimedia.org/r/1310031 (https://phabricator.wikimedia.org/T431946) (owner: 10Majavah) [10:14:37] (03CR) 10Majavah: [C:03+2] "At the moment my theory is that would've mostly been caused by running out of file descriptors, so the previous patch should help with tha" [puppet] - 10https://gerrit.wikimedia.org/r/1310029 (https://phabricator.wikimedia.org/T431946) (owner: 10Majavah) [10:14:56] (03PS2) 10Majavah: P:toolforge::k8s::haproxy: Set aggressive timeouts on HTTP redirects [puppet] - 10https://gerrit.wikimedia.org/r/1310030 (https://phabricator.wikimedia.org/T431946) [10:14:56] (03PS2) 10Majavah: P:toolforge::k8s::haproxy: Decrease client timeouts [puppet] - 10https://gerrit.wikimedia.org/r/1310031 (https://phabricator.wikimedia.org/T431946) [10:15:23] FIRING: [13x] CertAlmostExpired: gNMI TLS certificate for lsw1-c2-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [10:16:20] (03PS1) 10Majavah: P:toolforge::k8s::haproxy: Fix duplicate HAProxy service [puppet] - 10https://gerrit.wikimedia.org/r/1310032 [10:16:57] (03CR) 10Majavah: "Internal API clients, and the K8s API, is why this patch doesn't adjust the global timeout, but just the HTTP client facing ones." [puppet] - 10https://gerrit.wikimedia.org/r/1310031 (https://phabricator.wikimedia.org/T431946) (owner: 10Majavah) [10:17:05] (03CR) 10CI reject: [V:04-1] P:toolforge::k8s::haproxy: Fix duplicate HAProxy service [puppet] - 10https://gerrit.wikimedia.org/r/1310032 (owner: 10Majavah) [10:17:32] (03PS2) 10Majavah: P:toolforge::k8s::haproxy: Fix duplicate HAProxy service [puppet] - 10https://gerrit.wikimedia.org/r/1310032 [10:17:32] (03CR) 10Majavah: [C:03+2] P:toolforge::k8s::haproxy: Set aggressive timeouts on HTTP redirects [puppet] - 10https://gerrit.wikimedia.org/r/1310030 (https://phabricator.wikimedia.org/T431946) (owner: 10Majavah) [10:18:33] (03CR) 10Majavah: [C:03+2] P:toolforge::k8s::haproxy: Decrease client timeouts [puppet] - 10https://gerrit.wikimedia.org/r/1310031 (https://phabricator.wikimedia.org/T431946) (owner: 10Majavah) [10:19:37] (03CR) 10David Caro: "yep, what I meant is that I tried setting it in both (internal and http client) and did not reduce that 10min timeout on the `toolforge jo" [puppet] - 10https://gerrit.wikimedia.org/r/1310031 (https://phabricator.wikimedia.org/T431946) (owner: 10Majavah) [10:19:54] (03CR) 10Priyankar22: "Thanks for reviewing! Since this logo needs to go live on 15 July 2026," [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309865 (https://phabricator.wikimedia.org/T431094) (owner: 10Priyankar22) [10:20:10] (03CR) 10Majavah: [C:03+2] P:toolforge::k8s::haproxy: Fix duplicate HAProxy service [puppet] - 10https://gerrit.wikimedia.org/r/1310032 (owner: 10Majavah) [10:20:26] (03CR) 10Atsuko: [C:03+2] Revert "opensearch-ttmserver: increase memory to 1x index" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1302089 (owner: 10Atsuko) [10:21:20] (03CR) 10Cathal Mooney: Nokia SR-Linux: add export policy for unicast IBGP clusters (031 comment) [homer/public] - 10https://gerrit.wikimedia.org/r/1309643 (https://phabricator.wikimedia.org/T423430) (owner: 10Cathal Mooney) [10:24:25] (03CR) 10Anzx: [C:03+1] "ok, I may be available to schedule this patch on 15th afternoon backport window" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309865 (https://phabricator.wikimedia.org/T431094) (owner: 10Priyankar22) [10:30:03] (03Merged) 10jenkins-bot: Revert "opensearch-ttmserver: increase memory to 1x index" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1302089 (owner: 10Atsuko) [10:31:08] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host thanos-be2007.codfw.wmnet with OS trixie [10:31:19] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12113962 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host thanos-be2007.codfw.wmnet with OS trixie [10:33:00] !log marostegui@cumin1003 dbctl commit (dc=all): 'Push x4 initial dbctl config T404715', diff saved to https://phabricator.wikimedia.org/P94803 and previous config saved to /var/cache/conftool/dbconfig/20260713-103259-marostegui.json [10:33:03] T404715: Setup x4 section - https://phabricator.wikimedia.org/T404715 [10:34:06] !log atsuko@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply [10:34:14] !log atsuko@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply [10:35:43] !log trueg@deploy2003 helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply [10:35:46] !log trueg@deploy2003 helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply [10:37:07] !log gmodena@deploy2003 helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply [10:37:10] !log gmodena@deploy2003 helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply [10:38:27] (03CR) 10Btullis: [C:03+2] kubernetes: run lvmd as a systemd service on selected dse-k8s workers [puppet] - 10https://gerrit.wikimedia.org/r/1309121 (https://phabricator.wikimedia.org/T429325) (owner: 10Btullis) [10:40:17] (03PS1) 10Nik Gkountas: Update cxserver to 2026-07-10-150240-production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310042 (https://phabricator.wikimedia.org/T290730) [10:40:47] (03PS2) 10Nik Gkountas: Update cxserver to 2026-07-10-150240-production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310042 (https://phabricator.wikimedia.org/T290730) [10:42:49] !log marostegui@cumin1003 dbctl commit (dc=all): 'Change x4 masters T404715', diff saved to https://phabricator.wikimedia.org/P94804 and previous config saved to /var/cache/conftool/dbconfig/20260713-104248-marostegui.json [10:42:52] T404715: Setup x4 section - https://phabricator.wikimedia.org/T404715 [10:43:54] !log btullis@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1023 [10:44:08] !log btullis@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1023 [10:50:41] !log cmooney@cumin1003 START - Cookbook sre.network.host-bgp for host dse-k8s-worker1023 [10:51:28] !log cmooney@cumin1003 END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host dse-k8s-worker1023 [10:52:39] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage [10:53:28] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, July 13 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308656 (https://phabricator.wikimedia.org/T430675) (owner: 10JavierMonton) [10:54:39] 06SRE, 10Observability-Alerting: alertmanager silence confirmation page links to localhost - https://phabricator.wikimedia.org/T328869#12114089 (10hnowlan) [10:57:09] PROBLEM - Docker registry HTTPS interface on registry1005 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/Docker [10:57:59] RECOVERY - Docker registry HTTPS interface on registry1005 is OK: HTTP OK: HTTP/1.1 200 OK - 3746 bytes in 0.233 second response time https://wikitech.wikimedia.org/wiki/Docker [10:59:36] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage [11:00:40] !log marostegui@cumin1003 conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s5 [11:00:53] !log marostegui@cumin1003 conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=x3 [11:04:32] I have two patches to deploy (one private and one on Gerrit). Would anyone mind if I deployed them now? [11:04:44] (03PS2) 10Hnowlan: hadoop: add hdfs alert for HA status [alerts] - 10https://gerrit.wikimedia.org/r/1304769 (https://phabricator.wikimedia.org/T407138) [11:05:25] (03CR) 10Hnowlan: hadoop: add hdfs alert for HA status (031 comment) [alerts] - 10https://gerrit.wikimedia.org/r/1304769 (https://phabricator.wikimedia.org/T407138) (owner: 10Hnowlan) [11:06:02] !log marostegui@cumin1003 conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s8 [11:07:14] Seeing no objections, I'll proceed with deploying [11:07:42] (03CR) 10TrainBranchBot: [C:03+2] "Approved by mszwarc@deploy2003 using scap backport" [extensions/CheckUser] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1310027 (https://phabricator.wikimedia.org/T431977) (owner: 10Mszwarc) [11:08:56] (03CR) 10Hnowlan: [C:03+1] thumbor: Allow 2 connection per worker [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309832 (https://phabricator.wikimedia.org/T431767) (owner: 10Ladsgroup) [11:09:26] (03Merged) 10jenkins-bot: SI: Fix client side instrumentation [extensions/CheckUser] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1310027 (https://phabricator.wikimedia.org/T431977) (owner: 10Mszwarc) [11:09:38] !log mszwarc@deploy2003 Started scap sync-world: Backport for [[gerrit:1310027|SI: Fix client side instrumentation (T431977)]] [11:09:41] T431977: Suggested investigations: Events instrumented client-side are not emitted to event streams - https://phabricator.wikimedia.org/T431977 [11:12:15] Scap exited with error "'mwscript eval.php --wiki aawiki' generated unexpected output: Warning: Undefined array key "x4" in /srv/mediawiki-staging/src/etcd.php on line 135". I see some changes related to x4 have been done earlier, maybe I should better wait for the backport window... [11:13:07] (03PS1) 10Mszwarc: Revert "SI: Fix client side instrumentation" [extensions/CheckUser] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1310055 [11:13:34] (03CR) 10Mszwarc: [C:03+2] "To keep the repo in the same state as prod" [extensions/CheckUser] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1310055 (owner: 10Mszwarc) [11:13:48] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, July 13 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309684 (https://phabricator.wikimedia.org/T416623) (owner: 10Dreamy Jazz) [11:16:10] (03PS2) 10Mszwarc: Revert "SI: Fix client side instrumentation" [extensions/CheckUser] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1310055 (https://phabricator.wikimedia.org/T431977) [11:16:42] (03CR) 10Mszwarc: Revert "SI: Fix client side instrumentation" [extensions/CheckUser] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1310055 (https://phabricator.wikimedia.org/T431977) (owner: 10Mszwarc) [11:16:45] (03CR) 10Mszwarc: [C:03+2] Revert "SI: Fix client side instrumentation" [extensions/CheckUser] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1310055 (https://phabricator.wikimedia.org/T431977) (owner: 10Mszwarc) [11:17:48] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2007.codfw.wmnet with OS trixie [11:17:59] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12114169 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host thanos-be2007.codfw.wmnet with OS trixie compl... [11:22:45] FIRING: SwiftLowContainerAvailability: Swift eqiad container availability low - https://wikitech.wikimedia.org/wiki/Swift/How_To - https://grafana.wikimedia.org/d/OPgmB1Eiz/swift?panelId=8&fullscreen&orgId=1&var-DC=eqiad - https://alerts.wikimedia.org/?q=alertname%3DSwiftLowContainerAvailability [11:24:35] (03Merged) 10jenkins-bot: Revert "SI: Fix client side instrumentation" [extensions/CheckUser] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1310055 (https://phabricator.wikimedia.org/T431977) (owner: 10Mszwarc) [11:26:42] RECOVERY - jenkins_service_running on contint1002 is OK: PROCS OK: 1 process with regex args .*/bin/java .*-jar /usr/share/java/jenkins.war https://wikitech.wikimedia.org/wiki/Jenkins [11:27:01] (03PS1) 10Mszwarc: Revert^2 "SI: Fix client side instrumentation" [extensions/CheckUser] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1310058 (https://phabricator.wikimedia.org/T431977) [11:27:20] RESOLVED: HelmReleaseBadStatus: Helm release wdqs/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [11:27:26] !log gmodena@deploy2003 helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply [11:28:56] !log gmodena@deploy2003 helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply [11:29:42] PROBLEM - jenkins_service_running on contint1002 is CRITICAL: PROCS CRITICAL: 0 processes with regex args .*/bin/java .*-jar /usr/share/java/jenkins.war https://wikitech.wikimedia.org/wiki/Jenkins [11:30:07] !log gmodena@deploy2003 helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply [11:30:42] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [11:31:18] (03PS1) 10Aude: Enable ChartWizard on the beta cluster [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310059 (https://phabricator.wikimedia.org/T431990) [11:31:22] (03CR) 10CWilliams: mysql: update replication source (031 comment) [cookbooks] - 10https://gerrit.wikimedia.org/r/1238368 (https://phabricator.wikimedia.org/T373436) (owner: 10Federico Ceratto) [11:31:45] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, July 13 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310059 (https://phabricator.wikimedia.org/T431990) (owner: 10Aude) [11:32:31] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, July 13 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1307438 (https://phabricator.wikimedia.org/T420883) (owner: 10AikoChou) [11:33:52] !log gmodena@deploy2003 helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply [11:34:28] !log gmodena@deploy2003 helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply [11:34:35] (03PS1) 10Zabe: etcd: Add support for x4 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310061 (https://phabricator.wikimedia.org/T431989) [11:35:45] !log gmodena@deploy2003 helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply [11:37:33] (03CR) 10Ladsgroup: [C:03+1] etcd: Add support for x4 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310061 (https://phabricator.wikimedia.org/T431989) (owner: 10Zabe) [11:37:45] RESOLVED: SwiftLowContainerAvailability: Swift eqiad container availability low - https://wikitech.wikimedia.org/wiki/Swift/How_To - https://grafana.wikimedia.org/d/OPgmB1Eiz/swift?panelId=8&fullscreen&orgId=1&var-DC=eqiad - https://alerts.wikimedia.org/?q=alertname%3DSwiftLowContainerAvailability [11:40:49] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, July 13 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-" [extensions/CheckUser] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1310058 (https://phabricator.wikimedia.org/T431977) (owner: 10Mszwarc) [11:42:49] FIRING: HelmReleaseBadStatus: Helm release wdqs/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [11:43:09] !log gmodena@deploy2003 helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply [11:46:31] jouncebot: nowandnext [11:46:31] No deployments scheduled for the next 1 hour(s) and 13 minute(s) [11:46:31] In 1 hour(s) and 13 minute(s): UTC afternoon backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260713T1300) [11:46:57] thanks zabe. Let me know if I can help on anything [11:47:21] Msz2001: are you done? [11:47:54] I haven't deployed anything eventually, scap failed and I haven't tried again [11:48:17] So, if you're planning to deploy, no obstacles from my side [11:48:39] Alright [11:48:48] (03PS1) 10Zabe: Do not apply elwiki abusefilter settings to dewiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310064 (https://phabricator.wikimedia.org/T431934) [11:49:01] (03CR) 10Zabe: [C:03+2] etcd: Add support for x4 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310061 (https://phabricator.wikimedia.org/T431989) (owner: 10Zabe) [11:49:20] (03CR) 10Zabe: [C:03+2] Do not apply elwiki abusefilter settings to dewiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310064 (https://phabricator.wikimedia.org/T431934) (owner: 10Zabe) [11:50:55] (03Merged) 10jenkins-bot: etcd: Add support for x4 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310061 (https://phabricator.wikimedia.org/T431989) (owner: 10Zabe) [11:51:33] (03Merged) 10jenkins-bot: Do not apply elwiki abusefilter settings to dewiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310064 (https://phabricator.wikimedia.org/T431934) (owner: 10Zabe) [11:51:57] !log gmodena@deploy2003 helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply [11:52:39] (03CR) 10Marostegui: "thanks for the quick fix Zabe!!!" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310061 (https://phabricator.wikimedia.org/T431989) (owner: 10Zabe) [11:52:40] !log zabe@deploy2003 Started scap sync-world: Backport for [[gerrit:1310064|Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061|etcd: Add support for x4 (T431989)]] [11:52:44] T431989: PHP Warning: Undefined array key "x4" - https://phabricator.wikimedia.org/T431989 [11:52:49] RESOLVED: HelmReleaseBadStatus: Helm release wdqs/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [11:53:36] (03PS1) 10Gmodena: wdqs: load qlever index from qlever-index/ S3 prefix [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310066 (https://phabricator.wikimedia.org/T429150) [11:53:59] (03PS1) 10Majavah: P:toolforge::redis_sentinel: Migrate to keepalived::failover [puppet] - 10https://gerrit.wikimedia.org/r/1310067 (https://phabricator.wikimedia.org/T360704) [11:54:02] Thanks zabe [11:54:23] !log zabe@deploy2003 zabe: Backport for [[gerrit:1310064|Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061|etcd: Add support for x4 (T431989)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [11:55:58] !log zabe@deploy2003 zabe: Continuing with deployment [11:56:08] (03CR) 10CI reject: [V:04-1] P:toolforge::redis_sentinel: Migrate to keepalived::failover [puppet] - 10https://gerrit.wikimedia.org/r/1310067 (https://phabricator.wikimedia.org/T360704) (owner: 10Majavah) [11:56:55] (03PS2) 10Majavah: P:toolforge::redis_sentinel: Migrate to keepalived::failover [puppet] - 10https://gerrit.wikimedia.org/r/1310067 (https://phabricator.wikimedia.org/T360704) [11:57:19] (03CR) 10Btullis: [C:03+1] wdqs: load qlever index from qlever-index/ S3 prefix [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310066 (https://phabricator.wikimedia.org/T429150) (owner: 10Gmodena) [11:59:05] (03CR) 10Majavah: [V:03+1] "PCC SUCCESS (CORE_DIFF 5): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/8982/co" [puppet] - 10https://gerrit.wikimedia.org/r/1310067 (https://phabricator.wikimedia.org/T360704) (owner: 10Majavah) [11:59:19] FIRING: [2x] HelmReleaseBadStatus: Helm release wdqs/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [12:00:18] !log zabe@deploy2003 Finished scap sync-world: Backport for [[gerrit:1310064|Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061|etcd: Add support for x4 (T431989)]] (duration: 07m 37s) [12:00:22] T431989: PHP Warning: Undefined array key "x4" - https://phabricator.wikimedia.org/T431989 [12:01:33] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host thanos-be2008.codfw.wmnet with OS trixie [12:01:51] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12114324 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host thanos-be2008.codfw.wmnet with OS trixie [12:02:56] Thanks, zabe! Can I deploy my patch now (the one that failed originally)? [12:03:26] Msz2001: yes:) [12:03:53] !log mszwarc@deploy2003 Started scap sync-world: Backport for [[gerrit:1310027|SI: Fix client side instrumentation (T431977)]] [12:03:56] T431977: Suggested investigations: Events instrumented client-side are not emitted to event streams - https://phabricator.wikimedia.org/T431977 [12:04:18] Thanks zabe! [12:04:22] !log mszwarc@deploy2003 sync-world aborted: Backport for [[gerrit:1310027|SI: Fix client side instrumentation (T431977)]] (duration: 00m 29s) [12:04:32] sorry for the oopsie [12:04:48] yw [12:04:49] (03CR) 10TrainBranchBot: [C:03+2] "Approved by mszwarc@deploy2003 using scap backport" [extensions/CheckUser] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1310058 (https://phabricator.wikimedia.org/T431977) (owner: 10Mszwarc) [12:04:57] although the error rate is not to zero yet [12:05:40] (aborted, because I noticed that I tried to deploy the original patch and not revert^2, now should be the correct one) [12:07:38] (03Merged) 10jenkins-bot: Revert^2 "SI: Fix client side instrumentation" [extensions/CheckUser] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1310058 (https://phabricator.wikimedia.org/T431977) (owner: 10Mszwarc) [12:07:49] !log mszwarc@deploy2003 Started scap sync-world: Backport for [[gerrit:1310058|Revert^2 "SI: Fix client side instrumentation" (T431977)]] [12:09:26] !log mszwarc@deploy2003 mszwarc: Backport for [[gerrit:1310058|Revert^2 "SI: Fix client side instrumentation" (T431977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [12:09:29] T431977: Suggested investigations: Events instrumented client-side are not emitted to event streams - https://phabricator.wikimedia.org/T431977 [12:10:41] !log mszwarc@deploy2003 mszwarc: Continuing with deployment [12:14:02] (03PS1) 10Majavah: P:toolforge::redis_sentinel: Bind on IPv6 [puppet] - 10https://gerrit.wikimedia.org/r/1310072 (https://phabricator.wikimedia.org/T360704) [12:15:04] !log mszwarc@deploy2003 Finished scap sync-world: Backport for [[gerrit:1310058|Revert^2 "SI: Fix client side instrumentation" (T431977)]] (duration: 07m 14s) [12:15:08] T431977: Suggested investigations: Events instrumented client-side are not emitted to event streams - https://phabricator.wikimedia.org/T431977 [12:15:37] I'll deploy one more patch, in a private code [12:16:29] !log gmodena@deploy2003 helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply [12:17:00] !log atsuko@deploy2003 helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply [12:17:05] !log atsuko@deploy2003 helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply [12:19:19] RESOLVED: HelmReleaseBadStatus: Helm release wdqs/scholarly-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [12:19:26] !log gmodena@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply [12:20:47] !log gmodena@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply [12:20:55] (03CR) 10CI reject: [V:04-1] Localisation updates from https://translatewiki.net. [phabricator/translations] (wmf/stable) - 10https://gerrit.wikimedia.org/r/1310074 (owner: 10L10n-bot) [12:22:01] Msz2001: Can you ping me when done? [12:22:06] Sure [12:22:50] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage [12:23:39] !log Deployed changes to private code for Suggested Investigations [12:23:40] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [12:23:46] Dreamy_Jazz: over to you [12:24:11] Thanks [12:26:55] (03PS1) 10Gmodena: wdqs: bump init container resources. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310078 (https://phabricator.wikimedia.org/T429150) [12:27:02] (03CR) 10TrainBranchBot: [C:03+2] "Approved by dreamyjazz@deploy2003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309684 (https://phabricator.wikimedia.org/T416623) (owner: 10Dreamy Jazz) [12:28:15] (03CR) 10Btullis: [C:03+2] admin_ng: configure topolvm storage classes and node placement (032 comments) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309261 (https://phabricator.wikimedia.org/T428235) (owner: 10Btullis) [12:28:36] FIRING: [5x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-3/1/2 (Transport: Hurricane Electric (dc4841.dal5) {#122601}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [12:29:17] (03Abandoned) 10Btullis: datahub-next: reach OpenSearch in plaintext via the mesh [deployment-charts] - 10https://gerrit.wikimedia.org/r/1307073 (https://phabricator.wikimedia.org/T402408) (owner: 10Btullis) [12:29:21] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage [12:29:24] (03CR) 10Brouberol: "Looking at https://grafana.wikimedia.org/d/-D2KNUEGk/kubernetes-pod-details?orgId=1&from=2026-07-13T11%3A41%3A54.528Z&to=2026-07-13T12%3A1" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310078 (https://phabricator.wikimedia.org/T429150) (owner: 10Gmodena) [12:30:31] (03CR) 10Btullis: [C:03+1] archiva: migrate to blackbox HTTP check [puppet] - 10https://gerrit.wikimedia.org/r/1307158 (https://phabricator.wikimedia.org/T407117) (owner: 10Hnowlan) [12:32:19] FIRING: [2x] HelmReleaseBadStatus: Helm release wdqs/scholarly-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [12:32:28] jouncebot: nowandnext [12:32:28] No deployments scheduled for the next 0 hour(s) and 27 minute(s) [12:32:28] In 0 hour(s) and 27 minute(s): UTC afternoon backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260713T1300) [12:32:36] (03CR) 10Btullis: [C:03+1] "Yes the profile is not currently used, but let's keep it around for now, while we decide whether or not to make a second turnilo on bare-m" [puppet] - 10https://gerrit.wikimedia.org/r/1307102 (https://phabricator.wikimedia.org/T407117) (owner: 10Hnowlan) [12:32:40] Dreamy_Jazz: let me know once you're done :D [12:33:06] Sure, CI jobs are taking their time :D [12:33:31] oh okay, I deployed my thumbor change then (it's not mediawiki-config) just wanted to avoid overlap [12:33:48] 👍 [12:33:48] (03CR) 10Ladsgroup: [C:03+2] "let's gooooo" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309832 (https://phabricator.wikimedia.org/T431767) (owner: 10Ladsgroup) [12:34:34] FIRING: [3x] HelmReleaseBadStatus: Helm release wdqs/scholarly-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [12:35:35] 06SRE, 06MediaWiki-Media-Platform-Team, 06ServiceOps new, 10ServiceOps-Services-Oids, 10Thumbor: Thumbor-k8s performance improvements - https://phabricator.wikimedia.org/T333445#12114470 (10Ladsgroup) 05Open→03Resolved Thanks. Closing it then. Further improvements should be tracked in {T431718} [12:35:58] (03Merged) 10jenkins-bot: ProductionServices: Drop NodeJS iPoid URL [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309684 (https://phabricator.wikimedia.org/T416623) (owner: 10Dreamy Jazz) [12:36:08] !log dreamyjazz@deploy2003 Started scap sync-world: Backport for [[gerrit:1309684|ProductionServices: Drop NodeJS iPoid URL (T416623)]] [12:36:13] T416623: Decommission NodeJS IPoid service - https://phabricator.wikimedia.org/T416623 [12:37:47] !log dreamyjazz@deploy2003 dreamyjazz: Backport for [[gerrit:1309684|ProductionServices: Drop NodeJS iPoid URL (T416623)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [12:38:53] !log dreamyjazz@deploy2003 dreamyjazz: Continuing with deployment [12:39:09] (03PS1) 10Majavah: P:toolforge::redis_sentinel: Fix ordering issues [puppet] - 10https://gerrit.wikimedia.org/r/1310086 [12:42:05] I think the CI for deployment charts are a bit unhappy [12:43:06] (03Merged) 10jenkins-bot: admin_ng: configure topolvm storage classes and node placement [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309261 (https://phabricator.wikimedia.org/T428235) (owner: 10Btullis) [12:43:10] !log dreamyjazz@deploy2003 Finished scap sync-world: Backport for [[gerrit:1309684|ProductionServices: Drop NodeJS iPoid URL (T416623)]] (duration: 07m 02s) [12:43:14] T416623: Decommission NodeJS IPoid service - https://phabricator.wikimedia.org/T416623 [12:43:24] (03PS2) 10Gmodena: wdqs: bump init container resources. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310078 (https://phabricator.wikimedia.org/T429150) [12:43:34] (03CR) 10Gmodena: wdqs: bump init container resources. (032 comments) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310078 (https://phabricator.wikimedia.org/T429150) (owner: 10Gmodena) [12:43:37] (03Merged) 10jenkins-bot: thumbor: Allow 2 connection per worker [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309832 (https://phabricator.wikimedia.org/T431767) (owner: 10Ladsgroup) [12:43:53] (03CR) 10Brouberol: [C:03+1] wdqs: bump init container resources. (032 comments) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310078 (https://phabricator.wikimedia.org/T429150) (owner: 10Gmodena) [12:44:14] !log gmodena@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply [12:44:34] (03PS3) 10Gmodena: wdqs: bump init container resources. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310078 (https://phabricator.wikimedia.org/T429150) [12:45:17] !log gmodena@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply [12:45:37] !log ladsgroup@deploy2003 helmfile [staging] START helmfile.d/services/thumbor: apply [12:45:51] !log atsuko@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. [12:45:59] !log ladsgroup@deploy2003 helmfile [staging] DONE helmfile.d/services/thumbor: apply [12:46:45] 10SRE-Access-Requests, 06Data-Engineering: Request to be added to analytics_wmde_users_members - https://phabricator.wikimedia.org/T432002 (10catherine.kelsey.wmde) 03NEW [12:47:18] !log atsuko@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. [12:47:19] RESOLVED: [2x] HelmReleaseBadStatus: Helm release wdqs/main-external on k8s-dse@eqiad in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=eqiad&var-cluster=k8s-dse&var-namespace=wdqs - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [12:47:43] !log ladsgroup@deploy2003 helmfile [staging] START helmfile.d/services/thumbor: apply [12:47:49] !log atsuko@deploy2003 helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. [12:47:58] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2008.codfw.wmnet with OS trixie [12:48:09] !log ladsgroup@deploy2003 helmfile [staging] DONE helmfile.d/services/thumbor: apply [12:48:10] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12114579 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host thanos-be2008.codfw.wmnet with OS trixie compl... [12:48:11] !log atsuko@deploy2003 helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. [12:51:20] !log ladsgroup@deploy2003 helmfile [codfw] START helmfile.d/services/thumbor: apply [12:52:09] (03PS1) 10Mvolz: Update update_version.py to be compatible with ruamel >=0.18.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310090 [12:52:43] !log ladsgroup@deploy2003 helmfile [codfw] DONE helmfile.d/services/thumbor: apply [12:52:45] FIRING: SwiftLowContainerAvailability: Swift eqiad container availability low - https://wikitech.wikimedia.org/wiki/Swift/How_To - https://grafana.wikimedia.org/d/OPgmB1Eiz/swift?panelId=8&fullscreen&orgId=1&var-DC=eqiad - https://alerts.wikimedia.org/?q=alertname%3DSwiftLowContainerAvailability [12:56:23] (03PS2) 10Mvolz: Update update_version.py to be compatible with ruamel >=0.18.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310090 [12:58:07] (03PS2) 10Aude: Enable ChartWizard on the beta cluster [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310059 (https://phabricator.wikimedia.org/T431990) [12:58:08] (03CR) 10Gmodena: [C:03+2] wdqs: bump init container resources. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310078 (https://phabricator.wikimedia.org/T429150) (owner: 10Gmodena) [12:59:37] !log gmodena@deploy2003 helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply [13:00:05] urbanecm and TheresNoTime: May I have your attention please! UTC afternoon backport window. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260713T1300) [13:00:05] gkm563, javiermonton, Dreamy_Jazz, aude, and aiko: A patch you scheduled for UTC afternoon backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [13:00:09] hi [13:00:11] !log ladsgroup@deploy2003 helmfile [eqiad] START helmfile.d/services/thumbor: apply [13:00:34] o/ [13:00:41] o/ [13:00:44] (03PS3) 10Mvolz: Update update_version.py to be compatible with ruamel >=0.18.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310090 [13:01:08] !log gmodena@deploy2003 helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply [13:01:50] (03PS4) 10Mvolz: Update update_version.py to be compatible with ruamel >=0.18.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310090 [13:01:57] !log ladsgroup@deploy2003 helmfile [eqiad] DONE helmfile.d/services/thumbor: apply [13:07:02] (03PS1) 10Bartosz Wójtowicz: ml-services: bump llm-qwen36-27b image and set fp8 KV cache dtype [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310091 (https://phabricator.wikimedia.org/T431978) [13:07:21] (03CR) 10Ayounsi: Nokia SR-Linux: add export policy for unicast IBGP clusters (031 comment) [homer/public] - 10https://gerrit.wikimedia.org/r/1309643 (https://phabricator.wikimedia.org/T423430) (owner: 10Cathal Mooney) [13:07:45] RESOLVED: SwiftLowContainerAvailability: Swift eqiad container availability low - https://wikitech.wikimedia.org/wiki/Swift/How_To - https://grafana.wikimedia.org/d/OPgmB1Eiz/swift?panelId=8&fullscreen&orgId=1&var-DC=eqiad - https://alerts.wikimedia.org/?q=alertname%3DSwiftLowContainerAvailability [13:08:42] (03CR) 10Btullis: [C:03+2] admin_ng: enable the topolvm CSI driver on dse-k8s [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305978 (https://phabricator.wikimedia.org/T429331) (owner: 10Btullis) [13:11:49] FIRING: HelmReleaseBadStatus: Helm release wdqs/scholarly-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [13:12:17] you can self serve? [13:12:23] If noone is around, I can do it [13:12:48] I can deploy mine and maybe include the other patches [13:13:19] sounds good. Let me know if I can help on anything [13:14:24] (03CR) 10Ayounsi: [C:03+1] Nokia SR-Linux: fix issues with BGP policies matching communities [homer/public] - 10https://gerrit.wikimedia.org/r/1308689 (https://phabricator.wikimedia.org/T430810) (owner: 10Cathal Mooney) [13:14:32] aiko JavierMonton I will deploy your patches, but want to make sure you are still aroudn to spot check them [13:15:03] thank you! [13:15:06] thanks aude! [13:15:21] (03CR) 10TrainBranchBot: [C:03+2] "Approved by aude@deploy2003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310059 (https://phabricator.wikimedia.org/T431990) (owner: 10Aude) [13:15:21] (03CR) 10TrainBranchBot: [C:03+2] "Approved by aude@deploy2003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308656 (https://phabricator.wikimedia.org/T430675) (owner: 10JavierMonton) [13:15:21] (03CR) 10TrainBranchBot: [C:03+2] "Approved by aude@deploy2003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1307438 (https://phabricator.wikimedia.org/T420883) (owner: 10AikoChou) [13:15:21] (03CR) 10TrainBranchBot: [C:03+2] "Approved by aude@deploy2003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309894 (https://phabricator.wikimedia.org/T431959) (owner: 10Gkm563) [13:16:49] FIRING: [2x] HelmReleaseBadStatus: Helm release wdqs/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [13:17:01] (03Abandoned) 10Mvolz: Update update_version.py to be compatible with ruamel >=0.18.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310090 (owner: 10Mvolz) [13:17:03] (03Merged) 10jenkins-bot: Enable ChartWizard on the beta cluster [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310059 (https://phabricator.wikimedia.org/T431990) (owner: 10Aude) [13:17:12] (03Merged) 10jenkins-bot: streams: webrequest - pageview - trending [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308656 (https://phabricator.wikimedia.org/T430675) (owner: 10JavierMonton) [13:17:16] (03Merged) 10jenkins-bot: EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1307438 (https://phabricator.wikimedia.org/T420883) (owner: 10AikoChou) [13:17:20] (03Merged) 10jenkins-bot: Remove nonexistent autopatrolled group from Outreach Wiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309894 (https://phabricator.wikimedia.org/T431959) (owner: 10Gkm563) [13:17:33] !log aude@deploy2003 Started scap sync-world: Backport for [[gerrit:1310059|Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656|streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438|EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894|Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] [13:17:43] T431990: Enable ChartWizard on the beta cluster - https://phabricator.wikimedia.org/T431990 [13:17:44] T430675: Rename streams, schemas and applications - https://phabricator.wikimedia.org/T430675 [13:17:45] T420883: Add Wikidata RevertRisk predictions to mediawiki.page_revert_risk_prediction_change - https://phabricator.wikimedia.org/T420883 [13:17:45] T431959: Modify user group configuration settings for Outreach Wiki - https://phabricator.wikimedia.org/T431959 [13:17:58] (03Merged) 10jenkins-bot: admin_ng: enable the topolvm CSI driver on dse-k8s [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305978 (https://phabricator.wikimedia.org/T429331) (owner: 10Btullis) [13:19:14] !log aude@deploy2003 aikochou, javiermonton, aude, gkm563: Backport for [[gerrit:1310059|Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656|streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438|EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894|Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] synced to the testservers [13:19:14] (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [13:19:38] !log gmodena@deploy2003 helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply [13:20:13] (03Restored) 10Mvolz: Update update_version.py to be compatible with ruamel >=0.18.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310090 (owner: 10Mvolz) [13:20:40] aiko JavierMonton please check (if there is anything to check) [13:21:49] FIRING: [2x] HelmReleaseBadStatus: Helm release wdqs/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [13:21:54] mine are just configs, and they appear now in the debug cluster [13:22:05] ok thanks [13:22:10] !log gmodena@deploy2003 helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply [13:22:33] continuing [13:22:38] !log aude@deploy2003 aikochou, javiermonton, aude, gkm563: Continuing with deployment [13:23:11] (03CR) 10Filippo Giunchedi: [C:03+1] P:toolforge::redis_sentinel: Migrate to keepalived::failover [puppet] - 10https://gerrit.wikimedia.org/r/1310067 (https://phabricator.wikimedia.org/T360704) (owner: 10Majavah) [13:23:18] (03CR) 10Filippo Giunchedi: [C:03+1] P:toolforge::redis_sentinel: Bind on IPv6 [puppet] - 10https://gerrit.wikimedia.org/r/1310072 (https://phabricator.wikimedia.org/T360704) (owner: 10Majavah) [13:23:32] mine are configs too [13:23:53] !log gmodena@deploy2003 helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply [13:24:06] (03CR) 10Kevin Bazira: [C:03+1] ml-services: bump llm-qwen36-27b image and set fp8 KV cache dtype [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310091 (https://phabricator.wikimedia.org/T431978) (owner: 10Bartosz Wójtowicz) [13:24:52] (03PS5) 10Mvolz: Update update_version.py to be compatible with ruamel >=0.18.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310090 [13:25:43] (03CR) 10Majavah: [V:03+1 C:03+2] P:toolforge::redis_sentinel: Migrate to keepalived::failover [puppet] - 10https://gerrit.wikimedia.org/r/1310067 (https://phabricator.wikimedia.org/T360704) (owner: 10Majavah) [13:25:53] (03PS6) 10Mvolz: Update update_version.py to be compatible with ruamel >=0.15.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310090 [13:25:54] (03CR) 10Majavah: [C:03+2] P:toolforge::redis_sentinel: Bind on IPv6 [puppet] - 10https://gerrit.wikimedia.org/r/1310072 (https://phabricator.wikimedia.org/T360704) (owner: 10Majavah) [13:26:05] (03CR) 10Majavah: [C:03+2] P:toolforge::redis_sentinel: Fix ordering issues [puppet] - 10https://gerrit.wikimedia.org/r/1310086 (owner: 10Majavah) [13:26:41] (03CR) 10Mvolz: "Wasn't quite sure who to add as a reviewer, sorry for spam if not relevant." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310090 (owner: 10Mvolz) [13:26:49] RESOLVED: [2x] HelmReleaseBadStatus: Helm release wdqs/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [13:27:44] there are some errors [96 hits] PHP Warning: Undefined array key "x4" [96 hits] PHP Warning: foreach() argument must be of type array|object, null given but think they are unrelated [13:27:52] Amir1 ^ [13:28:19] yeah, it was fixed by za.be but it takes time for everything to flush [13:28:23] like maint scripts, etc. [13:28:35] ok thanks, I will continue [13:28:45] !log aude@deploy2003 Finished scap sync-world: Backport for [[gerrit:1310059|Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656|streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438|EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894|Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] (duration: 11m 12s) [13:28:55] T431990: Enable ChartWizard on the beta cluster - https://phabricator.wikimedia.org/T431990 [13:28:55] T430675: Rename streams, schemas and applications - https://phabricator.wikimedia.org/T430675 [13:28:56] T420883: Add Wikidata RevertRisk predictions to mediawiki.page_revert_risk_prediction_change - https://phabricator.wikimedia.org/T420883 [13:28:56] T431959: Modify user group configuration settings for Outreach Wiki - https://phabricator.wikimedia.org/T431959 [13:29:13] these were occurring before the deploy [13:30:24] !log ladsgroup@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1070.eqiad.wmnet with reason: vacuum overlarge container dbs [13:30:36] 06SRE, 10SRE-swift-storage: Disk near-full warnings on ms swift backends for container filesystems due to some bloated sqlite files - https://phabricator.wikimedia.org/T377827#12114836 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=8dc7cbdc-eccc-43ce-a361-b8b404665284) set by ladsgroup@cum... [13:31:17] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host thanos-be2009.codfw.wmnet with OS trixie [13:31:29] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12114842 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host thanos-be2009.codfw.wmnet with OS trixie [13:32:01] (03CR) 10Mvolz: "oh nevermind, publish is still failing, so this is needed! see https://integration.wikimedia.org/ci/job/citoid-pipeline-publish/215/consol" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310090 (owner: 10Mvolz) [13:33:36] (03PS1) 10Joal: Revert webrequest (+other) retention to 90 days [puppet] - 10https://gerrit.wikimedia.org/r/1310099 (https://phabricator.wikimedia.org/T431968) [13:33:49] !log gmodena@deploy2003 helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply [13:34:33] (03CR) 10Bartosz Wójtowicz: [C:03+2] ml-services: bump llm-qwen36-27b image and set fp8 KV cache dtype [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310091 (https://phabricator.wikimedia.org/T431978) (owner: 10Bartosz Wójtowicz) [13:36:30] (03CR) 10Kamila Součková: [C:03+1] "+1 to a temporary revert so it is not blocking releases" [puppet] - 10https://gerrit.wikimedia.org/r/1309689 (https://phabricator.wikimedia.org/T428909) (owner: 10Scott French) [13:38:30] (03CR) 10Brouberol: [C:03+1] Revert webrequest (+other) retention to 90 days [puppet] - 10https://gerrit.wikimedia.org/r/1310099 (https://phabricator.wikimedia.org/T431968) (owner: 10Joal) [13:39:19] (03Merged) 10jenkins-bot: ml-services: bump llm-qwen36-27b image and set fp8 KV cache dtype [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310091 (https://phabricator.wikimedia.org/T431978) (owner: 10Bartosz Wójtowicz) [13:40:40] !log bwojtowicz@deploy2003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . [13:40:41] (03CR) 10Hnowlan: [C:03+2] turnilo: migrate to use blackbox check to replace nrpe check [puppet] - 10https://gerrit.wikimedia.org/r/1307102 (https://phabricator.wikimedia.org/T407117) (owner: 10Hnowlan) [13:44:28] (03PS1) 10Effie Mouzeli: WIP: exclude rest.php from debug routing in some cases [puppet] - 10https://gerrit.wikimedia.org/r/1310103 [13:45:20] (03Abandoned) 10Hnowlan: turnilo: remove all monitoring cruft [puppet] - 10https://gerrit.wikimedia.org/r/1307103 (https://phabricator.wikimedia.org/T407117) (owner: 10Hnowlan) [13:45:31] (03PS2) 10Effie Mouzeli: WIP: exclude rest.php from debug routing in some cases [puppet] - 10https://gerrit.wikimedia.org/r/1310103 [13:47:45] !log rscout@deploy2003 helmfile [codfw] START helmfile.d/services/miscweb: apply [13:48:07] !log rscout@deploy2003 helmfile [codfw] DONE helmfile.d/services/miscweb: apply [13:48:21] !log rscout@deploy2003 helmfile [eqiad] START helmfile.d/services/miscweb: apply [13:48:34] (03CR) 10Ssingh: [C:03+1] hieradata: issue fixes [puppet] - 10https://gerrit.wikimedia.org/r/1308150 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [13:48:44] !log rscout@deploy2003 helmfile [eqiad] DONE helmfile.d/services/miscweb: apply [13:50:39] (03PS4) 10Cathal Mooney: Adjust order BGP_outfilter is applied to routes announced [homer/public] - 10https://gerrit.wikimedia.org/r/1309707 (https://phabricator.wikimedia.org/T431849) [13:52:03] (03CR) 10Cathal Mooney: "ACK yeah not sure what the best way to do that might be. I guess we can just manually reset a peer and see how long it takes before we ge" [homer/public] - 10https://gerrit.wikimedia.org/r/1309707 (https://phabricator.wikimedia.org/T431849) (owner: 10Cathal Mooney) [13:52:12] (03PS1) 10Ladsgroup: cache: Reduce webp threshold to 90 [puppet] - 10https://gerrit.wikimedia.org/r/1310105 (https://phabricator.wikimedia.org/T431150) [13:52:37] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage [13:55:46] (03CR) 10CDobbins: [V:03+1 C:03+2] hieradata: issue fixes [puppet] - 10https://gerrit.wikimedia.org/r/1308150 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [13:57:45] !log cdobbins@cumin2002 conftool action : set/pooled=no; selector: name=dns7002.* [13:58:24] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage [13:58:33] (03PS1) 10Marostegui: wmnet: Failover m2-master [dns] - 10https://gerrit.wikimedia.org/r/1310106 (https://phabricator.wikimedia.org/T431660) [14:00:05] !log ladsgroup@cumin1003 START - Cookbook sre.hosts.remove-downtime for ms-be1070.eqiad.wmnet [14:00:06] !log ladsgroup@cumin1003 END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1070.eqiad.wmnet [14:00:19] (03CR) 10Marostegui: [C:03+2] wmnet: Failover m2-master [dns] - 10https://gerrit.wikimedia.org/r/1310106 (https://phabricator.wikimedia.org/T431660) (owner: 10Marostegui) [14:00:28] !log marostegui@dns1004 START - running authdns-update [14:01:41] 06SRE, 10SRE-swift-storage: Disk near-full warnings on ms swift backends for container filesystems due to some bloated sqlite files - https://phabricator.wikimedia.org/T377827#12114989 (10Ladsgroup) {F93397285} {F93397224} [14:02:18] !log marostegui@dns1004 END - running authdns-update [14:05:06] !log cdobbins@cumin2002 START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie [14:05:17] !log disable-puppet on A:cp for ATS config change - T428909 T431838 [14:05:22] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:05:22] T428909: Investigate X-Wikimedia-Debug routing for api.wikimedia.org - https://phabricator.wikimedia.org/T428909 [14:05:23] T431838: mw-parsoid endpoints are returning 404 for the rest.php endpoints used by rt-testing - https://phabricator.wikimedia.org/T431838 [14:06:25] PROBLEM - Docker registry HTTPS interface on registry1005 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/Docker [14:07:15] RECOVERY - Docker registry HTTPS interface on registry1005 is OK: HTTP OK: HTTP/1.1 200 OK - 3746 bytes in 0.188 second response time https://wikitech.wikimedia.org/wiki/Docker [14:08:14] (03CR) 10Scott French: [C:03+2] "Deploying now in order to unblock RT testing." [puppet] - 10https://gerrit.wikimedia.org/r/1309689 (https://phabricator.wikimedia.org/T428909) (owner: 10Scott French) [14:09:05] PROBLEM - BFD status on asw1-b4-magru.mgmt is CRITICAL: Down: 2 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [14:10:10] FIRING: [2x] BFDdown: BFD session down between asw1-b4-magru and 195.200.68.37 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=asw1-b4-magru:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [14:14:37] !log start rolling run-puppet-agent on A:cp for ATS config change - T428909 T431838 [14:14:41] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:14:42] T428909: Investigate X-Wikimedia-Debug routing for api.wikimedia.org - https://phabricator.wikimedia.org/T428909 [14:14:42] T431838: mw-parsoid endpoints are returning 404 for the rest.php endpoints used by rt-testing - https://phabricator.wikimedia.org/T431838 [14:15:38] FIRING: CertAlmostExpired: gNMI TLS certificate for lsw1-f6-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [14:15:42] FIRING: [3x] JobUnavailable: Reduced availability for job haproxy in ops@magru - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [14:16:05] PROBLEM - Postgres Replication Lag on puppetdb2003 is CRITICAL: POSTGRES_HOT_STANDBY_DELAY CRITICAL: DB puppetdb (host:localhost) 308429688 and 25 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [14:16:34] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2009.codfw.wmnet with OS trixie [14:16:51] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12115032 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host thanos-be2009.codfw.wmnet with OS trixie compl... [14:18:01] !log btullis@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. [14:18:27] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12115039 (10MatthewVernon) [14:18:34] 10ops-codfw, 06SRE, 06DBA, 06DC-Ops: db2209 firmware upgrade after a crash - https://phabricator.wikimedia.org/T431952#12115041 (10Jhancock.wm) a:03Jhancock.wm working on this. [14:20:05] RECOVERY - Postgres Replication Lag on puppetdb2003 is OK: POSTGRES_HOT_STANDBY_DELAY OK: DB puppetdb (host:localhost) 2526896 and 0 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [14:21:45] FIRING: SwiftLowContainerAvailability: Swift eqiad container availability low - https://wikitech.wikimedia.org/wiki/Swift/How_To - https://grafana.wikimedia.org/d/OPgmB1Eiz/swift?panelId=8&fullscreen&orgId=1&var-DC=eqiad - https://alerts.wikimedia.org/?q=alertname%3DSwiftLowContainerAvailability [14:23:52] !log marostegui@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1023.eqiad.wmnet with reason: reboot [14:24:54] (03PS1) 10Marostegui: Revert "wmnet: Failover m2-master" [dns] - 10https://gerrit.wikimedia.org/r/1310111 [14:27:00] (03CR) 10Marostegui: [C:03+2] Revert "wmnet: Failover m2-master" [dns] - 10https://gerrit.wikimedia.org/r/1310111 (owner: 10Marostegui) [14:27:05] !log marostegui@dns1004 START - running authdns-update [14:28:23] !log btullis@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. [14:28:54] !log marostegui@dns1004 END - running authdns-update [14:28:57] (03PS1) 10Marostegui: check_private_data_report: Add clouddb1029 [puppet] - 10https://gerrit.wikimedia.org/r/1310112 (https://phabricator.wikimedia.org/T409557) [14:29:41] (03CR) 10Marostegui: [C:03+2] check_private_data_report: Add clouddb1029 [puppet] - 10https://gerrit.wikimedia.org/r/1310112 (https://phabricator.wikimedia.org/T409557) (owner: 10Marostegui) [14:29:54] !log cdobbins@cumin2002 START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage [14:30:04] Deploy window Test Kitchen Experiment Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260713T1430) [14:30:54] (03CR) 10Majavah: [C:03+1] wmcs: switch to production systemd unit alerting [alerts] - 10https://gerrit.wikimedia.org/r/1300745 (https://phabricator.wikimedia.org/T428873) (owner: 10Filippo Giunchedi) [14:32:46] !log jgiannelos@deploy2003 helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply [14:33:26] !log jgiannelos@deploy2003 helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply [14:33:28] !log jgiannelos@deploy2003 helmfile [codfw] START helmfile.d/services/mw-parsoid: apply [14:34:04] !log jgiannelos@deploy2003 helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply [14:34:38] (03PS1) 10Btullis: topolvm: run the node plugin pod as root [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310116 (https://phabricator.wikimedia.org/T428895) [14:34:59] (03CR) 10TChin: [C:03+1] stream: pageview-trending-relative [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309908 (https://phabricator.wikimedia.org/T430675) (owner: 10JavierMonton) [14:35:02] (03PS1) 10Atsuko: service: change pybal check for cloudelastic-psi [puppet] - 10https://gerrit.wikimedia.org/r/1310117 (https://phabricator.wikimedia.org/T431538) [14:35:27] !log cdobbins@cumin2002 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage [14:35:44] (03CR) 10Hnowlan: [C:03+2] archiva: migrate to blackbox HTTP check [puppet] - 10https://gerrit.wikimedia.org/r/1307158 (https://phabricator.wikimedia.org/T407117) (owner: 10Hnowlan) [14:35:55] (03CR) 10TChin: [C:03+1] topic: webrequest-page-view-next [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310025 (https://phabricator.wikimedia.org/T430675) (owner: 10JavierMonton) [14:36:45] RESOLVED: SwiftLowContainerAvailability: Swift eqiad container availability low - https://wikitech.wikimedia.org/wiki/Swift/How_To - https://grafana.wikimedia.org/d/OPgmB1Eiz/swift?panelId=8&fullscreen&orgId=1&var-DC=eqiad - https://alerts.wikimedia.org/?q=alertname%3DSwiftLowContainerAvailability [14:42:03] PROBLEM - Recursive DNS on 195.200.68.37 is CRITICAL: DNS_QUERY CRITICAL - query timed out https://wikitech.wikimedia.org/wiki/DNS [14:43:16] (03CR) 10Ssingh: [C:03+1] service: change pybal check for cloudelastic-psi [puppet] - 10https://gerrit.wikimedia.org/r/1310117 (https://phabricator.wikimedia.org/T431538) (owner: 10Atsuko) [14:47:03] PROBLEM - Recursive DNS on 2a02:ec80:700:2:195:200:68:37 is CRITICAL: DNS_QUERY CRITICAL - query timed out https://wikitech.wikimedia.org/wiki/DNS [14:47:18] (03CR) 10Brouberol: [C:03+1] topolvm: run the node plugin pod as root [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310116 (https://phabricator.wikimedia.org/T428895) (owner: 10Btullis) [14:48:54] (03CR) 10Btullis: [C:03+2] topolvm: run the node plugin pod as root [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310116 (https://phabricator.wikimedia.org/T428895) (owner: 10Btullis) [14:53:42] 06SRE, 10SRE-Access-Requests, 06Data-Engineering: Request to be added to analytics_wmde_users_members - https://phabricator.wikimedia.org/T432002#12115215 (10jcrespo) p:05Triage→03High [14:55:42] FIRING: [3x] JobUnavailable: Reduced availability for job haproxy in ops@magru - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [14:55:47] (03PS1) 10Elukey: profile::puppetserver::volatile: add metrics for airflow webrequest [puppet] - 10https://gerrit.wikimedia.org/r/1310118 (https://phabricator.wikimedia.org/T402512) [14:56:15] (03CR) 10Elukey: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1310118 (https://phabricator.wikimedia.org/T402512) (owner: 10Elukey) [14:58:35] (03Merged) 10jenkins-bot: topolvm: run the node plugin pod as root [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310116 (https://phabricator.wikimedia.org/T428895) (owner: 10Btullis) [14:58:58] (03PS2) 10Elukey: profile::puppetserver::volatile: add metrics for airflow webrequest [puppet] - 10https://gerrit.wikimedia.org/r/1310118 (https://phabricator.wikimedia.org/T402512) [15:01:08] !log btullis@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. [15:01:34] !log btullis@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. [15:02:24] (03PS1) 10Ayounsi: Add profile::server_depool for centrallog and graphite [puppet] - 10https://gerrit.wikimedia.org/r/1310119 (https://phabricator.wikimedia.org/T327300) [15:03:56] (03CR) 10Ayounsi: "From the comment on https://phabricator.wikimedia.org/T430918#12114149" [puppet] - 10https://gerrit.wikimedia.org/r/1310119 (https://phabricator.wikimedia.org/T327300) (owner: 10Ayounsi) [15:06:44] !log btullis@deploy2003 helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. [15:06:57] 10SRE-Access-Requests, 06Data-Engineering: Request to be added to analytics_wmde_users_members - https://phabricator.wikimedia.org/T432002#12115301 (10jcrespo) Hi, @catherine.kelsey.wmde, I see you already have access and a NDA filed with WMF. I am not aware of any LDAP group with that name (I checked and it d... [15:08:13] (03CR) 10Elukey: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1310118 (https://phabricator.wikimedia.org/T402512) (owner: 10Elukey) [15:08:38] !log btullis@deploy2003 helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. [15:08:52] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 - https://phabricator.wikimedia.org/T431682#12115314 (10fnegri) [15:08:55] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Ensure cloudvirt capacity is more evenly spread out among racks - https://phabricator.wikimedia.org/T424658#12115317 (10fnegri) [15:09:00] RECOVERY - Recursive DNS on 2a02:ec80:700:2:195:200:68:37 is OK: DNS_QUERY OK - Success https://wikitech.wikimedia.org/wiki/DNS [15:09:00] RECOVERY - Recursive DNS on 195.200.68.37 is OK: DNS_QUERY OK - Success https://wikitech.wikimedia.org/wiki/DNS [15:09:59] (03PS1) 10Jcrespo: admin: Add catherinekelsey to the analytics-wmde-users group [puppet] - 10https://gerrit.wikimedia.org/r/1310121 (https://phabricator.wikimedia.org/T432002) [15:10:42] FIRING: [2x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [15:11:55] 10SRE-Access-Requests, 06Data-Engineering, 13Patch-For-Review: Request to be added to analytics_wmde_users_members - https://phabricator.wikimedia.org/T432002#12115362 (10catherine.kelsey.wmde) Hey @jcrespo - thanks for your speedy response! Yes exactly - https://phabricator.wikimedia.org/T215938 - this pre... [15:16:54] 10SRE-SLO, 06Abstract Wikipedia team, 06ServiceOps new, 07Essential-Work: wikifunctions-backend-combined-v1 SLI error budget has been rapidly dropping over Feb 2026 - https://phabricator.wikimedia.org/T418160#12115401 (10Jdforrester-WMF) 05Open→03Resolved a:03Jdforrester-WMF Oops, we never closed... [15:17:13] (03CR) 10AikoChou: [C:03+2] ml-services: enable revertrisk-wikidata event stream predictions [deployment-charts] - 10https://gerrit.wikimedia.org/r/1307441 (https://phabricator.wikimedia.org/T420883) (owner: 10AikoChou) [15:18:23] 10SRE-Access-Requests, 06Data-Engineering, 13Patch-For-Review: Add catherinekelsey to analytics-wmde-users group - https://phabricator.wikimedia.org/T432002#12115407 (10jcrespo) [15:19:04] 10SRE-Access-Requests, 06Data-Engineering, 13Patch-For-Review: Add catherinekelsey to analytics-wmde-users group - https://phabricator.wikimedia.org/T432002#12115412 (10jcrespo) [15:19:48] (03Merged) 10jenkins-bot: ml-services: enable revertrisk-wikidata event stream predictions [deployment-charts] - 10https://gerrit.wikimedia.org/r/1307441 (https://phabricator.wikimedia.org/T420883) (owner: 10AikoChou) [15:21:18] (03CR) 10Jcrespo: [C:04-1] "No manager or service owner approvals yet." [puppet] - 10https://gerrit.wikimedia.org/r/1310121 (https://phabricator.wikimedia.org/T432002) (owner: 10Jcrespo) [15:25:17] (03PS3) 10Effie Mouzeli: WIP: exclude rest.php from debug routing in some cases [puppet] - 10https://gerrit.wikimedia.org/r/1310103 [15:26:27] 10SRE-Access-Requests, 06Data-Engineering, 13Patch-For-Review: Add catherinekelsey to analytics-wmde-users group - https://phabricator.wikimedia.org/T432002#12115472 (10jcrespo) >>! In T432002#12115362, @catherine.kelsey.wmde wrote: > Hey @jcrespo - thanks for your speedy response! > > Yes exactly - https:/... [15:27:14] PROBLEM - Ensure traffic_server is running for instance backend on cp3077 is CRITICAL: PROCS CRITICAL: 3 processes with args /usr/bin/traffic_server -M --httpport 3128 https://wikitech.wikimedia.org/wiki/Apache_Traffic_Server [15:27:26] uh? [15:27:46] brett: ^ do we know anything about this host? [15:27:52] 06SRE, 10Horizon, 10Prod-Kubernetes, 06ServiceOps new, and 2 others: Move cloudweb to Ganeti VMs and repurpose the servers as wikikube nodes - https://phabricator.wikimedia.org/T392478#12115482 (10fnegri) [15:28:05] seems to be running fine though? [15:28:14] RECOVERY - Ensure traffic_server is running for instance backend on cp3077 is OK: PROCS OK: 1 process with args /usr/bin/traffic_server -M --httpport 3128 https://wikitech.wikimedia.org/wiki/Apache_Traffic_Server [15:28:27] ha [15:28:55] 06SRE-OnFire, 10Cloud-VPS, 06tools-infrastructure-team, 13Patch-For-Review, 07Sustainability (Incident Followup): Add external meta-monitoring for metricsinfra - https://phabricator.wikimedia.org/T288053#12115498 (10fnegri) [15:29:51] (03PS1) 10Btullis: Correct typo in hiera for an-test-coord1002 [puppet] - 10https://gerrit.wikimedia.org/r/1310122 (https://phabricator.wikimedia.org/T406746) [15:30:05] jan_drewniak: That opportune time for a Wikimedia Portals Update deploy is upon us again. Don't be afraid. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260713T1530). [15:32:31] FIRING: Traffic bill over quota: Alert for device cr2-eqsin.wikimedia.org - Traffic bill over quota - https://alerts.wikimedia.org/?q=alertname%3DTraffic+bill+over+quota [15:32:48] (03PS4) 10Effie Mouzeli: WIP: exclude rest.php from debug routing in some cases [puppet] - 10https://gerrit.wikimedia.org/r/1310103 [15:33:10] PROBLEM - Postgres Replication Lag on puppetdb2003 is CRITICAL: POSTGRES_HOT_STANDBY_DELAY CRITICAL: DB puppetdb (host:localhost) 82593928 and 10 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [15:34:01] !log elukey@cumin1003 START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet [15:34:08] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet [15:34:10] RECOVERY - Postgres Replication Lag on puppetdb2003 is OK: POSTGRES_HOT_STANDBY_DELAY OK: DB puppetdb (host:localhost) 2760320 and 0 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [15:36:24] !log restart pybal on lvs1020 [15:36:25] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [15:40:08] !log elukey@cumin1003 START - Cookbook sre.hosts.bmc-user-mgmt for host netbox2003.codfw.wmnet [15:40:09] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox2003.codfw.wmnet [15:40:27] !log elukey@cumin1003 START - Cookbook sre.hosts.bmc-user-mgmt for host netbox1003.eqiad.wmnet [15:40:28] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox1003.eqiad.wmnet [15:40:56] !log elukey@cumin1003 START - Cookbook sre.hosts.bmc-user-mgmt for host ganeti-test[2001-2003].codfw.wmnet [15:41:40] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host ganeti-test[2001-2003].codfw.wmnet [15:42:15] (03CR) 10Atsuko: [C:03+2] service: change pybal check for cloudelastic-psi [puppet] - 10https://gerrit.wikimedia.org/r/1310117 (https://phabricator.wikimedia.org/T431538) (owner: 10Atsuko) [15:42:21] (03PS5) 10Effie Mouzeli: WIP: exclude rest.php from debug routing in some cases [puppet] - 10https://gerrit.wikimedia.org/r/1310103 [15:43:05] !log elukey@cumin1003 START - Cookbook sre.hosts.bmc-user-mgmt for host kafka-logging1006.eqiad.wmnet [15:43:15] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host kafka-logging1006.eqiad.wmnet [15:46:47] !log aikochou@deploy2003 helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . [15:48:59] 10ops-codfw, 06SRE, 06DBA, 06DC-Ops: db2209 firmware upgrade after a crash - https://phabricator.wikimedia.org/T431952#12115591 (10Jhancock.wm) i think i pinged the wrong person @Marostegui but it's updated now. sorry about that. [15:50:25] !log restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310117 [15:50:26] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [15:50:37] 10ops-codfw, 06SRE, 06DBA, 06DC-Ops: db2209 firmware upgrade after a crash - https://phabricator.wikimedia.org/T431952#12115598 (10Marostegui) Thanks! I'll take it from here [15:52:09] (03PS2) 10Esanders: Enable DiscussionTools visual enhancements on enwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1212157 (https://phabricator.wikimedia.org/T409297) [15:52:14] sukhe: Looked into cp3077 and I don't see anything wrong. bad check? [15:52:31] RESOLVED: Traffic bill over quota: Alert for device cr2-eqsin.wikimedia.org - Traffic bill over quota - https://alerts.wikimedia.org/?q=alertname%3DTraffic+bill+over+quota [15:52:57] brett: yeah probably. sorry, forgot to update you here [15:53:01] nothing to do here unless you find something wrong [15:55:48] !log aikochou@deploy2003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . [15:59:51] !log restarting pybal on lvs1018 for https://gerrit.wikimedia.org/r/1310117 [15:59:52] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [16:07:52] PROBLEM - OpenSearch health check for shards on 9600 on cloudelastic1007 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:10:42] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:13:46] RECOVERY - OpenSearch health check for shards on 9600 on cloudelastic1007 is OK: OK - elasticsearch status cloudelastic-psi-eqiad: cluster_name: cloudelastic-psi-eqiad, status: yellow, timed_out: False, number_of_nodes: 6, number_of_data_nodes: 6, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 817, active_shards: 1555, relocating_shards: 0, initializing_shards: 9, unassigned_shards: 74, delayed_unassigne [16:13:46] : 0, number_of_pending_tasks: 6, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 936, active_shards_percent_as_number: 94.93284493284493 https://wikitech.wikimedia.org/wiki/Search%23Administration [16:15:42] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:17:48] !log javiermonton@deploy2003 helmfile [eqiad] START helmfile.d/services/eventgate-main: sync [16:18:14] !log javiermonton@deploy2003 helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync [16:19:15] !log javiermonton@deploy2003 helmfile [codfw] START helmfile.d/services/eventgate-main: sync [16:19:37] !log javiermonton@deploy2003 helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync [16:21:01] !log javiermonton@deploy2003 helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync [16:21:32] !log javiermonton@deploy2003 helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync [16:21:59] !log javiermonton@deploy2003 helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync [16:22:32] !log javiermonton@deploy2003 helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync [16:22:54] (03PS1) 10Atsuko: service: change pybal check for opensearch clusters [puppet] - 10https://gerrit.wikimedia.org/r/1310129 (https://phabricator.wikimedia.org/T431538) [16:25:59] (03CR) 10Atsuko: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1310129 (https://phabricator.wikimedia.org/T431538) (owner: 10Atsuko) [16:26:33] !log javiermonton@deploy2003 helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: sync [16:26:56] !log javiermonton@deploy2003 helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: sync [16:26:59] (03PS13) 10Federico Ceratto: sre.mysql: split pool/depool [cookbooks] - 10https://gerrit.wikimedia.org/r/1295480 (https://phabricator.wikimedia.org/T422361) [16:27:05] !log javiermonton@deploy2003 helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: sync [16:27:26] !log javiermonton@deploy2003 helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: sync [16:27:40] (03PS1) 10Elukey: redfish: add delete_account method [software/spicerack] - 10https://gerrit.wikimedia.org/r/1310131 (https://phabricator.wikimedia.org/T426180) [16:27:53] !log dancy@deploy2003 Installing scap version "4.273.0" for 159 host(s) [16:28:36] FIRING: [5x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-3/1/2 (Transport: Hurricane Electric (dc4841.dal5) {#122601}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [16:28:45] (03CR) 10Elukey: "Tested with spicerack-shell and the PYTHONPATH trick. The only weird thing that I found is that on a host the first time i got a HTTP 400 " [software/spicerack] - 10https://gerrit.wikimedia.org/r/1310131 (https://phabricator.wikimedia.org/T426180) (owner: 10Elukey) [16:29:49] !log ladsgroup@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1065.eqiad.wmnet with reason: vacuum overlarge container dbs [16:29:59] 06SRE, 10SRE-swift-storage: Disk near-full warnings on ms swift backends for container filesystems due to some bloated sqlite files - https://phabricator.wikimedia.org/T377827#12115907 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=67b2b7ef-499a-4a70-9dc5-3f16c4974d6d) set by ladsgroup@cum... [16:31:16] (03CR) 10Elukey: "Cole, Hugh: Hi! Let me know if this is the approach that you suggested on IRC, it seems sound to me but I may be missing something. Thanks" [puppet] - 10https://gerrit.wikimedia.org/r/1310118 (https://phabricator.wikimedia.org/T402512) (owner: 10Elukey) [16:31:54] !log dancy@deploy2003 Installation of scap version "4.273.0" completed for 159 hosts [16:32:03] (03CR) 10CI reject: [V:04-1] redfish: add delete_account method [software/spicerack] - 10https://gerrit.wikimedia.org/r/1310131 (https://phabricator.wikimedia.org/T426180) (owner: 10Elukey) [16:32:40] (03CR) 10Federico Ceratto: sre.mysql: split pool/depool (031 comment) [cookbooks] - 10https://gerrit.wikimedia.org/r/1295480 (https://phabricator.wikimedia.org/T422361) (owner: 10Federico Ceratto) [16:33:30] !log fceratto@cumin1003 START - Cookbook sre.mysql.depool depool pc1023: Depool test [16:33:30] !log fceratto@cumin1003 START - Cookbook sre.mysql.parsercache [16:33:40] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) [16:33:40] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test [16:34:14] (03PS2) 10Elukey: redfish: add delete_account method [software/spicerack] - 10https://gerrit.wikimedia.org/r/1310131 (https://phabricator.wikimedia.org/T426180) [16:34:33] !log fceratto@cumin1003 START - Cookbook sre.mysql.pool pool pc1023: Pool test [16:34:33] !log fceratto@cumin1003 START - Cookbook sre.mysql.parsercache [16:34:49] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) [16:34:49] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test [16:36:04] PROBLEM - Bird Internet Routing Daemon on dns7002 is CRITICAL: PROCS CRITICAL: 0 processes with command name bird https://wikitech.wikimedia.org/wiki/Anycast%23Bird_daemon_not_running [16:39:11] (03CR) 10CI reject: [V:04-1] redfish: add delete_account method [software/spicerack] - 10https://gerrit.wikimedia.org/r/1310131 (https://phabricator.wikimedia.org/T426180) (owner: 10Elukey) [16:40:14] PROBLEM - NTP peers and stratum check on dns7002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/NTP [16:42:04] (03CR) 10Hnowlan: "Overall approach lgtm! Just one minor concern" [puppet] - 10https://gerrit.wikimedia.org/r/1310118 (https://phabricator.wikimedia.org/T402512) (owner: 10Elukey) [16:42:24] !log mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model damaging --old (T431159) [16:42:26] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [16:42:27] T431159: ORES flagging not working on simplewiki - https://phabricator.wikimedia.org/T431159 [16:42:35] !log jgiannelos@deploy2003 helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply [16:42:38] !log jgiannelos@deploy2003 helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply [16:42:39] !log jgiannelos@deploy2003 helmfile [codfw] START helmfile.d/services/mw-parsoid: apply [16:42:42] !log jgiannelos@deploy2003 helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply [16:45:03] !log restarting pybal on lvs1019 to flush IP address for `cirrussearch1122.eqiad.wmnet` after moving the vlan T431311 [16:45:06] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [16:45:07] T431311: Migrate production OpenSearch clusters from 1.x-2.x - EQIAD - https://phabricator.wikimedia.org/T431311 [16:47:08] (03CR) 10JHathaway: [C:03+1] redfish: skip setting AccountTypes if the admin user doesn't expose it (031 comment) [software/spicerack] - 10https://gerrit.wikimedia.org/r/1309120 (https://phabricator.wikimedia.org/T426180) (owner: 10Elukey) [16:47:12] jouncebot now [16:47:12] No deployments scheduled for the next 0 hour(s) and 12 minute(s) [16:47:30] !log dancy@deploy2003 Started scap sync-world: testing T428971 [16:47:33] T428971: Allow configuration of canary and production checks based on deployment target - https://phabricator.wikimedia.org/T428971 [16:48:00] * swfrench-wmf is excited to see this deployment [16:49:49] (03CR) 10JHathaway: [C:03+1] redfish: add delete_account method [software/spicerack] - 10https://gerrit.wikimedia.org/r/1310131 (https://phabricator.wikimedia.org/T426180) (owner: 10Elukey) [16:51:07] !log dancy@deploy2003 Finished scap sync-world: testing T428971 (duration: 03m 37s) [16:51:45] (03CR) 10Scott French: "Thanks for the review!" [puppet] - 10https://gerrit.wikimedia.org/r/1309712 (https://phabricator.wikimedia.org/T431108) (owner: 10Scott French) [16:52:11] (03PS1) 10Ssingh: admin: add SSH key for sukhe [puppet] - 10https://gerrit.wikimedia.org/r/1310133 [16:52:30] (03CR) 10Scott French: [C:03+2] linked-artifacts: Use /healthz handler in blackbox probes [puppet] - 10https://gerrit.wikimedia.org/r/1309712 (https://phabricator.wikimedia.org/T431108) (owner: 10Scott French) [16:53:21] (03PS14) 10Federico Ceratto: sre.mysql: split pool/depool [cookbooks] - 10https://gerrit.wikimedia.org/r/1295480 (https://phabricator.wikimedia.org/T422361) [16:53:24] (03CR) 10JHathaway: [C:03+1] sre.hosts.bmc-user-mgmt: add more logging [cookbooks] - 10https://gerrit.wikimedia.org/r/1309226 (https://phabricator.wikimedia.org/T426180) (owner: 10Elukey) [16:55:26] (03CR) 10Ssingh: "@dzahn@wikimedia.org: Yes, we discussed it and this is fine for the purpose of the rollout/testing. Once done, we can revert back to 180 o" [dns] - 10https://gerrit.wikimedia.org/r/1306459 (https://phabricator.wikimedia.org/T425441) (owner: 10Arnaudb) [16:58:19] (03CR) 10Dzahn: [C:03+1] "gotcha! thanks. I will deploy it." [dns] - 10https://gerrit.wikimedia.org/r/1306459 (https://phabricator.wikimedia.org/T425441) (owner: 10Arnaudb) [16:59:00] (03CR) 10Dzahn: [C:03+2] gitlab: lower TTL for CNAMEs [dns] - 10https://gerrit.wikimedia.org/r/1306459 (https://phabricator.wikimedia.org/T425441) (owner: 10Arnaudb) [16:59:19] !log dzahn@dns1006 START - running authdns-update [16:59:35] mutante: though one thing I guess [16:59:42] we decided to do this when we would do the testing, not right now I think [16:59:50] sorry if I my CR text was not clear [17:00:05] Deploy window MediaWiki infrastructure (UTC late) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260713T1700) [17:00:05] ryankemper: I, the Bot under the Fountain, call upon thee, The Deployer, to do Wikidata Query Service weekly deploy deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260713T1700). [17:01:13] !log dzahn@dns1006 END - running authdns-update [17:01:58] sukhe: well, that was the question though.. if it's actually not ok permanently then we should not do it. [17:02:06] reverting [17:02:22] in that case I vote -1 to have a special case [17:03:05] (03PS1) 10Dzahn: Revert "gitlab: lower TTL for CNAMEs" [dns] - 10https://gerrit.wikimedia.org/r/1310134 [17:03:22] mutante: yeah let's discuss it on the CR I guess. the special case for just for the test in a way, so that user-impact if any can be kept to a minimum. that's for the team to decide though, I was simply commenting on if it is acceptable to go to 60, which for a special case it is [17:03:31] (03CR) 10Dzahn: [C:03+2] Revert "gitlab: lower TTL for CNAMEs" [dns] - 10https://gerrit.wikimedia.org/r/1310134 (owner: 10Dzahn) [17:03:33] I will comment on the task [17:04:08] ACK. I am against having special cases for one service we never do for other services then. [17:04:13] thanks [17:04:27] !log dzahn@dns1006 START - running authdns-update [17:04:36] reverted. no problem. [17:04:41] thanks [17:05:35] (03PS1) 10Dzahn: Revert^2 "gitlab: lower TTL for CNAMEs" [dns] - 10https://gerrit.wikimedia.org/r/1310135 [17:06:21] !log dzahn@dns1006 END - running authdns-update [17:06:32] (03CR) 10Dzahn: [C:03+1] "cool, thank you" [puppet] - 10https://gerrit.wikimedia.org/r/1308659 (https://phabricator.wikimedia.org/T431582) (owner: 10Hashar) [17:06:45] (03CR) 10Dzahn: [C:03+2] ci: add buildx on production Jenkins agents [puppet] - 10https://gerrit.wikimedia.org/r/1308659 (https://phabricator.wikimedia.org/T431582) (owner: 10Hashar) [17:07:19] !log ladsgroup@cumin1003 START - Cookbook sre.hosts.remove-downtime for ms-be1065.eqiad.wmnet [17:07:20] !log ladsgroup@cumin1003 END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1065.eqiad.wmnet [17:07:45] (03PS6) 10Effie Mouzeli: trafficserver: Remove XWD routing for /w/rest.php mw-debug backend [puppet] - 10https://gerrit.wikimedia.org/r/1310103 (https://phabricator.wikimedia.org/T428909) [17:09:52] (03PS7) 10Effie Mouzeli: trafficserver: Remove XWD routing for /w/rest.php mw-debug backend [puppet] - 10https://gerrit.wikimedia.org/r/1310103 (https://phabricator.wikimedia.org/T428909) [17:10:16] (03CR) 10Effie Mouzeli: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1310103 (https://phabricator.wikimedia.org/T428909) (owner: 10Effie Mouzeli) [17:11:56] (03CR) 10Dzahn: [C:03+1] "I think we should probably take a step back and independently from a specific scraping event just ask ourselves:" [puppet] - 10https://gerrit.wikimedia.org/r/1308081 (https://phabricator.wikimedia.org/T431398) (owner: 10Arnaudb) [17:13:01] !log ladsgroup@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1071.eqiad.wmnet with reason: vacuum overlarge container dbs [17:13:09] 06SRE, 10SRE-swift-storage: Disk near-full warnings on ms swift backends for container filesystems due to some bloated sqlite files - https://phabricator.wikimedia.org/T377827#12116071 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=8f59effe-3eb2-4acc-8cd9-40706e564a81) set by ladsgroup@cum... [17:13:15] (03CR) 10Dzahn: "How long do you think we should wait before we can do this now?" [puppet] - 10https://gerrit.wikimedia.org/r/1308249 (https://phabricator.wikimedia.org/T418521) (owner: 10Dzahn) [17:18:57] (03CR) 10Btullis: [C:03+2] Revert webrequest (+other) retention to 90 days [puppet] - 10https://gerrit.wikimedia.org/r/1310099 (https://phabricator.wikimedia.org/T431968) (owner: 10Joal) [17:20:53] (03CR) 10Dzahn: [V:03+1] "verified in compiler it's a NOOP https://puppet-compiler.wmflabs.org/output/1309271/8983/" [puppet] - 10https://gerrit.wikimedia.org/r/1309271 (https://phabricator.wikimedia.org/T240266) (owner: 10Dzahn) [17:22:06] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, July 13 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1212157 (https://phabricator.wikimedia.org/T409297) (owner: 10Esanders) [17:23:55] (03PS1) 10Ahmon Dancy: scap.cfg.erb: Add logstash_url settings [puppet] - 10https://gerrit.wikimedia.org/r/1310137 (https://phabricator.wikimedia.org/T431635) [17:27:12] (03CR) 10Dzahn: [V:03+1 C:03+2] gerrit: add support for multiple ssh host keys [puppet] - 10https://gerrit.wikimedia.org/r/1309271 (https://phabricator.wikimedia.org/T240266) (owner: 10Dzahn) [17:27:32] (03PS1) 10Santiago Faci: Test Kitchen UI: Deploying v1.4.8 release to production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310139 (https://phabricator.wikimedia.org/T428984) [17:27:54] 10ops-eqiad, 06DC-Ops: Find console port for leaf - https://phabricator.wikimedia.org/T432044 (10Jhancock.wm) 03NEW [17:28:02] 10ops-eqiad, 06DC-Ops: Find console port for leaf - https://phabricator.wikimedia.org/T432044#12116157 (10Jhancock.wm) p:05Triage→03Low [17:28:54] (03CR) 10Dzahn: [C:03+2] scap.cfg.erb: Add logstash_url settings [puppet] - 10https://gerrit.wikimedia.org/r/1310137 (https://phabricator.wikimedia.org/T431635) (owner: 10Ahmon Dancy) [17:30:24] mutante: Thanks! [17:31:50] 06SRE, 10SRE-SLO, 10Observability-Metrics: Rework the Pyrra list dashboard - https://phabricator.wikimedia.org/T394415#12116195 (10Aklapper) @elukey: Would you know? ^ [17:32:04] 06SRE, 06ServiceOps new, 07Service-deployment-requests: New Service Request: Headless VE - https://phabricator.wikimedia.org/T431497#12116199 (10ppelberg) [17:33:18] dancy: np, I happaned to be merging something else [17:34:41] (03CR) 10Dzahn: [C:04-1] "now that we have support for multiple keys I am going to change this to just add a new key in addition to the old key" [puppet] - 10https://gerrit.wikimedia.org/r/1282395 (https://phabricator.wikimedia.org/T240266) (owner: 10Dzahn) [17:37:39] 06SRE, 10SRE-SLO, 10Observability-Metrics: Rework the Pyrra list dashboard - https://phabricator.wikimedia.org/T394415#12116237 (10hnowlan) 05Open→03Declined [17:46:28] !log ladsgroup@cumin1003 START - Cookbook sre.hosts.remove-downtime for ms-be1071.eqiad.wmnet [17:46:29] !log ladsgroup@cumin1003 END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1071.eqiad.wmnet [17:50:04] PROBLEM - Check if ntpsec.service has been restarted after /etc/ntpsec/ntp.conf was changed on dns7002 is CRITICAL: CRITICAL: Service ntpsec.service has not been restarted after /etc/ntpsec/ntp.conf was changed (gt 2h). https://wikitech.wikimedia.org/wiki/NTP%23Monitoring [17:51:42] (03PS1) 10Hnowlan: karma: automatically linkify Phab task IDs in messages [puppet] - 10https://gerrit.wikimedia.org/r/1310141 (https://phabricator.wikimedia.org/T431980) [17:54:03] (03PS2) 10Dzahn: gerrit: ssh_host_keys: use openssh format and add ed25519 key [puppet] - 10https://gerrit.wikimedia.org/r/1282395 (https://phabricator.wikimedia.org/T240266) [17:54:30] (03CR) 10CI reject: [V:04-1] gerrit: ssh_host_keys: use openssh format and add ed25519 key [puppet] - 10https://gerrit.wikimedia.org/r/1282395 (https://phabricator.wikimedia.org/T240266) (owner: 10Dzahn) [17:54:53] (03PS3) 10Dzahn: gerrit: ssh_host_keys: use openssh format and add ed25519 key [puppet] - 10https://gerrit.wikimedia.org/r/1282395 (https://phabricator.wikimedia.org/T240266) [18:01:04] RECOVERY - Bird Internet Routing Daemon on dns7002 is OK: PROCS OK: 1 process with command name bird https://wikitech.wikimedia.org/wiki/Anycast%23Bird_daemon_not_running [18:01:08] RECOVERY - BFD status on asw1-b4-magru.mgmt is OK: UP: 2 AdminDown: 0 Down: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [18:01:13] (03CR) 10Dzahn: "@adancy@wikimedia.org This is an older patch now amended on top of the patch to add support for multiple keys. So it removes the existing" [puppet] - 10https://gerrit.wikimedia.org/r/1282395 (https://phabricator.wikimedia.org/T240266) (owner: 10Dzahn) [18:05:10] RESOLVED: [2x] BFDdown: BFD session down between asw1-b4-magru and 195.200.68.37 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=asw1-b4-magru:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [18:06:55] (03CR) 10Ahmon Dancy: "To help with the transition, we should ensure that the ed25519 key is added to wherever operations/debs/wmf-laptop gets its known hosts in" [puppet] - 10https://gerrit.wikimedia.org/r/1282395 (https://phabricator.wikimedia.org/T240266) (owner: 10Dzahn) [18:08:39] (03CR) 10Ahmon Dancy: "The IRC messages from hashar:" [puppet] - 10https://gerrit.wikimedia.org/r/1282395 (https://phabricator.wikimedia.org/T240266) (owner: 10Dzahn) [18:09:28] PROBLEM - Check unit status of httpbb_kubernetes_mw-web_hourly on cumin2003 is CRITICAL: CRITICAL: Status of the systemd unit httpbb_kubernetes_mw-web_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [18:09:28] PROBLEM - Check unit status of httpbb_kubernetes_mw-api-ext_hourly on cumin2003 is CRITICAL: CRITICAL: Status of the systemd unit httpbb_kubernetes_mw-api-ext_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [18:14:07] (03CR) 10Ahmon Dancy: "I can handle this part if you supply the text." [puppet] - 10https://gerrit.wikimedia.org/r/1282395 (https://phabricator.wikimedia.org/T240266) (owner: 10Dzahn) [18:14:29] (03PS1) 10Kamila Součková: ml-services/*: don't hardcode helmBinary [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310146 (https://phabricator.wikimedia.org/T388390) [18:15:24] PROBLEM - Check unit status of httpbb_kubernetes_mw-api-int_hourly on cumin2003 is CRITICAL: CRITICAL: Status of the systemd unit httpbb_kubernetes_mw-api-int_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [18:15:24] PROBLEM - Check unit status of httpbb_kubernetes_mw-api-ext-next_hourly on cumin2003 is CRITICAL: CRITICAL: Status of the systemd unit httpbb_kubernetes_mw-api-ext-next_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [18:15:38] FIRING: CertAlmostExpired: gNMI TLS certificate for lsw1-f6-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [18:16:26] PROBLEM - Check unit status of httpbb_kubernetes_mw-web-next_hourly on cumin2003 is CRITICAL: CRITICAL: Status of the systemd unit httpbb_kubernetes_mw-web-next_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [18:19:58] !log cdobbins@cumin2002 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie [18:22:51] !log lvextend vg0/srv +500g on centrallog hosts [18:22:52] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [18:24:23] (03CR) 10Bking: [C:03+1] service: change pybal check for opensearch clusters [puppet] - 10https://gerrit.wikimedia.org/r/1310129 (https://phabricator.wikimedia.org/T431538) (owner: 10Atsuko) [18:24:33] (03PS1) 10BCornwall: Enable auth rate_limiting_flags on upload [puppet] - 10https://gerrit.wikimedia.org/r/1310151 (https://phabricator.wikimedia.org/T431627) [18:27:12] (03CR) 10Btullis: [C:03+2] Correct typo in hiera for an-test-coord1002 [puppet] - 10https://gerrit.wikimedia.org/r/1310122 (https://phabricator.wikimedia.org/T406746) (owner: 10Btullis) [18:28:36] 06SRE, 10SRE-SLO, 06Traffic: Page on ATS backend errors relative to traffic - https://phabricator.wikimedia.org/T400675#12116461 (10ssingh) Thanks for the update @fgiunchedi. We should certainly move forward on this, one way or the other. @hnowlan: Would appreciate some input from you, or someone on olly on... [18:29:34] (03CR) 10BCornwall: [C:03+1] cache: Reduce webp threshold to 90 [puppet] - 10https://gerrit.wikimedia.org/r/1310105 (https://phabricator.wikimedia.org/T431150) (owner: 10Ladsgroup) [18:31:43] (03CR) 10BCornwall: [V:03+1 C:03+1] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/8984/co" [puppet] - 10https://gerrit.wikimedia.org/r/1310105 (https://phabricator.wikimedia.org/T431150) (owner: 10Ladsgroup) [18:32:07] (03CR) 10Kamila Součková: "recheck" [puppet] - 10https://gerrit.wikimedia.org/r/1305986 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [18:33:47] (03CR) 10Andrea Denisse: [C:03+1] "LGTM, thank you!!" [puppet] - 10https://gerrit.wikimedia.org/r/1310119 (https://phabricator.wikimedia.org/T327300) (owner: 10Ayounsi) [18:34:12] (03CR) 10BCornwall: "I cannot find any context here. What's going on with these changes?" [dns] - 10https://gerrit.wikimedia.org/r/1310135 (owner: 10Dzahn) [18:46:41] 06SRE, 10SRE-SLO, 06Traffic: Page on ATS backend errors relative to traffic - https://phabricator.wikimedia.org/T400675#12116524 (10ssingh) We will be discussing this tomorrow in the Traffic meeting as well and will follow up. [18:47:46] (03CR) 10Ssingh: "[discussing this tomorrow in Traffic on who gets to be the lucky review :)]" [puppet] - 10https://gerrit.wikimedia.org/r/1306896 (https://phabricator.wikimedia.org/T427465) (owner: 10Slyngshede) [18:47:58] (03CR) 10Ssingh: "reviewer." [puppet] - 10https://gerrit.wikimedia.org/r/1306896 (https://phabricator.wikimedia.org/T427465) (owner: 10Slyngshede) [18:49:42] (03CR) 10Kamila Součková: [C:03+1] "LGTM, but we should probably fix the CI failure. It's not related, so I don't insist on it, but somebody should do something :D" [puppet] - 10https://gerrit.wikimedia.org/r/1305986 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [18:50:08] (03CR) 10Dzahn: "pretend this is still https://gerrit.wikimedia.org/r/c/operations/dns/+/1306459 and that was never merged - this was discussion between Ar" [dns] - 10https://gerrit.wikimedia.org/r/1310135 (owner: 10Dzahn) [18:54:46] (03CR) 10Dzahn: "or just remove the etcd part from this patch that has the problem but get everything else out of the way" [puppet] - 10https://gerrit.wikimedia.org/r/1305986 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [18:55:22] (03CR) 10Dzahn: "etcd parameters getting undefined sounds kinda dangerous if real" [puppet] - 10https://gerrit.wikimedia.org/r/1305986 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [18:56:17] (03PS1) 10Majavah: P:openstack: cumin: master: Add support for Trixie [puppet] - 10https://gerrit.wikimedia.org/r/1310159 [18:58:33] (03CR) 10Majavah: [C:03+2] P:openstack: cumin: master: Add support for Trixie [puppet] - 10https://gerrit.wikimedia.org/r/1310159 (owner: 10Majavah) [19:01:08] (03PS1) 10CDobbins: bird: ensure that the `bird` user owns `/etc/bird` [puppet] - 10https://gerrit.wikimedia.org/r/1310160 [19:01:42] (03CR) 10CI reject: [V:04-1] bird: ensure that the `bird` user owns `/etc/bird` [puppet] - 10https://gerrit.wikimedia.org/r/1310160 (owner: 10CDobbins) [19:02:31] FIRING: Not accepting/receiving prefixes from anycast BGP peer: Alert for device asw1-b4-magru.mgmt.magru.wmnet - Not accepting/receiving prefixes from anycast BGP peer - https://alerts.wikimedia.org/?q=alertname%3DNot+accepting%2Freceiving+prefixes+from+anycast+BGP+peer [19:02:39] ^ yeah this is fine [19:02:41] it's depooled [19:02:52] we will keep it like this for one more day :> [19:09:20] (03PS2) 10CDobbins: bird: ensure that the `bird` user owns `/etc/bird` [puppet] - 10https://gerrit.wikimedia.org/r/1310160 [19:09:28] RECOVERY - Check unit status of httpbb_kubernetes_mw-web_hourly on cumin2003 is OK: OK: Status of the systemd unit httpbb_kubernetes_mw-web_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [19:09:28] RECOVERY - Check unit status of httpbb_kubernetes_mw-api-ext_hourly on cumin2003 is OK: OK: Status of the systemd unit httpbb_kubernetes_mw-api-ext_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [19:10:23] FIRING: [2x] CertAlmostExpired: gNMI TLS certificate for lsw1-d2-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [19:11:22] (03CR) 10CDobbins: [V:03+1] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/8987/co" [puppet] - 10https://gerrit.wikimedia.org/r/1310160 (owner: 10CDobbins) [19:15:01] (03CR) 10Ssingh: bird: ensure that the `bird` user owns `/etc/bird` (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1310160 (owner: 10CDobbins) [19:15:24] RECOVERY - Check unit status of httpbb_kubernetes_mw-api-ext-next_hourly on cumin2003 is OK: OK: Status of the systemd unit httpbb_kubernetes_mw-api-ext-next_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [19:15:24] RECOVERY - Check unit status of httpbb_kubernetes_mw-api-int_hourly on cumin2003 is OK: OK: Status of the systemd unit httpbb_kubernetes_mw-api-int_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [19:16:21] (03PS1) 10DLynch: Disable mobile "exit the editor" survey phase 2 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310161 (https://phabricator.wikimedia.org/T426135) [19:16:26] RECOVERY - Check unit status of httpbb_kubernetes_mw-web-next_hourly on cumin2003 is OK: OK: Status of the systemd unit httpbb_kubernetes_mw-web-next_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [19:16:33] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, July 13 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310161 (https://phabricator.wikimedia.org/T426135) (owner: 10DLynch) [19:24:05] 06SRE, 10Wikimedia-Mailing-lists: Some req staff not receiving emails to wmfreqs@lists.wikimedia.org - https://phabricator.wikimedia.org/T430308#12116642 (10Aklapper) 05Open→03Resolved No ticket updates in two weeks; boldly accepting the theory posted in the last comment until proven wrong (please reop... [19:24:45] (03PS3) 10CDobbins: bird: ensure that the `bird` user owns `/etc/bird` [puppet] - 10https://gerrit.wikimedia.org/r/1310160 [19:25:45] (03CR) 10CDobbins: bird: ensure that the `bird` user owns `/etc/bird` (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1310160 (owner: 10CDobbins) [19:26:29] (03CR) 10CDobbins: bird: ensure that the `bird` user owns `/etc/bird` (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1310160 (owner: 10CDobbins) [19:28:44] (03CR) 10Dzahn: "the script inside that wmf-laptop package reads from https://config-master.wikimedia.org/known_hosts but I don't know how it gets in there" [puppet] - 10https://gerrit.wikimedia.org/r/1282395 (https://phabricator.wikimedia.org/T240266) (owner: 10Dzahn) [19:31:28] (03PS3) 10JHathaway: Puppet 8: Replace legacy facts in misc kafka [puppet] - 10https://gerrit.wikimedia.org/r/1308760 (https://phabricator.wikimedia.org/T372666) [19:32:02] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308760 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [19:32:21] (03CR) 10CI reject: [V:04-1] Puppet 8: Replace legacy facts in misc kafka [puppet] - 10https://gerrit.wikimedia.org/r/1308760 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [19:32:40] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 6/7 UP : OSPFv3: 6/7 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [19:32:46] PROBLEM - OSPF status on cr1-drmrs is CRITICAL: OSPFv2: 3/4 UP : OSPFv3: 3/4 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [19:32:53] (03PS4) 10JHathaway: Puppet 8: Replace legacy facts in misc kafka [puppet] - 10https://gerrit.wikimedia.org/r/1308760 (https://phabricator.wikimedia.org/T372666) [19:33:01] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308760 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [19:33:10] FIRING: [2x] BFDdown: BFD session down between cr1-drmrs and 185.15.58.138 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-drmrs:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [19:33:39] (03CR) 10CI reject: [V:04-1] Puppet 8: Replace legacy facts in misc kafka [puppet] - 10https://gerrit.wikimedia.org/r/1308760 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [19:33:39] FIRING: CoreBGPDown: Core BGP session down between cr1-drmrs and cr2-eqiad (185.15.58.138) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=drmrs&var-device=cr1-drmrs:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [19:34:40] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 7/7 UP : OSPFv3: 7/7 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [19:34:46] RECOVERY - OSPF status on cr1-drmrs is OK: OSPFv2: 4/4 UP : OSPFv3: 4/4 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [19:34:49] (03PS5) 10JHathaway: Puppet 8: Replace legacy facts in misc kafka [puppet] - 10https://gerrit.wikimedia.org/r/1308760 (https://phabricator.wikimedia.org/T372666) [19:34:52] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308760 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [19:36:31] (03PS1) 10Jasmine: mesh: Copy mesh.deployment 1.3.2 to 1.3.3 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310164 (https://phabricator.wikimedia.org/T427024) [19:38:10] RESOLVED: [4x] BFDdown: BFD session down between cr1-drmrs and 185.15.58.138 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [19:38:39] RESOLVED: [4x] CoreBGPDown: Core BGP session down between cr1-drmrs and cr2-eqiad (185.15.58.138) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [19:39:26] (03CR) 10SD0001: [C:03+1] InitialiseSettings: change enwiki extendedconfirmed autopromote settings [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1307270 (https://phabricator.wikimedia.org/T431060) (owner: 10Novem Linguae) [19:40:47] (03CR) 10Ssingh: bird: ensure that the `bird` user owns `/etc/bird` (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1310160 (owner: 10CDobbins) [19:43:09] (03PS4) 10CDobbins: bird: ensure that the `bird` user owns `/etc/bird` [puppet] - 10https://gerrit.wikimedia.org/r/1310160 [19:43:43] (03CR) 10CDobbins: bird: ensure that the `bird` user owns `/etc/bird` (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1310160 (owner: 10CDobbins) [19:44:26] (03CR) 10Ssingh: bird: ensure that the `bird` user owns `/etc/bird` (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1310160 (owner: 10CDobbins) [19:48:29] (03CR) 10Scott French: [C:03+2] P:etcd::v3: Add support for explicitly enabling the v2 API [puppet] - 10https://gerrit.wikimedia.org/r/1309315 (https://phabricator.wikimedia.org/T428495) (owner: 10Scott French) [19:50:29] (03CR) 10CDobbins: [V:03+1] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/8988/co" [puppet] - 10https://gerrit.wikimedia.org/r/1310160 (owner: 10CDobbins) [19:55:02] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, July 13 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1304630 (https://phabricator.wikimedia.org/T53980) (owner: 10Sohom Datta) [20:00:05] RoanKattouw, urbanecm, TheresNoTime, kindrobot, and cjming: How many deployers does it take to do UTC late backport window deploy? (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260713T2000). [20:00:05] kemayo and Sohom_Datta: A patch you scheduled for UTC late backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [20:00:14] o/ [20:00:20] o/ [20:01:17] I can deploy mine. [20:01:26] Sohom_Datta: want me to throw yours in as well? [20:01:55] Definitely (I don't have deployer permissions 😀) [20:02:07] (03CR) 10TrainBranchBot: [C:03+2] "Approved by kemayo@deploy2003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1212157 (https://phabricator.wikimedia.org/T409297) (owner: 10Esanders) [20:02:07] (03CR) 10TrainBranchBot: [C:03+2] "Approved by kemayo@deploy2003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310161 (https://phabricator.wikimedia.org/T426135) (owner: 10DLynch) [20:02:07] (03CR) 10TrainBranchBot: [C:03+2] "Approved by kemayo@deploy2003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1304630 (https://phabricator.wikimedia.org/T53980) (owner: 10Sohom Datta) [20:02:12] Away we go, then. [20:03:32] (03Merged) 10jenkins-bot: Enable DiscussionTools visual enhancements on enwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1212157 (https://phabricator.wikimedia.org/T409297) (owner: 10Esanders) [20:03:35] (03Merged) 10jenkins-bot: Disable mobile "exit the editor" survey phase 2 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310161 (https://phabricator.wikimedia.org/T426135) (owner: 10DLynch) [20:03:39] (03Merged) 10jenkins-bot: Add source tab to ukwikisource's "Архів" (Archive) namespace [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1304630 (https://phabricator.wikimedia.org/T53980) (owner: 10Sohom Datta) [20:03:53] !log kemayo@deploy2003 Started scap sync-world: Backport for [[gerrit:1212157|Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161|Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630|Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] [20:04:02] T409297: Phase 6: Offer Usability Improvements as default-on feature at English Wikipedia - https://phabricator.wikimedia.org/T409297 [20:04:03] T426135: Deploy config change to stop "Exit the editor" survey (Phase 2) - https://phabricator.wikimedia.org/T426135 [20:04:03] T53980: Source tab not showing up in the Translation namespace - https://phabricator.wikimedia.org/T53980 [20:05:00] (03PS7) 10Scott French: zookeeper: Support installation from component/zookeeper34 [puppet] - 10https://gerrit.wikimedia.org/r/1309257 (https://phabricator.wikimedia.org/T428495) [20:05:00] (03PS5) 10Scott French: hieradata: Use component/zookeeper34 on conf* hosts (codfw) [puppet] - 10https://gerrit.wikimedia.org/r/1309260 (https://phabricator.wikimedia.org/T428495) [20:05:37] !log kemayo@deploy2003 soda, esanders, kemayo: Backport for [[gerrit:1212157|Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161|Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630|Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there [20:05:37] . [20:06:29] Sohom_Datta: Do you need to verify anything? [20:07:29] It should be mostly visual, but yeah gimme a sec! [20:07:39] Ah, I see it on the test servers, lgtm! [20:07:48] Great, thanks! [20:07:55] !log kemayo@deploy2003 soda, esanders, kemayo: Continuing with deployment [20:08:00] (03CR) 10Scott French: "Thanks for the review!" [puppet] - 10https://gerrit.wikimedia.org/r/1309257 (https://phabricator.wikimedia.org/T428495) (owner: 10Scott French) [20:10:59] (03CR) 10Scott French: [C:03+2] zookeeper: Support installation from component/zookeeper34 [puppet] - 10https://gerrit.wikimedia.org/r/1309257 (https://phabricator.wikimedia.org/T428495) (owner: 10Scott French) [20:12:09] (03PS2) 10Kamila Součková: Puppet 8: Replace legacy facts [puppet] - 10https://gerrit.wikimedia.org/r/1305986 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [20:12:18] !log kemayo@deploy2003 Finished scap sync-world: Backport for [[gerrit:1212157|Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161|Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630|Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] (duration: 08m 25s) [20:12:26] T409297: Phase 6: Offer Usability Improvements as default-on feature at English Wikipedia - https://phabricator.wikimedia.org/T409297 [20:12:26] T426135: Deploy config change to stop "Exit the editor" survey (Phase 2) - https://phabricator.wikimedia.org/T426135 [20:12:26] T53980: Source tab not showing up in the Translation namespace - https://phabricator.wikimedia.org/T53980 [20:12:30] Okay, everything listed in the window has been deployed. Things are free if anyone else needs to use it. [20:14:14] 🎉 Thank you for the deploy! [20:15:23] Also /offtopic but I hate that irccloud likes to bring up this emoji: 🤬 when I type "celebration" for some reason [20:15:39] (03CR) 10Kamila Součková: "I believe it was the rspec that wasn't dealing well with the newly multilevel dict, but someone who knows this better than me should doubl" [puppet] - 10https://gerrit.wikimedia.org/r/1305986 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [20:15:43] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [20:16:01] (03CR) 10Scott French: [C:03+2] hieradata: Use component/zookeeper34 on conf* hosts (codfw) [puppet] - 10https://gerrit.wikimedia.org/r/1309260 (https://phabricator.wikimedia.org/T428495) (owner: 10Scott French) [20:16:29] I will deploy a scap update. [20:16:39] !log dancy@deploy2003 Installing scap version "4.274.0" for 3 host(s) [20:18:31] !log dancy@deploy2003 Installation of scap version "4.274.0" completed for 3 hosts [20:19:30] !log dancy@deploy2003 Started scap sync-world: Testing T431635 [20:19:34] T431635: scap triggers a FutureWarning on a deployment - https://phabricator.wikimedia.org/T431635 [20:20:38] (03CR) 10Ssingh: "Looks good. Please run PCC for the affected host (dns7002) and one other DNS host." [puppet] - 10https://gerrit.wikimedia.org/r/1310160 (owner: 10CDobbins) [20:23:07] !log dancy@deploy2003 Finished scap sync-world: Testing T431635 (duration: 03m 36s) [20:26:51] !log reprepro include etcd-mirror_0.0.12-1+deb12u1 into main for bookworm-wikimedia - T428495 [20:26:54] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [20:26:55] T428495: Migrate conf* hosts away from bullseye - https://phabricator.wikimedia.org/T428495 [20:28:37] FIRING: [5x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-3/1/2 (Transport: Hurricane Electric (dc4841.dal5) {#122601}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [20:28:43] !log reprepro include etcd-mirror_0.0.12-1+deb13u1 into main for trixie-wikimedia - T424266 [20:28:45] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [20:28:46] T424266: Develop a plan for integrating conf200[7-9] - https://phabricator.wikimedia.org/T424266 [20:29:21] (03PS1) 10Ahmon Dancy: scap.cfg.erb: Remove deprecated logstash_host parameters [puppet] - 10https://gerrit.wikimedia.org/r/1310189 (https://phabricator.wikimedia.org/T431635) [20:30:43] FIRING: [2x] JobUnavailable: Reduced availability for job atlas_exporter in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [20:35:42] FIRING: [2x] JobUnavailable: Reduced availability for job atlas_exporter in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [20:55:44] (03CR) 10Dzahn: [C:03+2] scap.cfg.erb: Remove deprecated logstash_host parameters [puppet] - 10https://gerrit.wikimedia.org/r/1310189 (https://phabricator.wikimedia.org/T431635) (owner: 10Ahmon Dancy) [21:00:05] alexsanford, Reedy, sbassett, Maryum, and manfredi: #bothumor Q:How do functions break up? A:They stop calling each other. Rise for Weekly Security deployment window deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260713T2100). [21:00:26] preparing to deploy one security issue [21:00:35] doesn't look like anyone is deploying right now [21:02:36] (03CR) 10Scott French: "Thanks, Blake!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309641 (https://phabricator.wikimedia.org/T427668) (owner: 10Blake) [21:05:24] (03CR) 10Gmodena: [C:03+2] wdqs: load qlever index from qlever-index/ S3 prefix [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310066 (https://phabricator.wikimedia.org/T429150) (owner: 10Gmodena) [21:07:42] (03Merged) 10jenkins-bot: wdqs: load qlever index from qlever-index/ S3 prefix [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310066 (https://phabricator.wikimedia.org/T429150) (owner: 10Gmodena) [21:09:17] preparing to run scap [21:13:54] scap running [21:18:23] !log Deployed security fix for T321092 [21:18:25] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [21:20:03] (03CR) 10SBassett: varnish: put testwiki back into enforce mode (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1309263 (owner: 10CDobbins) [21:34:52] (03CR) 10Ahmon Dancy: "I'm investigating this. In the meantime, can you drop the public part of the ed25519 here please?" [puppet] - 10https://gerrit.wikimedia.org/r/1282395 (https://phabricator.wikimedia.org/T240266) (owner: 10Dzahn) [21:50:55] (03CR) 10Dzahn: "Thank you! I was also asking on #wikimedia-sre-foundations which pointed me to "class ssh::publish_fingerprints" in the puppet repo. So th" [puppet] - 10https://gerrit.wikimedia.org/r/1282395 (https://phabricator.wikimedia.org/T240266) (owner: 10Dzahn) [21:57:59] (03PS1) 10Ahmon Dancy: Add Gerrit's ed25519 public key [puppet] - 10https://gerrit.wikimedia.org/r/1310197 (https://phabricator.wikimedia.org/T240266) [21:59:54] (03CR) 10Dzahn: "was just going to update here that it's about that custom profile class for gerrit - but you are faster and already have a patch" [puppet] - 10https://gerrit.wikimedia.org/r/1282395 (https://phabricator.wikimedia.org/T240266) (owner: 10Dzahn) [22:24:35] (03CR) 10Scott French: "Nice! A couple of minor comments, but I think this should hold us over nicely until X-W-D support for REST gateway routed traffic lands." [puppet] - 10https://gerrit.wikimedia.org/r/1310103 (https://phabricator.wikimedia.org/T428909) (owner: 10Effie Mouzeli) [22:34:54] !log ladsgroup@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1069.eqiad.wmnet with reason: vacuum overlarge container dbs [22:35:03] 06SRE, 10SRE-swift-storage: Disk near-full warnings on ms swift backends for container filesystems due to some bloated sqlite files - https://phabricator.wikimedia.org/T377827#12117309 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=142aff76-01c7-4a7b-99f0-ca9594c6e151) set by ladsgroup@cum... [22:47:16] (03PS2) 10Dzahn: gerrit: refactor profile::gerrit::sshkey to handle multiple keys [puppet] - 10https://gerrit.wikimedia.org/r/1310197 (https://phabricator.wikimedia.org/T240266) (owner: 10Ahmon Dancy) [22:55:31] (03PS3) 10Dzahn: ssh/gerrit: add support for handling multiple ssh keys [puppet] - 10https://gerrit.wikimedia.org/r/1310197 (https://phabricator.wikimedia.org/T240266) (owner: 10Ahmon Dancy) [22:55:51] (03CR) 10Dzahn: [C:03+2] ssh/gerrit: add support for handling multiple ssh keys [puppet] - 10https://gerrit.wikimedia.org/r/1310197 (https://phabricator.wikimedia.org/T240266) (owner: 10Ahmon Dancy) [22:56:08] (03PS2) 10Ladsgroup: cache: Reduce webp threshold to 90 [puppet] - 10https://gerrit.wikimedia.org/r/1310105 (https://phabricator.wikimedia.org/T431150) [22:56:13] (03CR) 10Ladsgroup: [V:03+2 C:03+2] cache: Reduce webp threshold to 90 [puppet] - 10https://gerrit.wikimedia.org/r/1310105 (https://phabricator.wikimedia.org/T431150) (owner: 10Ladsgroup) [22:57:13] (03CR) 10Ladsgroup: [V:03+2 C:03+2] "Thanks for the PCC!" [puppet] - 10https://gerrit.wikimedia.org/r/1310105 (https://phabricator.wikimedia.org/T431150) (owner: 10Ladsgroup) [23:00:04] Deploy window Readers deployment window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260713T2300) [23:02:46] FIRING: Not accepting/receiving prefixes from anycast BGP peer: Alert for device asw1-b4-magru.mgmt.magru.wmnet - Not accepting/receiving prefixes from anycast BGP peer - https://alerts.wikimedia.org/?q=alertname%3DNot+accepting%2Freceiving+prefixes+from+anycast+BGP+peer [23:06:39] !log ladsgroup@cumin1003 START - Cookbook sre.hosts.remove-downtime for ms-be1069.eqiad.wmnet [23:06:40] !log ladsgroup@cumin1003 END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1069.eqiad.wmnet [23:08:20] !log ladsgroup@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1064.eqiad.wmnet with reason: vacuum overlarge container dbs [23:08:33] 06SRE, 10SRE-swift-storage: Disk near-full warnings on ms swift backends for container filesystems due to some bloated sqlite files - https://phabricator.wikimedia.org/T377827#12117368 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=749df0e8-2e0d-426d-86b0-90be81a3a8fd) set by ladsgroup@cum... [23:10:38] FIRING: [2x] CertAlmostExpired: gNMI TLS certificate for lsw1-d2-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [23:33:44] !log ladsgroup@cumin1003 START - Cookbook sre.hosts.remove-downtime for ms-be1064.eqiad.wmnet [23:33:45] !log ladsgroup@cumin1003 END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1064.eqiad.wmnet [23:38:05] (03PS11) 10Cwhite: opensearch_dashboards: add settings required to enable the security plugin [puppet] - 10https://gerrit.wikimedia.org/r/1305774 (https://phabricator.wikimedia.org/T350516) [23:42:10] (03PS1) 10TrainBranchBot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1310208 [23:42:11] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1310208 (owner: 10TrainBranchBot) [23:43:11] (03CR) 10Cwhite: "PCC OK: https://puppet-compiler.wmflabs.org/output/1305774/8989/" [puppet] - 10https://gerrit.wikimedia.org/r/1305774 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [23:47:20] (03CR) 10Dzahn: [C:03+2] "it works - https://config-master.wikimedia.org/known_hosts now has multiple lines for gerrit" [puppet] - 10https://gerrit.wikimedia.org/r/1310197 (https://phabricator.wikimedia.org/T240266) (owner: 10Ahmon Dancy) [23:50:18] (03Merged) 10jenkins-bot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1310208 (owner: 10TrainBranchBot)