[00:00:17] (03PS1) 10Aleksandar Mastilovic: Use a separate severity value to route to data engineering Slack [puppet] - 10https://gerrit.wikimedia.org/r/1319190 (https://phabricator.wikimedia.org/T432312) [00:00:56] (03CR) 10Aleksandar Mastilovic: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1319190 (https://phabricator.wikimedia.org/T432312) (owner: 10Aleksandar Mastilovic) [00:09:22] PROBLEM - Host asw1-eqsin is DOWN: PING CRITICAL - Packet loss = 100% [00:09:23] FIRING: [2x] CertAlmostExpired: gNMI TLS certificate for asw1-603-eqsin.mgmt.eqsin.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqsin - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [00:09:40] PROBLEM - Host ps1-603-eqsin is DOWN: PING CRITICAL - Packet loss = 100% [00:10:22] PROBLEM - Host ps1-604-eqsin is DOWN: PING CRITICAL - Packet loss = 100% [00:10:44] FIRING: RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [00:12:52] (03PS2) 10Aleksandar Mastilovic: Change severity of data eng Airflow alerts to "page" [alerts] - 10https://gerrit.wikimedia.org/r/1319189 (https://phabricator.wikimedia.org/T432312) [00:13:39] FIRING: [2x] CoreBGPDown: Core BGP session down between cr2-eqsin and mr1-eqsin (103.102.166.143) - group Management - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=eqsin&var-device=cr2-eqsin:9804&var-bgp_group=Management&var-bgp_neighbor=mr1-eqsin - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [00:14:28] FIRING: [2x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [00:20:44] RESOLVED: RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [00:27:24] RECOVERY - Host ps1-604-eqsin is UP: PING OK - Packet loss = 0%, RTA = 212.89 ms [00:27:28] RECOVERY - Host ps1-603-eqsin is UP: PING OK - Packet loss = 0%, RTA = 212.12 ms [00:27:40] RECOVERY - Host asw1-eqsin is UP: PING OK - Packet loss = 0%, RTA = 211.17 ms [00:28:58] FIRING: [2x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [00:29:23] FIRING: [2x] CertAlmostExpired: gNMI TLS certificate for asw1-603-eqsin.mgmt.eqsin.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqsin - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [00:33:39] RESOLVED: [2x] CoreBGPDown: Core BGP session down between cr2-eqsin and mr1-eqsin (103.102.166.143) - group Management - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=eqsin&var-device=cr2-eqsin:9804&var-bgp_group=Management&var-bgp_neighbor=mr1-eqsin - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [00:34:22] PROBLEM - Host ps1-604-eqsin is DOWN: PING CRITICAL - Packet loss = 100% [00:36:38] PROBLEM - Host cr2-eqsin.mgmt is DOWN: PING CRITICAL - Packet loss = 100% [00:36:38] PROBLEM - Host scs-eqsin is DOWN: PING CRITICAL - Packet loss = 100% [00:38:58] FIRING: CoreRouterInterfaceDown: Core router interface down - cr2-eqsin:fxp0 (Core: msw1-eqsin:24 {#1029}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqsin:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [00:38:58] FIRING: [2x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [00:43:58] RESOLVED: CoreRouterInterfaceDown: Core router interface down - cr2-eqsin:fxp0 (Core: msw1-eqsin:24 {#1029}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqsin:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [00:45:16] RECOVERY - Host ps1-604-eqsin is UP: PING OK - Packet loss = 0%, RTA = 212.74 ms [00:46:52] RECOVERY - Host scs-eqsin is UP: PING OK - Packet loss = 0%, RTA = 211.12 ms [00:46:52] RECOVERY - Host cr2-eqsin.mgmt is UP: PING OK - Packet loss = 0%, RTA = 215.91 ms [00:47:39] 10ops-eqsin, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 10netops: EQSIN:Switch refresh diagram and wiring - https://phabricator.wikimedia.org/T423724#12169791 (10Papaul) [00:48:58] FIRING: [2x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [00:54:23] RESOLVED: CertAlmostExpired: gNMI TLS certificate for asw1-604-eqsin.mgmt.eqsin.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqsin - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [00:58:46] PROBLEM - Host asw1-eqsin is DOWN: PING CRITICAL - Packet loss = 100% [01:00:02] PROBLEM - Host ps1-604-eqsin is DOWN: PING CRITICAL - Packet loss = 100% [01:00:24] PROBLEM - Host ps1-603-eqsin is DOWN: PING CRITICAL - Packet loss = 100% [01:02:39] FIRING: CoreBGPDown: Core BGP session down between cr2-eqsin and mr1-eqsin (2001:df2:e500:fe04::2) - group Management - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=eqsin&var-device=cr2-eqsin:9804&var-bgp_group=Management&var-bgp_neighbor=mr1-eqsin - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [01:03:23] FIRING: CertAlmostExpired: gNMI TLS certificate for asw1-604-eqsin.mgmt.eqsin.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqsin - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [01:03:58] FIRING: [2x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [01:05:24] RECOVERY - Host ps1-604-eqsin is UP: PING OK - Packet loss = 0%, RTA = 212.81 ms [01:05:26] RECOVERY - Host ps1-603-eqsin is UP: PING OK - Packet loss = 0%, RTA = 211.76 ms [01:05:30] RECOVERY - Host asw1-eqsin is UP: PING OK - Packet loss = 0%, RTA = 211.25 ms [01:07:39] RESOLVED: [2x] CoreBGPDown: Core BGP session down between cr2-eqsin and mr1-eqsin (103.102.166.143) - group Management - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=eqsin&var-device=cr2-eqsin:9804&var-bgp_group=Management&var-bgp_neighbor=mr1-eqsin - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [01:08:23] RESOLVED: [2x] CertAlmostExpired: gNMI TLS certificate for asw1-603-eqsin.mgmt.eqsin.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqsin - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [01:08:58] FIRING: [2x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [01:12:21] (03PS1) 10TrainBranchBot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1319197 [01:12:21] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1319197 (owner: 10TrainBranchBot) [01:17:44] PROBLEM - Host asw1-eqsin is DOWN: PING CRITICAL - Packet loss = 100% [01:18:22] PROBLEM - Host ps1-604-eqsin is DOWN: PING CRITICAL - Packet loss = 100% [01:18:22] PROBLEM - Host ps1-603-eqsin is DOWN: PING CRITICAL - Packet loss = 100% [01:21:39] FIRING: [2x] CoreBGPDown: Core BGP session down between cr2-eqsin and mr1-eqsin (103.102.166.143) - group Management - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=eqsin&var-device=cr2-eqsin:9804&var-bgp_group=Management&var-bgp_neighbor=mr1-eqsin - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [01:22:23] FIRING: [2x] CertAlmostExpired: gNMI TLS certificate for asw1-603-eqsin.mgmt.eqsin.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqsin - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [01:23:58] FIRING: [2x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [01:24:24] RECOVERY - Host ps1-604-eqsin is UP: PING OK - Packet loss = 0%, RTA = 211.91 ms [01:24:28] RECOVERY - Host ps1-603-eqsin is UP: PING OK - Packet loss = 0%, RTA = 211.72 ms [01:24:40] RECOVERY - Host asw1-eqsin is UP: PING OK - Packet loss = 0%, RTA = 211.34 ms [01:25:04] (03Merged) 10jenkins-bot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1319197 (owner: 10TrainBranchBot) [01:26:39] RESOLVED: [2x] CoreBGPDown: Core BGP session down between cr2-eqsin and mr1-eqsin (103.102.166.143) - group Management - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=eqsin&var-device=cr2-eqsin:9804&var-bgp_group=Management&var-bgp_neighbor=mr1-eqsin - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [01:27:23] RESOLVED: [2x] CertAlmostExpired: gNMI TLS certificate for asw1-603-eqsin.mgmt.eqsin.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqsin - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [01:28:58] FIRING: [2x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [01:38:28] !log brett@cumin2002 START - Cookbook sre.dns.admin DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, T433097] [01:38:32] T433097: eqsin downtime scheduling for switch upgrade - https://phabricator.wikimedia.org/T433097 [01:38:33] !log brett@cumin2002 END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, T433097] [01:47:05] 10ops-eqsin, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 10netops: EQSIN:Switch refresh diagram and wiring - https://phabricator.wikimedia.org/T423724#12169848 (10Papaul) [01:53:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [01:57:56] PROBLEM - Router interfaces on mr1-eqsin is CRITICAL: CRITICAL: host 103.102.166.128, interfaces up: 37, down: 1, dormant: 0, excluded: 0, unused: 0: https://wikitech.wikimedia.org/wiki/Network_monitoring%23Router_interface_down [02:00:38] !log mwpresync@deploy1003 Started scap build-images: Publishing wmf/next image [02:00:56] RECOVERY - Router interfaces on mr1-eqsin is OK: OK: host 103.102.166.128, interfaces up: 37, down: 0, dormant: 0, excluded: 0, unused: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23Router_interface_down [02:07:16] !log mwpresync@deploy1003 Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) [02:08:58] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:14:28] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:14:44] FIRING: RipeAtlasAnchorUnreachable: ipv6 ping to esams RIPE Atlas anchor: failures over threshold for measurement 59940414 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [02:19:44] FIRING: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to esams RIPE Atlas anchor: failures over threshold for measurement 59940414 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [02:39:36] 10ops-eqsin, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 10netops: EQSIN:New switch setup/configuration - https://phabricator.wikimedia.org/T418439#12169932 (10Papaul) [02:43:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at codfw: 24.66% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=codfw%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [02:49:04] FIRING: HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-rollback - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs-next - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [02:53:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at codfw: 24.38% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=codfw%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [02:54:50] PROBLEM - Router interfaces on mr1-drmrs is CRITICAL: CRITICAL: host 185.15.58.130, interfaces up: 34, down: 1, dormant: 0, excluded: 0, unused: 0: https://wikitech.wikimedia.org/wiki/Network_monitoring%23Router_interface_down [02:55:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at codfw: 23.45% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=codfw%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [02:55:54] PROBLEM - Host mr1-drmrs.oob is DOWN: CRITICAL - Host Unreachable (193.251.154.146) [02:56:58] PROBLEM - Host mr1-drmrs.oob IPv6 is DOWN: PING CRITICAL - Packet loss = 100% [02:58:56] PROBLEM - Router interfaces on mr1-eqsin is CRITICAL: CRITICAL: host 103.102.166.128, interfaces up: 37, down: 1, dormant: 0, excluded: 0, unused: 0: https://wikitech.wikimedia.org/wiki/Network_monitoring%23Router_interface_down [03:00:25] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at codfw: 23.45% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=codfw%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [03:03:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at codfw: 23.83% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=codfw%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [03:04:56] RECOVERY - Router interfaces on mr1-eqsin is OK: OK: host 103.102.166.128, interfaces up: 38, down: 0, dormant: 0, excluded: 0, unused: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23Router_interface_down [03:06:50] RECOVERY - Router interfaces on mr1-drmrs is OK: OK: host 185.15.58.130, interfaces up: 35, down: 0, dormant: 0, excluded: 0, unused: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23Router_interface_down [03:07:12] RECOVERY - Host mr1-drmrs.oob IPv6 is UP: PING WARNING - Packet loss = 0%, RTA = 504.84 ms [03:08:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at codfw: 24.14% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=codfw%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [03:11:18] RECOVERY - Host mr1-drmrs.oob is UP: PING OK - Packet loss = 0%, RTA = 86.42 ms [03:17:25] RESOLVED: SystemdUnitFailed: generate_vrts_aliases.service on mx-in2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [03:20:25] FIRING: SystemdUnitFailed: generate_vrts_aliases.service on mx-in2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [03:22:52] (03PS1) 10Papaul: Add new eqsin ASN and mr1-ge-0/0/3 to prod zone [homer/public] - 10https://gerrit.wikimedia.org/r/1319205 (https://phabricator.wikimedia.org/T418439) [03:34:46] 10ops-eqsin, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, and 2 others: EQSIN:New switch setup/configuration - https://phabricator.wikimedia.org/T418439#12169969 (10Papaul) I didn't disable one of the BGP session from mr1 to cr since we didn't make the connection from asw1-603-eqsin to cr2-cr3 et-0/0/2 (... [03:50:26] FIRING: [2x] GanetiBGPNoInPrefixes: ganeti2033 is not sending any prefix to lsw1-b7-codfw - group Ganeti4 - https://wikitech.wikimedia.org/wiki/Ganeti#GanetiBGPNoInPrefixes - https://alerts.wikimedia.org/?q=alertname%3DGanetiBGPNoInPrefixes [04:03:30] !log cwilliams@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance [04:03:38] !log cwilliams@cumin1003 dbctl commit (dc=all): 'Depooling db1252 (T431660)', diff saved to https://phabricator.wikimedia.org/P95725 and previous config saved to /var/cache/conftool/dbconfig/20260730-040337-cwilliams.json [04:22:13] !log pt1979@cumin2002 START - Cookbook sre.dns.netbox [04:28:23] pt1979@cumin2002 netbox (PID 2467158) is awaiting input [04:29:05] !log cwilliams@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1167.eqiad.wmnet with reason: Maintenance [04:29:16] !log cwilliams@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance [04:29:23] !log cwilliams@cumin1003 dbctl commit (dc=all): 'Depooling db1167 (T431660)', diff saved to https://phabricator.wikimedia.org/P95726 and previous config saved to /var/cache/conftool/dbconfig/20260730-042923-cwilliams.json [04:32:35] 06SRE, 10SRE-Access-Requests: Requesting access to Jenkins server for PWangai-WMF - https://phabricator.wikimedia.org/T433563 (10pwangai) 03NEW [04:35:30] !log root@cumin1003 START - Cookbook sre.mysql.pool pool db1167: Maintenance [04:47:34] !log pt1979@cumin2002 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" [04:47:40] !log pt1979@cumin2002 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002" [04:47:40] !log pt1979@cumin2002 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [04:48:23] 06SRE, 10SRE-Access-Requests: Requesting access to Jenkins server for PWangai-WMF - https://phabricator.wikimedia.org/T433563#12170028 (10pwangai) [05:03:59] !log cwilliams@cumin1003 dbctl commit (dc=all): 'Repooling after maintenance db1252 (T431660)', diff saved to https://phabricator.wikimedia.org/P95729 and previous config saved to /var/cache/conftool/dbconfig/20260730-050358-cwilliams.json [05:04:44] FIRING: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to esams RIPE Atlas anchor: failures over threshold for measurement 59940414 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [05:06:46] (03CR) 10ArielGlenn: [C:03+1] "nice catch!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310654 (owner: 10Aaron Schulz) [05:11:52] (03CR) 10ArielGlenn: [C:03+1] "untested but looks right to me." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1312695 (owner: 10Aaron Schulz) [05:14:07] !log cwilliams@cumin1003 dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95731 and previous config saved to /var/cache/conftool/dbconfig/20260730-051406-cwilliams.json [05:19:44] FIRING: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to esams RIPE Atlas anchor: failures over threshold for measurement 59940414 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [05:21:39] FIRING: [2x] CoreBGPDown: Core BGP session down between cr2-eqord and cr3-ulsfo (198.35.26.128) - group Confed_ulsfo - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [05:23:25] !log root@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1167: Maintenance [05:23:47] !log cwilliams@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1172.eqiad.wmnet with reason: Maintenance [05:23:55] !log cwilliams@cumin1003 dbctl commit (dc=all): 'Depooling db1172 (T431660)', diff saved to https://phabricator.wikimedia.org/P95733 and previous config saved to /var/cache/conftool/dbconfig/20260730-052354-cwilliams.json [05:23:58] FIRING: [4x] CoreRouterInterfaceDown: Core router interface down - cr2-codfw:xe-0/0/1:1 (Transport: cr2-eqord:xe-0/1/0 (Arelion, IC-314534 29ms 10Gbps wave) {#10694_12249-2}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [05:24:15] !log cwilliams@cumin1003 dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95734 and previous config saved to /var/cache/conftool/dbconfig/20260730-052414-cwilliams.json [05:30:27] !log root@cumin1003 START - Cookbook sre.mysql.pool pool db1172: Maintenance [05:33:56] PROBLEM - Host db1218 #page is DOWN: PING CRITICAL - Packet loss = 100% [05:34:23] !log cwilliams@cumin1003 dbctl commit (dc=all): 'Repooling after maintenance db1252 (T431660)', diff saved to https://phabricator.wikimedia.org/P95736 and previous config saved to /var/cache/conftool/dbconfig/20260730-053422-cwilliams.json [05:34:55] !incidents [05:34:56] 8230 (UNACKED) Host db1218 (paged) [05:34:56] 8229 (RESOLVED) [2x] ProbeDown sre (text-https:443 probes/service magru) [05:35:04] !ack [05:35:05] 8230 (ACKED) Host db1218 (paged) [05:37:40] lets see whats going on. its not the host which had hardware issues which jynus mentioned yesterday (db1245) [05:39:06] hey [05:39:07] need help? [05:39:41] I am supposed to be ooo but I am around, I will check [05:40:01] any help is welcome. db1245 was reimaged yesterday to trixie (T431660) and is a replica(?) [05:40:22] db1245 has nothing to do with db1218 is a s1 replica [05:40:23] I am depooling [05:40:31] so I assume we can just depool it again ? [05:40:34] great thanks [05:40:35] I got a p.age for transport [05:40:46] tx manuel [05:40:47] !incidents [05:40:47] 8230 (ACKED) Host db1218 (paged) [05:40:47] 8229 (RESOLVED) [2x] ProbeDown sre (text-https:443 probes/service magru) [05:41:03] ok there was one for traffic [05:41:05] sigh [05:41:12] !log marostegui@cumin1003 dbctl commit (dc=all): 'Depool db1217 it crashed', diff saved to https://phabricator.wikimedia.org/P95737 and previous config saved to /var/cache/conftool/dbconfig/20260730-054111-marostegui.json [05:41:19] I got pag.ed for db1218 [05:41:34] no not db1217 [05:41:48] the pag.e was for db1218 [05:41:51] yeah [05:41:54] the commit message is wrong [05:41:56] the depool is ok [05:42:00] https://phabricator.wikimedia.org/P95737 [05:42:14] okay :) [05:44:01] !log marostegui@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: crashed [05:45:12] (03PS1) 10Marostegui: db1218: Disable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1319383 [05:45:43] I can't ssh to the host and metric are missing since 5:33 UTC [05:46:02] yep it seems hardcrash again without HW logs [05:46:40] should we try to powercycle it or leave the rest of troubleshooting to data persistence folks? [05:46:52] I am on it [05:46:58] RECOVERY - Host db1218 #page is UP: PING WARNING - Packet loss = 33%, RTA = 0.47 ms [05:47:55] ah db1218 is back, ssh also works and it has 0 mins uptime [05:49:11] marostegui: is it fine to leave it depooled and follow up with data persistence after breakfast? [05:49:22] yes [05:49:29] we'll take it from here [05:50:07] 10ops-eqiad, 06DBA, 06DC-Ops: db1218 crashed - https://phabricator.wikimedia.org/T433565#12170065 (10Marostegui) p:05Triage→03Medium [05:50:11] 10ops-eqiad, 06DBA, 06DC-Ops: db1218 crashed - https://phabricator.wikimedia.org/T433565#12170070 (10Marostegui) [05:50:42] great, thank you very much! [05:50:42] I'll head over to breakfast then [05:51:34] jelto: enjoy! [05:51:49] thanks 🥐 [05:51:54] (03CR) 10Marostegui: [C:03+2] db1218: Disable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1319383 (owner: 10Marostegui) [05:52:21] 10ops-eqiad, 06DBA, 06DC-Ops: db1218 crashed - https://phabricator.wikimedia.org/T433565#12170087 (10Papaul) @Marostegui thank you for letting us know, we will look in deep on this issue on our end [05:52:24] (03CR) 10Marostegui: [C:03+2] "Related task: https://phabricator.wikimedia.org/T433565" [puppet] - 10https://gerrit.wikimedia.org/r/1319383 (owner: 10Marostegui) [05:53:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [06:00:04] Deploy window MediaWiki infrastructure (UTC early) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260730T0600) [06:00:04] marostegui, Amir1, and federico3: Your horoscope predicts another Primary database switchover deploy. May Zuul be (nice) with you. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260730T0600). [06:01:19] 10ops-eqiad, 06DBA, 06DC-Ops: db1218 crashed - https://phabricator.wikimedia.org/T433565#12170097 (10Papaul) Quick look at the host seeing that it is running an old version of BIOS and IDRAC this can explain the reason why we are not getting in anything in getsel BIOS version 1.92 new version 1.21 IDRAC ver... [06:02:06] 10ops-eqiad, 06DBA, 06DC-Ops: db1218 crashed - https://phabricator.wikimedia.org/T433565#12170098 (10Marostegui) >>! In T433565#12170097, @Papaul wrote: > Quick look at the host seeing that it is running an old version of BIOS and IDRAC this can explain the reason why we are not getting in anything in getsel... [06:02:20] marostegui: hey what is the other DB node that did crash ? [06:06:08] papaul: we've had quite a bunch doing the same thing (rebooting itself and showing no errors) [06:06:21] papaul: do you want me to send a list of tasks later? [06:06:48] marostegui: if you can please do thank you [06:07:12] papaul: I will later today, I am ooo and I was planning on doing some stuff now, I will send it in a couple of hours to your mail [06:07:31] marostegui: no problem me too going to bed [06:07:37] papaul: sweet dreams! [06:07:41] thanks [06:14:28] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [06:17:09] !log root@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1172: Maintenance [06:17:29] !log cwilliams@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1177.eqiad.wmnet with reason: Maintenance [06:17:37] !log cwilliams@cumin1003 dbctl commit (dc=all): 'Depooling db1177 (T431660)', diff saved to https://phabricator.wikimedia.org/P95741 and previous config saved to /var/cache/conftool/dbconfig/20260730-061736-cwilliams.json [06:24:03] !log root@cumin1003 START - Cookbook sre.mysql.pool pool db1177: Maintenance [06:29:00] (03CR) 10Slyngshede: [C:03+2] Account blocking: Only unset email if requested [software/bitu] - 10https://gerrit.wikimedia.org/r/1318083 (https://phabricator.wikimedia.org/T431114) (owner: 10Slyngshede) [06:31:44] (03Merged) 10jenkins-bot: Account blocking: Only unset email if requested [software/bitu] - 10https://gerrit.wikimedia.org/r/1318083 (https://phabricator.wikimedia.org/T431114) (owner: 10Slyngshede) [06:35:12] (03PS3) 10KineticPelagic: rest: Add test server option to REST Sandbox for Wikipedia projects [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1313827 (https://phabricator.wikimedia.org/T408816) [06:42:39] (03PS4) 10KineticPelagic: rest: Add test server option to REST Sandbox for Wikipedia projects [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1313827 (https://phabricator.wikimedia.org/T408816) [06:44:44] (03PS5) 10KineticPelagic: rest: Add test server option to REST Sandbox for Wikipedia projects [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1313827 (https://phabricator.wikimedia.org/T408816) [06:48:17] (03CR) 10Filippo Giunchedi: "Doh of course, my bad I wasn't thinking 'images' as in 'container images'." [debs/pint] - 10https://gerrit.wikimedia.org/r/1319039 (owner: 10Hnowlan) [06:49:04] FIRING: HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-rollback - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs-next - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [06:54:41] (03PS1) 10Slyngshede: data.yaml: record LDAP access for cdiggs-ctr [puppet] - 10https://gerrit.wikimedia.org/r/1319385 [06:55:11] (03CR) 10Slyngshede: [C:03+2] data.yaml offboarding for suecarmol [puppet] - 10https://gerrit.wikimedia.org/r/1318055 (owner: 10Slyngshede) [07:00:05] Amir1, urbanecm, and awight: May I have your attention please! UTC morning backport window. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260730T0700) [07:00:05] No Gerrit patches in the queue for this window AFAICS. [07:03:58] FIRING: [6x] CoreRouterInterfaceDown: Core router interface down - cr1-magru:et-0/0/0 (Transport: Hurricane Electric (dc4841.sao4) {#changeme_magru_he_cct}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [07:10:45] !log root@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1177: Maintenance [07:11:04] !log cwilliams@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1178.eqiad.wmnet with reason: Maintenance [07:11:12] !log cwilliams@cumin1003 dbctl commit (dc=all): 'Depooling db1178 (T431660)', diff saved to https://phabricator.wikimedia.org/P95746 and previous config saved to /var/cache/conftool/dbconfig/20260730-071112-cwilliams.json [07:14:44] RESOLVED: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to esams RIPE Atlas anchor: failures over threshold for measurement 59940414 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [07:17:26] !log root@cumin1003 START - Cookbook sre.mysql.pool pool db1178: Maintenance [07:20:25] FIRING: SystemdUnitFailed: generate_vrts_aliases.service on mx-in2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [07:23:02] (03CR) 10Klausman: [C:03+2] Add tts-section-generator CNAMEs to k8s-ingress-ml-serve [dns] - 10https://gerrit.wikimedia.org/r/1318685 (https://phabricator.wikimedia.org/T430536) (owner: 10Kevin Bazira) [07:23:11] !log btullis@cumin1003 END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host an-test-master1003.eqiad.wmnet [07:23:29] 10ops-eqiad, 06SRE, 06DBA, 06DC-Ops: db1218 crashed - https://phabricator.wikimedia.org/T433565#12170146 (10CWilliams-WMF) This host was rebooted for a new kernel yesterday, with its post-reboot repool being: ` 2026-07-29 14:02:24 [s1 db1218] restarting prometheus-mysqld-exporter... [07:24:19] !log klausman@dns2004 START - running authdns-update [07:25:43] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host ms-be2075.codfw.wmnet with OS trixie [07:25:56] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12170148 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host ms-be2075.codfw.wmnet with OS trixie [07:25:58] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie [07:26:07] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12170149 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-be1080.eqiad.wmnet with OS trixie [07:26:30] !log klausman@dns2004 END - running authdns-update [07:32:53] 06SRE, 07SRE-Unowned, 10DNS, 07Kubernetes: 10.67.28.73 reverse DNS showing 2(SERVFAIL) - https://phabricator.wikimedia.org/T428573#12170153 (10Marostegui) We ran into this issue again today whilst troubleshooting a persistent connection that was preventing some maintenance being done, would it be possible... [07:35:10] !log marostegui@cumin1003 dbctl commit (dc=all): 'Depool db1252', diff saved to https://phabricator.wikimedia.org/P95748 and previous config saved to /var/cache/conftool/dbconfig/20260730-073510-marostegui.json [07:38:33] !log dcausse@deploy1003 helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply [07:38:41] !log dcausse@deploy1003 helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply [07:42:39] (03PS5) 10Aaron Schulz: Remove some API Gateway specific template files [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319114 (https://phabricator.wikimedia.org/T428625) [07:44:56] !log cwilliams@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance [07:45:09] (03CR) 10CI reject: [V:04-1] Remove some API Gateway specific template files [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319114 (https://phabricator.wikimedia.org/T428625) (owner: 10Aaron Schulz) [07:46:34] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage [07:48:44] FIRING: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to esams RIPE Atlas anchor: failures over threshold for measurement 59940414 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [07:50:26] FIRING: [2x] GanetiBGPNoInPrefixes: ganeti2033 is not sending any prefix to lsw1-b7-codfw - group Ganeti4 - https://wikitech.wikimedia.org/wiki/Ganeti#GanetiBGPNoInPrefixes - https://alerts.wikimedia.org/?q=alertname%3DGanetiBGPNoInPrefixes [07:50:41] !log cwilliams@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1263.eqiad.wmnet with reason: Maintenance [07:50:41] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage [07:50:59] !log cwilliams@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[1260-1262].eqiad.wmnet with reason: Maintenance [07:51:07] !log cwilliams@cumin1003 dbctl commit (dc=all): 'Depooling db1263 (T431660)', diff saved to https://phabricator.wikimedia.org/P95751 and previous config saved to /var/cache/conftool/dbconfig/20260730-075106-cwilliams.json [07:57:11] (03CR) 10Ilias Sarantopoulos: ml-services: revert temporary change (set jina embeddings staging to ml-serve1012) (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319161 (https://phabricator.wikimedia.org/T433361) (owner: 10Ozge) [07:57:37] (03CR) 10Filippo Giunchedi: "For 2 that'd be "depends" only, not build-depends since sysusers is not needed at build time." [debs/pint] - 10https://gerrit.wikimedia.org/r/1319039 (owner: 10Hnowlan) [07:57:38] !log root@cumin1003 START - Cookbook sre.mysql.pool pool db1263: Maintenance [08:05:35] !log vriley@cumin1003 START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS bullseye [08:05:52] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12170226 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by vriley@cumin1003 for host cloudvirt1048.eqiad.wmnet with OS bullseye [08:05:55] !log root@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1178: Maintenance [08:06:04] !log cwilliams@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1192.eqiad.wmnet with reason: Maintenance [08:06:12] !log cwilliams@cumin1003 dbctl commit (dc=all): 'Depooling db1192 (T431660)', diff saved to https://phabricator.wikimedia.org/P95754 and previous config saved to /var/cache/conftool/dbconfig/20260730-080611-cwilliams.json [08:09:26] (03CR) 10Filippo Giunchedi: "What I had in mind is the following:" [puppet] - 10https://gerrit.wikimedia.org/r/1307368 (https://phabricator.wikimedia.org/T431017) (owner: 10Dpogorzelski) [08:11:12] 10ops-eqiad, 06SRE, 06DC-Ops, 10decommission-hardware, 06Data-Platform-SRE (2026-07-03 - 2026-07-31): decommission an-test-master100[1-2] - https://phabricator.wikimedia.org/T433495#12170247 (10VRiley-WMF) a:03VRiley-WMF [08:11:34] !log cwilliams@cumin1003 START - Cookbook sre.mysql.depool depool db1252: Maintenance [08:11:41] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1252: Maintenance [08:12:20] !log cwilliams@cumin1003 START - Cookbook sre.mysql.pool pool db1252: Maintenance [08:13:11] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2075.codfw.wmnet with OS trixie [08:13:23] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12170250 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-be2075.codfw.wmnet with OS trixie completed: - ms-be2075 (**PASS*... [08:14:21] !log root@cumin1003 START - Cookbook sre.mysql.pool pool db1192: Maintenance [08:15:46] 10ops-eqiad, 06SRE, 06Data-Persistence, 06DC-Ops: eqiad row A&B host migration details request for Data Persistence - https://phabricator.wikimedia.org/T432644#12170259 (10MatthewVernon) By renumbering, I meant "will the IP address change?". So yes, you have answered my question, thank you :) [08:15:57] !log mvernon@cumin1003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie [08:16:04] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12170260 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-be1080.eqiad.wmnet with OS trixie executed with errors: - ms-be10... [08:17:21] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie [08:17:32] 10ops-eqiad, 06SRE, 06DC-Ops, 10decommission-hardware, 06Data-Platform-SRE (2026-07-03 - 2026-07-31): decommission an-test-master100[1-2] - https://phabricator.wikimedia.org/T433495#12170274 (10VRiley-WMF) [08:17:33] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12170275 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-be1080.eqiad.wmnet with OS trixie [08:18:54] 10ops-eqiad, 06SRE, 06DC-Ops, 10decommission-hardware: Physical uninstall of an-test-master100[1-2] - https://phabricator.wikimedia.org/T433576 (10VRiley-WMF) 03NEW [08:19:07] 10ops-eqiad, 06SRE, 06DC-Ops, 10decommission-hardware: Physical uninstall of an-test-master100[1-2] - https://phabricator.wikimedia.org/T433576#12170301 (10VRiley-WMF) 05Open→03Resolved This has been completed [08:19:32] 10ops-eqiad, 06SRE, 06DC-Ops, 10decommission-hardware, 06Data-Platform-SRE (2026-07-03 - 2026-07-31): decommission an-test-master100[1-2] - https://phabricator.wikimedia.org/T433495#12170304 (10VRiley-WMF) a:05VRiley-WMF→03BTullis [08:20:02] 10ops-eqiad, 06SRE, 06DC-Ops, 10decommission-hardware, 06Data-Platform-SRE (2026-07-03 - 2026-07-31): decommission an-test-master100[1-2] - https://phabricator.wikimedia.org/T433495#12170307 (10VRiley-WMF) Hey @BTullis my part has been completed. I'm tossing this your way to finish it out, thanks! [08:21:12] (03CR) 10Filippo Giunchedi: [C:03+2] hieradata: enable dumps-nfs.w.o usage [puppet] - 10https://gerrit.wikimedia.org/r/1319053 (https://phabricator.wikimedia.org/T432587) (owner: 10Filippo Giunchedi) [08:21:50] (03PS3) 10Jelto: admin_ng: pin calico chart to current version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311394 (https://phabricator.wikimedia.org/T427400) [08:21:50] (03PS6) 10Jelto: Update calico-crds to calico v3.30.7 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306165 (https://phabricator.wikimedia.org/T427400) [08:21:50] (03PS4) 10Jelto: Update calico to v3.30.7 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306307 (https://phabricator.wikimedia.org/T427400) [08:22:11] (03CR) 10Jelto: "relation chain updated" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311394 (https://phabricator.wikimedia.org/T427400) (owner: 10Jelto) [08:23:39] !log elukey@cumin1003 START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts [08:23:59] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts [08:24:15] !log elukey@cumin1003 START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts [08:24:45] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts [08:25:06] 10ops-eqiad, 06SRE, 06DC-Ops, 10decommission-hardware: Physical decommission of an-test-coord1001 - https://phabricator.wikimedia.org/T433578 (10VRiley-WMF) 03NEW [08:25:20] 10ops-eqiad, 06SRE, 06DC-Ops, 10decommission-hardware: Physical decommission of an-test-coord1001 - https://phabricator.wikimedia.org/T433578#12170341 (10VRiley-WMF) 05Open→03Resolved This is completed. [08:25:40] 10ops-eqiad, 06SRE, 06DC-Ops, 10decommission-hardware, 06Data-Platform-SRE (2026-07-03 - 2026-07-31): decommission an-test-coord1001.eqiad.wmnet - https://phabricator.wikimedia.org/T433494#12170345 (10VRiley-WMF) a:03BTullis [08:26:35] 10ops-eqiad, 06SRE, 06DC-Ops, 10decommission-hardware, 06Data-Platform-SRE (2026-07-03 - 2026-07-31): decommission an-test-coord1001.eqiad.wmnet - https://phabricator.wikimedia.org/T433494#12170350 (10VRiley-WMF) [08:27:06] 10ops-eqiad, 06SRE, 06DC-Ops, 10decommission-hardware, 06Data-Platform-SRE (2026-07-03 - 2026-07-31): decommission an-test-coord1001.eqiad.wmnet - https://phabricator.wikimedia.org/T433494#12170352 (10VRiley-WMF) Hey @BTullis this is all yours. It has been physically removed. [08:29:40] 10ops-eqiad, 06SRE, 06DC-Ops: Unresponsive management for ms-be1065.mgmt:22 - https://phabricator.wikimedia.org/T433368#12170356 (10VRiley-WMF) 05Open→03Resolved a:03VRiley-WMF Reseated cable and there is now activity on it. It should be good to go. [08:34:42] (03PS7) 10Jelto: Update calico-crds to calico v3.30.7 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306165 (https://phabricator.wikimedia.org/T427400) [08:38:31] 10ops-eqiad, 06SRE, 06DBA, 06DC-Ops: db1218 crashed - https://phabricator.wikimedia.org/T433565#12170391 (10VRiley-WMF) Thanks, I'll take a look at this and update the firmware in just a moment [08:43:27] (03CR) 10Kevin Bazira: [C:03+1] ml-services: Add minReplicas=1 to gpt-oss and qwen36 deployments. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319085 (https://phabricator.wikimedia.org/T432933) (owner: 10Bartosz Wójtowicz) [08:44:43] !log root@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1263: Maintenance [08:45:53] (03CR) 10Bartosz Wójtowicz: [C:03+2] ml-services: Add minReplicas=1 to gpt-oss and qwen36 deployments. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319085 (https://phabricator.wikimedia.org/T432933) (owner: 10Bartosz Wójtowicz) [08:48:05] (03Merged) 10jenkins-bot: ml-services: Add minReplicas=1 to gpt-oss and qwen36 deployments. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319085 (https://phabricator.wikimedia.org/T432933) (owner: 10Bartosz Wójtowicz) [08:49:37] (03CR) 10JMeybohm: [C:03+1] Update update_version.py to be compatible with ruamel >=0.15.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310090 (owner: 10Mvolz) [08:49:46] (03PS1) 10Elukey: sre.hosts.bmc: vary the attributes URI for IPMI on Dell iDRAC 10+ [cookbooks] - 10https://gerrit.wikimedia.org/r/1319427 (https://phabricator.wikimedia.org/T426180) [08:50:36] !log elukey@cumin1003 START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts [08:50:36] (03CR) 10JMeybohm: [C:03+1] admin_ng: pin calico chart to current version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311394 (https://phabricator.wikimedia.org/T427400) (owner: 10Jelto) [08:50:48] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts [08:51:22] !log bwojtowicz@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . [08:51:37] !log elukey@cumin1003 START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts [08:52:15] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts [08:53:01] !log elukey@cumin1003 START - Cookbook sre.hosts.bmc-user-mgmt for 23 hosts [08:56:32] !log jayme@deploy1003 helmfile [staging] START helmfile.d/services/citoid: apply [08:57:28] !log jayme@deploy1003 helmfile [staging] DONE helmfile.d/services/citoid: apply [08:57:47] 10ops-eqiad, 06SRE, 10Data-Persistence-Backup, 06DC-Ops: db1245 crashed - https://phabricator.wikimedia.org/T431115#12170432 (10jcrespo) 05Open→03Resolved p:05Triage→03Medium I've marked it as active on netbox and will be reimaging it to put it into service again. [08:57:49] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1252: Maintenance [08:59:29] (03CR) 10Cathal Mooney: "nice to simplify this code in any way we can good stuff!" [software/homer/deploy] - 10https://gerrit.wikimedia.org/r/1313174 (https://phabricator.wikimedia.org/T432689) (owner: 10Ayounsi) [09:00:52] (03CR) 10Ayounsi: [C:03+2] Remove cable label from interfaces descriptions [software/homer/deploy] - 10https://gerrit.wikimedia.org/r/1313174 (https://phabricator.wikimedia.org/T432689) (owner: 10Ayounsi) [09:01:06] !log root@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1192: Maintenance [09:01:26] !log cwilliams@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1209.eqiad.wmnet with reason: Maintenance [09:01:34] !log cwilliams@cumin1003 dbctl commit (dc=all): 'Depooling db1209 (T431660)', diff saved to https://phabricator.wikimedia.org/P95767 and previous config saved to /var/cache/conftool/dbconfig/20260730-090133-cwilliams.json [09:02:18] !log ayounsi@cumin1003 START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 [09:04:27] !log bwojtowicz@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . [09:04:43] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003 [09:06:04] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 23 hosts [09:07:08] mvernon@cumin1003 reimage (PID 3774590) is awaiting input [09:07:54] !log root@cumin1003 START - Cookbook sre.mysql.pool pool db1209: Maintenance [09:08:08] (03CR) 10David Caro: [C:03+1] "LGTM, did not try it though, let me know if you want me to test it too" [puppet] - 10https://gerrit.wikimedia.org/r/1318228 (https://phabricator.wikimedia.org/T401818) (owner: 10Majavah) [09:09:05] (03PS1) 10JMeybohm: admin/data: Add new YubiKey ssh key [puppet] - 10https://gerrit.wikimedia.org/r/1319430 [09:09:32] (03PS2) 10Blake: kube-state-metrics: update to upstream 7.3.0. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319069 (https://phabricator.wikimedia.org/T427405) [09:09:32] (03CR) 10Blake: "Hi Janis! Is there a way to test this? It looks to me like what's happening is that we're swapping some deprecated values for their new ty" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319069 (https://phabricator.wikimedia.org/T427405) (owner: 10Blake) [09:09:35] (03CR) 10Ayounsi: sre.hosts.bmc: vary the attributes URI for IPMI on Dell iDRAC 10+ (031 comment) [cookbooks] - 10https://gerrit.wikimedia.org/r/1319427 (https://phabricator.wikimedia.org/T426180) (owner: 10Elukey) [09:11:22] (03PS2) 10Elukey: sre.hosts.bmc: vary the attributes URI for IPMI on Dell iDRAC 10+ [cookbooks] - 10https://gerrit.wikimedia.org/r/1319427 (https://phabricator.wikimedia.org/T426180) [09:11:31] (03CR) 10Jelto: [C:03+1] "verified out of band (at the kitchen table)" [puppet] - 10https://gerrit.wikimedia.org/r/1319430 (owner: 10JMeybohm) [09:11:50] (03CR) 10Ayounsi: [C:03+1] sre.hosts.bmc: vary the attributes URI for IPMI on Dell iDRAC 10+ [cookbooks] - 10https://gerrit.wikimedia.org/r/1319427 (https://phabricator.wikimedia.org/T426180) (owner: 10Elukey) [09:13:24] (03CR) 10Blake: httpbb: Add a --request-header argument. (033 comments) [software/httpbb] - 10https://gerrit.wikimedia.org/r/1318682 (https://phabricator.wikimedia.org/T428972) (owner: 10Blake) [09:14:38] (03CR) 10Ozge: [V:03+2 C:03+2] ml-services: revert temporary change (set jina embeddings staging to ml-serve1012) (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319161 (https://phabricator.wikimedia.org/T433361) (owner: 10Ozge) [09:15:53] (03CR) 10Ayounsi: [C:03+2] Don't fetch the cable label [software/homer] - 10https://gerrit.wikimedia.org/r/1313177 (https://phabricator.wikimedia.org/T432689) (owner: 10Ayounsi) [09:15:58] (03CR) 10Elukey: [C:03+2] sre.hosts.bmc: vary the attributes URI for IPMI on Dell iDRAC 10+ (031 comment) [cookbooks] - 10https://gerrit.wikimedia.org/r/1319427 (https://phabricator.wikimedia.org/T426180) (owner: 10Elukey) [09:16:42] (03CR) 10Cathal Mooney: [C:03+1] "Looks good! Durations make sense." [software/netbox-extras] - 10https://gerrit.wikimedia.org/r/1311459 (https://phabricator.wikimedia.org/T432329) (owner: 10Ayounsi) [09:17:35] (03CR) 10Ayounsi: [C:03+2] cables report: add test_old_planned_cables [software/netbox-extras] - 10https://gerrit.wikimedia.org/r/1311459 (https://phabricator.wikimedia.org/T432329) (owner: 10Ayounsi) [09:17:55] !log mvernon@cumin1003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie [09:18:06] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12170489 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-be1080.eqiad.wmnet with OS trixie executed with errors: - ms-be10... [09:18:49] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie [09:18:57] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12170492 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-be1080.eqiad.wmnet with OS trixie [09:19:26] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host ms-be2076.codfw.wmnet with OS trixie [09:19:37] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12170495 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host ms-be2076.codfw.wmnet with OS trixie [09:21:39] FIRING: [2x] CoreBGPDown: Core BGP session down between cr2-eqord and cr3-ulsfo (198.35.26.128) - group Confed_ulsfo - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [09:22:01] (03Merged) 10jenkins-bot: cables report: add test_old_planned_cables [software/netbox-extras] - 10https://gerrit.wikimedia.org/r/1311459 (https://phabricator.wikimedia.org/T432329) (owner: 10Ayounsi) [09:24:58] (03CR) 10JMeybohm: [C:03+2] admin/data: Add new YubiKey ssh key [puppet] - 10https://gerrit.wikimedia.org/r/1319430 (owner: 10JMeybohm) [09:26:33] (03PS2) 10Dpogorzelski: ml-serve: enlarge kubelet LV on GPU nodes [puppet] - 10https://gerrit.wikimedia.org/r/1307368 (https://phabricator.wikimedia.org/T431017) [09:28:05] (03CR) 10Lucas Werkmeister (WMDE): [C:03+2] "I think +2 in this repo has to be done by the deployer." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1318731 (owner: 10Lucas Werkmeister (WMDE)) [09:28:13] * Lucas_WMDE deploying ^ [09:28:44] (03PS3) 10Dpogorzelski: ml-serve: enlarge kubelet LV on GPU nodes [puppet] - 10https://gerrit.wikimedia.org/r/1307368 (https://phabricator.wikimedia.org/T431017) [09:28:45] (03CR) 10CI reject: [V:04-1] ml-serve: enlarge kubelet LV on GPU nodes [puppet] - 10https://gerrit.wikimedia.org/r/1307368 (https://phabricator.wikimedia.org/T431017) (owner: 10Dpogorzelski) [09:29:03] !log cwilliams@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2035.codfw.wmnet with reason: Maintenance [09:29:11] !log cwilliams@cumin1003 dbctl commit (dc=all): 'Depooling es2035 (T431660)', diff saved to https://phabricator.wikimedia.org/P95771 and previous config saved to /var/cache/conftool/dbconfig/20260730-092910-cwilliams.json [09:31:00] (03CR) 10CI reject: [V:04-1] ml-serve: enlarge kubelet LV on GPU nodes [puppet] - 10https://gerrit.wikimedia.org/r/1307368 (https://phabricator.wikimedia.org/T431017) (owner: 10Dpogorzelski) [09:31:01] (03Merged) 10jenkins-bot: wikidata-query-gui: update staging custom-config.json [deployment-charts] - 10https://gerrit.wikimedia.org/r/1318731 (owner: 10Lucas Werkmeister (WMDE)) [09:31:34] 10ops-eqiad, 06SRE, 10Data-Persistence-Backup, 06DC-Ops: db1245 crashed - https://phabricator.wikimedia.org/T431115#12170547 (10jcrespo) 05Resolved→03Open Reopening because reimaging it failed: ` Running IPMI command: ipmitool -I lanplus -H db1245.mgmt.eqiad.wmnet -U root -E chassis power status Error... [09:32:45] !log lucaswerkmeister-wmde@deploy1003 helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply [09:32:54] !log lucaswerkmeister-wmde@deploy1003 helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply [09:34:39] !log root@cumin1003 START - Cookbook sre.mysql.pool pool es2035: Maintenance [09:34:45] (03Merged) 10jenkins-bot: Don't fetch the cable label [software/homer] - 10https://gerrit.wikimedia.org/r/1313177 (https://phabricator.wikimedia.org/T432689) (owner: 10Ayounsi) [09:35:15] !log lucaswerkmeister-wmde@deploy1003 helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply [09:35:21] !log lucaswerkmeister-wmde@deploy1003 helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply [09:35:26] !log lucaswerkmeister-wmde@deploy1003 helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply [09:35:31] !log lucaswerkmeister-wmde@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply [09:36:25] 10ops-eqiad, 06SRE, 10Data-Persistence-Backup, 06DC-Ops: db1245 crashed - https://phabricator.wikimedia.org/T431115#12170571 (10VRiley-WMF) Hey @jcrespo, I was able to log in via iDRAC and power it on. It was seemingly seeing an issue with tempature, which I may need to address. However, could you please... [09:37:18] hm, not quite working as expected… https://gerrit.wikimedia.org/r/c/operations/deployment-charts/+/1318731 should have changed the staging deploy, but I don’t see the change, and according to kubectl the pods are 2d22h old (despite helmfile saying that Release "wikidata-query-gui" has been upgraded.) [09:38:24] and I’m not allowed to use the `helm` commands printed out by helmfile (Error: query: failed to query with labels: secrets is forbidden: User "wikidata-query-gui" cannot list resource "secrets" in API group "" in the namespace "wikidata-query-gui") [09:38:59] FIRING: [10x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-3/1/2 (Transport: Hurricane Electric (dc4841.dal5)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [09:39:37] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage [09:39:37] !log ayounsi@cumin1003 START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox-canary [09:39:50] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox-canary [09:40:14] !log root@cumin1003 END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2035: Maintenance [09:41:42] !log ayounsi@cumin1003 START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox [09:42:12] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox [09:42:14] !log cwilliams@cumin1003 START - Cookbook sre.mysql.pool pool es2035: Maintenance [09:43:39] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage [09:43:58] FIRING: [10x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-3/1/2 (Transport: Hurricane Electric (dc4841.dal5)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [09:45:50] ⇒ T433588 (and T433589) [09:45:51] T433588: helmfile apply with new values*.yaml file did not deploy new k8s pods - https://phabricator.wikimedia.org/T433588 [09:45:52] T433589: helmfile prints broken helm get command - https://phabricator.wikimedia.org/T433589 [09:46:25] (03CR) 10Lucas Werkmeister (WMDE): [C:03+2] "Hm, it didn’t quite work due to T433588…" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1318731 (owner: 10Lucas Werkmeister (WMDE)) [09:47:22] 10ops-codfw, 10ops-drmrs, 10ops-eqdfw, 10ops-eqiad, and 6 others: Fix cables with placeholders names and planned status - https://phabricator.wikimedia.org/T432317#12170626 (10ayounsi) I added a new test in the cable Netbox report for Cable with 'planned' status for more than 90 days (warning) and 120 days... [09:47:55] (03PS4) 10Dpogorzelski: ml-serve: enlarge kubelet LV on GPU nodes [puppet] - 10https://gerrit.wikimedia.org/r/1307368 (https://phabricator.wikimedia.org/T431017) [09:48:58] FIRING: NELHigh: Elevated Network Error Logging events (tcp.timed_out) #page - https://wikitech.wikimedia.org/wiki/Network_monitoring#NEL_alerts - https://logstash.wikimedia.org/goto/5c8f4ca1413eda33128e5c5a35da7e28 - https://alerts.wikimedia.org/?q=alertname%3DNELHigh [09:48:58] FIRING: [12x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-3/1/2 (Transport: Hurricane Electric (dc4841.dal5)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [09:49:09] !incidents [09:49:10] 8231 (UNACKED) NELHigh sre (thanos-rule@main tcp.timed_out) [09:49:10] 8230 (RESOLVED) Host db1218 (paged) [09:49:10] 8229 (RESOLVED) [2x] ProbeDown sre (text-https:443 probes/service magru) [09:49:20] !ack [09:49:21] 8231 (ACKED) NELHigh sre (thanos-rule@main tcp.timed_out) [09:49:59] FIRING: NELByCountryHigh: Elevated Network Error Logging events (tcp.timed_out from RU) - https://wikitech.wikimedia.org/wiki/Network_monitoring#NEL_alerts - https://logstash.wikimedia.org/goto/5c8f4ca1413eda33128e5c5a35da7e28 - https://alerts.wikimedia.org/?q=alertname%3DNELByCountryHigh [09:53:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:53:58] FIRING: [13x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-3/1/2 (Transport: Hurricane Electric (dc4841.dal5)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [09:54:35] !log root@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1209: Maintenance [09:54:44] !log cwilliams@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1214.eqiad.wmnet with reason: Maintenance [09:54:52] !log cwilliams@cumin1003 dbctl commit (dc=all): 'Depooling db1214 (T431660)', diff saved to https://phabricator.wikimedia.org/P95775 and previous config saved to /var/cache/conftool/dbconfig/20260730-095451-cwilliams.json [09:56:21] PROBLEM - Postgres Replication Lag on puppetdb2003 is CRITICAL: POSTGRES_HOT_STANDBY_DELAY CRITICAL: DB puppetdb (host:localhost) 212205592 and 26 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [09:58:19] RECOVERY - Postgres Replication Lag on puppetdb2003 is OK: POSTGRES_HOT_STANDBY_DELAY OK: DB puppetdb (host:localhost) 91552 and 1 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [10:00:05] Deploy window MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260730T1000) [10:00:51] !log root@cumin1003 START - Cookbook sre.mysql.pool pool db1214: Maintenance [10:01:39] RESOLVED: [2x] CoreBGPDown: Core BGP session down between cr2-eqord and cr3-ulsfo (198.35.26.128) - group Confed_ulsfo - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [10:03:05] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2076.codfw.wmnet with OS trixie [10:03:14] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12170686 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-be2076.codfw.wmnet with OS trixie completed: - ms-be2076 (**PASS*... [10:03:58] FIRING: [10x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-3/1/2 (Transport: Hurricane Electric (dc4841.dal5)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [10:08:49] mvernon@cumin1003 reimage (PID 3782878) is awaiting input [10:08:58] FIRING: [8x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-3/1/2 (Transport: Hurricane Electric (dc4841.dal5)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [10:14:28] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [10:19:41] (03PS6) 10Jcrespo: dbbackups: Setup db1265 & db1285 as new backup sources [puppet] - 10https://gerrit.wikimedia.org/r/1319083 (https://phabricator.wikimedia.org/T407942) [10:22:11] (03CR) 10Dpogorzelski: "Acknowledged" [puppet] - 10https://gerrit.wikimedia.org/r/1307368 (https://phabricator.wikimedia.org/T431017) (owner: 10Dpogorzelski) [10:27:50] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: Maintenance [10:28:34] (03CR) 10Jcrespo: [C:03+2] dbbackups: Setup db1265 & db1285 as new backup sources [puppet] - 10https://gerrit.wikimedia.org/r/1319083 (https://phabricator.wikimedia.org/T407942) (owner: 10Jcrespo) [10:30:58] (03CR) 10Hnowlan: opensearch: add security-plugin specific configuration to opensearch.yml (033 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1318792 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [10:47:33] !log root@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1214: Maintenance [10:47:53] !log cwilliams@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1226.eqiad.wmnet with reason: Maintenance [10:48:01] !log cwilliams@cumin1003 dbctl commit (dc=all): 'Depooling db1226 (T431660)', diff saved to https://phabricator.wikimedia.org/P95783 and previous config saved to /var/cache/conftool/dbconfig/20260730-104801-cwilliams.json [10:49:05] FIRING: HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-rollback - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs-next - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [10:55:11] !log root@cumin1003 START - Cookbook sre.mysql.pool pool db1226: Maintenance [10:55:40] (03PS1) 10MVernon: installserver: use ms-be_simple.cfg for ms-be1080 [puppet] - 10https://gerrit.wikimedia.org/r/1319442 (https://phabricator.wikimedia.org/T429630) [10:58:31] (03PS1) 10Ozge: ml-services: Upgrade jina-embeddings-staging image to log GPU memory [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319443 (https://phabricator.wikimedia.org/T432717) [10:59:05] (03CR) 10Ozge: [V:03+2 C:03+2] "merging this as we are testing the memory issue." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319443 (https://phabricator.wikimedia.org/T432717) (owner: 10Ozge) [11:00:55] (03CR) 10CWilliams: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1319442 (https://phabricator.wikimedia.org/T429630) (owner: 10MVernon) [11:01:14] !log ozge@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . [11:02:58] (03CR) 10MVernon: [C:03+2] installserver: use ms-be_simple.cfg for ms-be1080 [puppet] - 10https://gerrit.wikimedia.org/r/1319442 (https://phabricator.wikimedia.org/T429630) (owner: 10MVernon) [11:03:27] !log mvernon@cumin1003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie [11:03:40] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12170849 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-be1080.eqiad.wmnet with OS trixie executed... [11:07:38] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie [11:07:58] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12170860 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-be1080.eqiad.wmnet with OS trixie [11:08:16] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host ms-be2077.codfw.wmnet with OS trixie [11:08:29] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12170863 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host ms-be2077.codfw.wmnet with OS trixie [11:15:23] (03CR) 10Dreamy Jazz: [C:03+1] WikimediaAntiAbuse: Document required load order after Echo [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319086 (https://phabricator.wikimedia.org/T432452) (owner: 10Mpostoronca) [11:16:25] !log cwilliams@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2036.codfw.wmnet with reason: Maintenance [11:16:33] !log cwilliams@cumin1003 dbctl commit (dc=all): 'Depooling es2036 (T431660)', diff saved to https://phabricator.wikimedia.org/P95786 and previous config saved to /var/cache/conftool/dbconfig/20260730-111633-cwilliams.json [11:20:40] FIRING: SystemdUnitFailed: generate_vrts_aliases.service on mx-in2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [11:21:58] !log root@cumin1003 START - Cookbook sre.mysql.pool pool es2036: Maintenance [11:24:26] !log mvernon@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage [11:26:04] (03CR) 10Jelto: "recheck" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1318569 (https://phabricator.wikimedia.org/T388390) (owner: 10Jelto) [11:27:32] !log root@cumin1003 END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2036: Maintenance [11:27:37] !log root@cumin1003 START - Cookbook sre.mysql.pool pool es2036: Maintenance [11:27:56] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage [11:28:51] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage [11:32:31] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage [11:39:45] (03PS1) 10Ozge: ml-services: Tune jina-embeddings vLLM dtype and attention backend [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319448 (https://phabricator.wikimedia.org/T432717) [11:40:16] (03CR) 10Ozge: [V:03+2 C:03+2] "testing a config." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319448 (https://phabricator.wikimedia.org/T432717) (owner: 10Ozge) [11:40:49] (03PS1) 10Michael Große: SpecialCreateAccount: make user policy-link available again [extensions/WikimediaMessages] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1319449 (https://phabricator.wikimedia.org/T430604) [11:41:00] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Thursday, July 30 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployca" [extensions/WikimediaMessages] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1319449 (https://phabricator.wikimedia.org/T430604) (owner: 10Michael Große) [11:41:09] !log ozge@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . [11:41:53] !log root@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1226: Maintenance [11:42:11] (03CR) 10Ayounsi: [C:03+1] Add new eqsin ASN and mr1-ge-0/0/3 to prod zone [homer/public] - 10https://gerrit.wikimedia.org/r/1319205 (https://phabricator.wikimedia.org/T418439) (owner: 10Papaul) [11:43:43] (03PS1) 10Arthur taylor: wikidata-query-gui: bump to latest version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319451 (https://phabricator.wikimedia.org/T433233) [11:44:46] (03CR) 10CI reject: [V:04-1] SpecialCreateAccount: make user policy-link available again [extensions/WikimediaMessages] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1319449 (https://phabricator.wikimedia.org/T430604) (owner: 10Michael Große) [11:45:21] (03PS2) 10Arthur taylor: wikidata-query-gui: bump to latest version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319451 (https://phabricator.wikimedia.org/T433233) [11:46:38] (03PS1) 10Ozge: ml-services: Tune jina-embeddings vLLM dtype and attention backend [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319452 (https://phabricator.wikimedia.org/T432717) [11:46:40] (03CR) 10Michael Große: "recheck (maybe this will make it run on a different host)" [extensions/WikimediaMessages] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1319449 (https://phabricator.wikimedia.org/T430604) (owner: 10Michael Große) [11:46:56] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1080.eqiad.wmnet with OS trixie [11:47:03] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12170999 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-be1080.eqiad.wmnet with OS trixie completed: - ms-be1080 (**PASS*... [11:47:09] (03CR) 10Ozge: [V:03+2 C:03+2] "testing one more ATTN_IMPLEMENTATION" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319452 (https://phabricator.wikimedia.org/T432717) (owner: 10Ozge) [11:47:22] (03CR) 10CI reject: [V:04-1] wikidata-query-gui: bump to latest version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319451 (https://phabricator.wikimedia.org/T433233) (owner: 10Arthur taylor) [11:47:58] (03PS1) 10Kamila Součková: php: rebuild to include Timo's APCU metrics fix [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1319453 (https://phabricator.wikimedia.org/T433312) [11:48:11] !log ozge@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . [11:48:30] (03CR) 10Arthur taylor: "recheck" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319451 (https://phabricator.wikimedia.org/T433233) (owner: 10Arthur taylor) [11:48:59] FIRING: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to esams RIPE Atlas anchor: failures over threshold for measurement 59940414 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [11:50:26] FIRING: [2x] GanetiBGPNoInPrefixes: ganeti2033 is not sending any prefix to lsw1-b7-codfw - group Ganeti4 - https://wikitech.wikimedia.org/wiki/Ganeti#GanetiBGPNoInPrefixes - https://alerts.wikimedia.org/?q=alertname%3DGanetiBGPNoInPrefixes [11:50:48] Anyone else seeing a failure on helm lint jobs? https://integration.wikimedia.org/ci/job/helm-lint/34422/console [11:50:58] "You need helm3.11 to run this task. Please install it or run "rake run_locally['default']" to run in a docker container" [11:51:17] (03CR) 10Kamila Součková: [C:03+2] aptrepo: add php85 component [puppet] - 10https://gerrit.wikimedia.org/r/1315940 (https://phabricator.wikimedia.org/T432983) (owner: 10Kamila Součková) [11:51:44] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2077.codfw.wmnet with OS trixie [11:51:49] (03CR) 10Kamila Součková: [C:03+2] package_builder: add pbuilder hook for component/php85 [puppet] - 10https://gerrit.wikimedia.org/r/1315948 (https://phabricator.wikimedia.org/T432983) (owner: 10Kamila Součková) [11:51:52] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12171015 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-be2077.codfw.wmnet with OS trixie completed: - ms-be2077 (**PASS*... [11:53:04] 06SRE, 06Infrastructure-Foundations, 10netops: Core router upgrades 2026 #2 - https://phabricator.wikimedia.org/T431748#12171023 (10ayounsi) 05Open→03Resolved a:03ayounsi All done here. [11:55:01] 06SRE, 06Infrastructure-Foundations, 10netops: cr2-esams rpd failure after enabling bgp 'graceful-shutdown' (June 2026) - https://phabricator.wikimedia.org/T429386#12171034 (10ayounsi) 05Open→03Resolved Closing this task as all the problematic 23.4R2-S7 have been upgraded. The remaining ones will go... [11:56:33] 06SRE, 06Infrastructure-Foundations, 10netops: Core router upgrades 2026 #2 - https://phabricator.wikimedia.org/T431748#12171036 (10ayounsi) 05Resolved→03Open a:05ayounsi→03None [11:59:34] (03CR) 10Ayounsi: [C:03+2] DNS discovery: split responses to eqsin servers based on rack [puppet] - 10https://gerrit.wikimedia.org/r/1313071 (https://phabricator.wikimedia.org/T411617) (owner: 10Ayounsi) [12:00:04] Deploy window Mobileapps/RESTBase/Wikifeeds (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260730T1200) [12:01:56] (03CR) 10Ayounsi: [C:03+2] DNS discovery: split responses to eqsin servers based on rack (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1313071 (https://phabricator.wikimedia.org/T411617) (owner: 10Ayounsi) [12:02:34] (03CR) 10Ayounsi: [C:03+2] eqsin geo-maps: match DNS discovery records [dns] - 10https://gerrit.wikimedia.org/r/1313072 (https://phabricator.wikimedia.org/T411617) (owner: 10Ayounsi) [12:02:49] !log ayounsi@dns1004 START - running authdns-update [12:05:19] !log ayounsi@dns1004 END - running authdns-update [12:07:43] 06SRE, 06Infrastructure-Foundations: Support LVS backend servers using nftables - https://phabricator.wikimedia.org/T433601 (10cmooney) 03NEW p:05Triage→03Medium [12:08:01] 06SRE, 06Infrastructure-Foundations: Support LVS backend servers using nftables - https://phabricator.wikimedia.org/T433601#12171071 (10cmooney) [12:08:02] 06SRE, 06Infrastructure-Foundations: Move URL downloaders to trixie - https://phabricator.wikimedia.org/T427282#12171072 (10cmooney) [12:08:02] Emperor: fyi I deployed https://gerrit.wikimedia.org/r/q/topic:%22T411617+-+eqsin%22 [12:08:32] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for iberker - https://phabricator.wikimedia.org/T433258#12171073 (10IBerker-WMF) Yes, in particular I am trying to see the Metrics FY26-27 dashboard (https://superset.wikimedia.org/superset/dashboard/2c53a25d-aa85-42b8-9d95-b2679822... [12:10:28] 10ops-eqiad, 06SRE, 10Data-Persistence-Backup, 06DC-Ops: db1245 crashed - https://phabricator.wikimedia.org/T431115#12171081 (10jcrespo) I am unable to access now through ssh or https, with both the old or new accounts. Could you check that the passwords have not been resetted to factory defaults, as well... [12:11:26] 10ops-eqsin: Unresponsive management for ganeti5007.mgmt:22 - https://phabricator.wikimedia.org/T433602 (10phaultfinder) 03NEW [12:13:36] !log root@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2036: Maintenance [12:13:57] !log cwilliams@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2047.codfw.wmnet with reason: Maintenance [12:14:05] !log cwilliams@cumin1003 dbctl commit (dc=all): 'Depooling es2047 (T431660)', diff saved to https://phabricator.wikimedia.org/P95793 and previous config saved to /var/cache/conftool/dbconfig/20260730-121404-cwilliams.json [12:14:57] !log dcausse@deploy1003 helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply [12:15:25] !log dcausse@deploy1003 helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply [12:18:31] (03CR) 10CI reject: [V:04-1] Localisation updates from https://translatewiki.net. [phabricator/translations] (wmf/stable) - 10https://gerrit.wikimedia.org/r/1319459 (owner: 10L10n-bot) [12:18:35] !log dcausse@deploy1003 helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply [12:18:43] !log dcausse@deploy1003 helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply [12:18:53] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host ms-be2078.codfw.wmnet with OS trixie [12:19:04] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12171108 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host ms-be2078.codfw.wmnet with OS trixie [12:19:06] !log root@cumin1003 START - Cookbook sre.mysql.pool pool es2047: Maintenance [12:19:36] 10ops-eqiad, 06SRE, 10Data-Persistence-Backup, 06DC-Ops: db1245 crashed - https://phabricator.wikimedia.org/T431115#12171109 (10jcrespo) I've also lost ping again with the main OS host, it was only up for some minutes before losing connectivity. [12:21:18] (03CR) 10FNegri: "Direction looks good to me, thanks for working on this! I left a comment inline." [cookbooks] - 10https://gerrit.wikimedia.org/r/1319088 (https://phabricator.wikimedia.org/T433497) (owner: 10CWilliams) [12:22:07] 10ops-eqiad, 06SRE, 06DC-Ops: wmcs PXE issues - https://phabricator.wikimedia.org/T433538#12171113 (10fgiunchedi) Hi @Jhancock.wm, thank you for reaching out -- I'll be OOO next week, we can postpone to the week of Aug 11th. In general though it is no problem for us to take these hosts out of service with co... [12:23:59] (03PS1) 10Jelto: databub: add dummy values for sql and kafka [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319461 (https://phabricator.wikimedia.org/T427403) [12:26:32] jouncebot: nowandnext [12:26:32] For the next 0 hour(s) and 33 minute(s): Mobileapps/RESTBase/Wikifeeds (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260730T1200) [12:26:32] In 0 hour(s) and 33 minute(s): UTC afternoon backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260730T1300) [12:27:09] (03CR) 10TrainBranchBot: [C:03+2] "Approved by ladsgroup@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319162 (https://phabricator.wikimedia.org/T119117) (owner: 10Ladsgroup) [12:29:05] (03Merged) 10jenkins-bot: Remove $wmg hack for UploadStashMaxAge [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319162 (https://phabricator.wikimedia.org/T119117) (owner: 10Ladsgroup) [12:29:50] (03PS1) 10STran: Instrument link_click server-side instead of client-side [extensions/CheckUser] (wmf/1.47.0-wmf.12) - 10https://gerrit.wikimedia.org/r/1319466 (https://phabricator.wikimedia.org/T433257) [12:30:06] (03PS1) 10STran: Instrument link_click server-side instead of client-side [extensions/CheckUser] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1319467 (https://phabricator.wikimedia.org/T433257) [12:30:15] !log ladsgroup@deploy1003 Started scap sync-world: Backport for [[gerrit:1319162|Remove $wmg hack for UploadStashMaxAge (T119117)]] [12:30:19] T119117: Get rid of $wg = $wmg hack - https://phabricator.wikimedia.org/T119117 [12:30:31] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Thursday, July 30 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployca" [extensions/CheckUser] (wmf/1.47.0-wmf.12) - 10https://gerrit.wikimedia.org/r/1319466 (https://phabricator.wikimedia.org/T433257) (owner: 10STran) [12:30:37] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Thursday, July 30 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployca" [extensions/CheckUser] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1319467 (https://phabricator.wikimedia.org/T433257) (owner: 10STran) [12:31:26] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. [12:31:59] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. [12:32:28] !log ladsgroup@deploy1003 ladsgroup: Backport for [[gerrit:1319162|Remove $wmg hack for UploadStashMaxAge (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [12:32:37] !log brouberol@deploy1003 helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. [12:32:57] !log ladsgroup@deploy1003 ladsgroup: Continuing with deployment [12:33:31] !log brouberol@deploy1003 helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. [12:37:06] !log ladsgroup@deploy1003 Finished scap sync-world: Backport for [[gerrit:1319162|Remove $wmg hack for UploadStashMaxAge (T119117)]] (duration: 06m 51s) [12:37:11] T119117: Get rid of $wg = $wmg hack - https://phabricator.wikimedia.org/T119117 [12:37:52] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage [12:38:55] (03CR) 10CI reject: [V:04-1] Instrument link_click server-side instead of client-side [extensions/CheckUser] (wmf/1.47.0-wmf.12) - 10https://gerrit.wikimedia.org/r/1319466 (https://phabricator.wikimedia.org/T433257) (owner: 10STran) [12:41:04] (03PS2) 10STran: Instrument link_click server-side instead of client-side [extensions/CheckUser] (wmf/1.47.0-wmf.12) - 10https://gerrit.wikimedia.org/r/1319466 (https://phabricator.wikimedia.org/T433257) [12:43:18] (03CR) 10CI reject: [V:04-1] Instrument link_click server-side instead of client-side [extensions/CheckUser] (wmf/1.47.0-wmf.12) - 10https://gerrit.wikimedia.org/r/1319466 (https://phabricator.wikimedia.org/T433257) (owner: 10STran) [12:43:34] (03PS3) 10STran: Instrument link_click server-side instead of client-side [extensions/CheckUser] (wmf/1.47.0-wmf.12) - 10https://gerrit.wikimedia.org/r/1319466 (https://phabricator.wikimedia.org/T433257) [12:43:43] (03PS1) 10Kevin Bazira: ml-services: update tts-section-generator image with pre-batch hardening (render_id, IPA strip, phoneme guard, NeMo 1.2.0) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319476 (https://phabricator.wikimedia.org/T433594) [12:44:06] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage [12:44:16] (03CR) 10CWilliams: sre.mysql: add MySQL-specific base classes (031 comment) [cookbooks] - 10https://gerrit.wikimedia.org/r/1319088 (https://phabricator.wikimedia.org/T433497) (owner: 10CWilliams) [12:45:55] (03CR) 10CI reject: [V:04-1] ml-services: update tts-section-generator image with pre-batch hardening (render_id, IPA strip, phoneme guard, NeMo 1.2.0) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319476 (https://phabricator.wikimedia.org/T433594) (owner: 10Kevin Bazira) [12:54:10] (03CR) 10Jelto: "This is related to https://github.com/helm/helm/pull/12879" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319461 (https://phabricator.wikimedia.org/T427403) (owner: 10Jelto) [12:55:58] (03CR) 10Mpostoronca: [C:03+2] Instrument link_click server-side instead of client-side [extensions/CheckUser] (wmf/1.47.0-wmf.12) - 10https://gerrit.wikimedia.org/r/1319466 (https://phabricator.wikimedia.org/T433257) (owner: 10STran) [12:58:04] (03Merged) 10jenkins-bot: Instrument link_click server-side instead of client-side [extensions/CheckUser] (wmf/1.47.0-wmf.12) - 10https://gerrit.wikimedia.org/r/1319466 (https://phabricator.wikimedia.org/T433257) (owner: 10STran) [12:58:47] (03CR) 10Kevin Bazira: "recheck" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319476 (https://phabricator.wikimedia.org/T433594) (owner: 10Kevin Bazira) [12:59:12] !log cwilliams@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2038.codfw.wmnet with reason: Maintenance [12:59:20] !log cwilliams@cumin1003 dbctl commit (dc=all): 'Depooling es2038 (T431660)', diff saved to https://phabricator.wikimedia.org/P95797 and previous config saved to /var/cache/conftool/dbconfig/20260730-125919-cwilliams.json [13:00:04] Lucas_WMDE, urbanecm, and TheresNoTime: I, the Bot under the Fountain, call upon thee, The Deployer, to do UTC afternoon backport window deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260730T1300). [13:00:04] MichaelG_WMF and Tran: A patch you scheduled for UTC afternoon backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [13:00:11] o/ [13:00:20] * MichaelG_WMF is here [13:00:24] * TheresNoTime cannot deploy today, meeting clash [13:00:32] Hi one of my cherry picks just got merged on accident so I assume at next sync it will deploy [13:01:13] !log filippo@cumin1003 START - Cookbook sre.dns.netbox [13:01:14] o/ [13:01:19] I can deploy if needed [13:01:40] MichaelG_WMF: So when you deploy your patch, mine will I think also incidentally sync. Can I test alongside you? [13:02:13] I can't deploy myself, but if Lucas could do it, that would be great! 🙏 [13:02:32] Note that my change is allmost completely i18n. So it will unfortunately take a long time [13:03:09] yeah, I would deploy the CheckUser backports separately first [13:03:25] RESOLVED: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:03:33] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2078.codfw.wmnet with OS trixie [13:03:42] (03CR) 10TrainBranchBot: [C:03+2] "Approved by lucaswerkmeister-wmde@deploy1003 using scap backport" [extensions/CheckUser] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1319467 (https://phabricator.wikimedia.org/T433257) (owner: 10STran) [13:03:48] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12171245 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-be2078.codfw.wmnet with OS trixie completed: - ms-be2078 (**PASS*... [13:03:49] That being said, there should not be a conflict between the backports and my change. So if needed, they could be tested at the same time from my side. [13:04:45] !log root@cumin1003 START - Cookbook sre.mysql.pool pool es2038: Maintenance [13:06:23] (03Merged) 10jenkins-bot: Instrument link_click server-side instead of client-side [extensions/CheckUser] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1319467 (https://phabricator.wikimedia.org/T433257) (owner: 10STran) [13:06:46] !log lucaswerkmeister-wmde@deploy1003 Started scap sync-world: Backport for [[gerrit:1319466|Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467|Instrument link_click server-side instead of client-side (T433257)]] [13:06:51] T433257: Instrument link_click server-side instead of client-side - https://phabricator.wikimedia.org/T433257 [13:07:45] !log filippo@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" [13:08:04] !log filippo@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003" [13:08:04] !log filippo@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [13:08:15] !log root@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2047: Maintenance [13:08:33] !log filippo@cumin1003 START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie [13:08:43] !log lucaswerkmeister-wmde@deploy1003 lucaswerkmeister-wmde, stran: Backport for [[gerrit:1319466|Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467|Instrument link_click server-side instead of client-side (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [13:08:49] Tran: please test ^^ [13:08:49] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12171261 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by filippo@cumin1003 for host cloudvirt1048.eqiad.wmnet with OS trixie [13:08:54] testing now [13:10:07] (03CR) 10Kareid: [C:03+1] Test Kitchen UI: Deploy v1.5.0 release to staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319178 (https://phabricator.wikimedia.org/T432304) (owner: 10Clare Ming) [13:10:21] !log root@cumin1003 END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2038: Maintenance [13:10:26] !log root@cumin1003 START - Cookbook sre.mysql.pool pool es2038: Maintenance [13:10:56] !log aokoth@cumin1003 START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - T433606 [13:11:43] Lucas_WMDE: lgtm [13:13:16] !log lucaswerkmeister-wmde@deploy1003 lucaswerkmeister-wmde, stran: Continuing with deployment [13:13:17] thanks [13:15:38] (03CR) 10JMeybohm: [C:03+1] databub: add dummy values for sql and kafka [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319461 (https://phabricator.wikimedia.org/T427403) (owner: 10Jelto) [13:16:16] !log bking@cumin2003 conftool action : set/pooled=yes; selector: name=wdqs2022\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main [13:16:44] (03CR) 10JMeybohm: [C:03+1] "You may remove the -1 vote from jenkins and override with C+2, V+2 to get out of this catch-22" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1318569 (https://phabricator.wikimedia.org/T388390) (owner: 10Jelto) [13:17:16] (03CR) 10Jelto: [C:03+2] databub: add dummy values for sql and kafka [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319461 (https://phabricator.wikimedia.org/T427403) (owner: 10Jelto) [13:17:17] !log lucaswerkmeister-wmde@deploy1003 Finished scap sync-world: Backport for [[gerrit:1319466|Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467|Instrument link_click server-side instead of client-side (T433257)]] (duration: 10m 31s) [13:17:20] !log cwilliams@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1036.eqiad.wmnet with reason: Maintenance [13:17:21] T433257: Instrument link_click server-side instead of client-side - https://phabricator.wikimedia.org/T433257 [13:17:28] !log cwilliams@cumin1003 dbctl commit (dc=all): 'Depooling es1036 (T431660)', diff saved to https://phabricator.wikimedia.org/P95800 and previous config saved to /var/cache/conftool/dbconfig/20260730-131727-cwilliams.json [13:17:33] (03CR) 10TrainBranchBot: [C:03+2] "Approved by lucaswerkmeister-wmde@deploy1003 using scap backport" [extensions/WikimediaMessages] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1319449 (https://phabricator.wikimedia.org/T430604) (owner: 10Michael Große) [13:18:00] (03PS3) 10Jelto: helmfile.d/* use helm3.17 and helm3.19 helmBinary [deployment-charts] - 10https://gerrit.wikimedia.org/r/1318569 (https://phabricator.wikimedia.org/T388390) [13:18:15] (03PS2) 10Jelto: databub: add dummy values for sql and kafka [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319461 (https://phabricator.wikimedia.org/T427403) [13:18:38] Thanks! [13:18:57] (03PS1) 10Ladsgroup: thumbor: Restrict access to thumb file types [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319481 (https://phabricator.wikimedia.org/T430528) [13:20:00] (03CR) 10CI reject: [V:04-1] helmfile.d/* use helm3.17 and helm3.19 helmBinary [deployment-charts] - 10https://gerrit.wikimedia.org/r/1318569 (https://phabricator.wikimedia.org/T388390) (owner: 10Jelto) [13:20:43] !log aokoth@cumin1003 END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - T433606 [13:21:25] (03CR) 10CI reject: [V:04-1] thumbor: Restrict access to thumb file types [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319481 (https://phabricator.wikimedia.org/T430528) (owner: 10Ladsgroup) [13:22:20] (03CR) 10Jelto: [V:03+2 C:03+2] helmfile.d/* use helm3.17 and helm3.19 helmBinary [deployment-charts] - 10https://gerrit.wikimedia.org/r/1318569 (https://phabricator.wikimedia.org/T388390) (owner: 10Jelto) [13:22:42] !log root@cumin1003 START - Cookbook sre.mysql.pool pool es1036: Maintenance [13:22:48] !log aokoth@cumin1003 START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - T433606 [13:23:24] (03PS2) 10Kevin Bazira: ml-services: update tts-section-generator image with pre-batch hardening (render_id, IPA strip, phoneme guard, NeMo 1.2.0) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319476 (https://phabricator.wikimedia.org/T433594) [13:25:40] (03CR) 10Ladsgroup: "recheck" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319481 (https://phabricator.wikimedia.org/T430528) (owner: 10Ladsgroup) [13:26:22] PROBLEM - Postgres Replication Lag on puppetdb2003 is CRITICAL: POSTGRES_HOT_STANDBY_DELAY CRITICAL: DB puppetdb (host:localhost) 44697856 and 2 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [13:26:40] (03Merged) 10jenkins-bot: databub: add dummy values for sql and kafka [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319461 (https://phabricator.wikimedia.org/T427403) (owner: 10Jelto) [13:27:07] (03PS3) 10Arthur taylor: wikidata-query-gui: bump to latest version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319451 (https://phabricator.wikimedia.org/T433233) [13:27:20] (03PS1) 10Mszwarc: plwikiquote: Set AutoConfirmCount to 25 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319482 (https://phabricator.wikimedia.org/T433541) [13:27:21] (03CR) 10Bking: [C:03+2] opensearch: ship the cirrus ECS server logs via rsyslog imfile [puppet] - 10https://gerrit.wikimedia.org/r/1311424 (https://phabricator.wikimedia.org/T324335) (owner: 10Btullis) [13:27:22] RECOVERY - Postgres Replication Lag on puppetdb2003 is OK: POSTGRES_HOT_STANDBY_DELAY OK: DB puppetdb (host:localhost) 46032 and 0 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [13:27:26] (03PS2) 10Ladsgroup: thumbor: Restrict access to thumb file types [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319481 (https://phabricator.wikimedia.org/T430528) [13:27:43] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Thursday, July 30 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployca" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319482 (https://phabricator.wikimedia.org/T433541) (owner: 10Mszwarc) [13:27:49] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Thursday, July 30 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployca" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309552 (https://phabricator.wikimedia.org/T430512) (owner: 10Mszwarc) [13:28:19] !log root@cumin1003 END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1036: Maintenance [13:28:24] !log root@cumin1003 START - Cookbook sre.mysql.pool pool es1036: Maintenance [13:29:26] Hi! I added two config patches to the window. I see there is a long (i18n) deployment ongoing, but if there's time for me in the end, please ping me :) [13:29:33] (03PS3) 10CWilliams: sre.mysql: add MySQL-specific base classes [cookbooks] - 10https://gerrit.wikimedia.org/r/1319088 (https://phabricator.wikimedia.org/T433497) [13:31:49] (03Merged) 10jenkins-bot: SpecialCreateAccount: make user policy-link available again [extensions/WikimediaMessages] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1319449 (https://phabricator.wikimedia.org/T430604) (owner: 10Michael Große) [13:32:14] !log lucaswerkmeister-wmde@deploy1003 Started scap sync-world: Backport for [[gerrit:1319449|SpecialCreateAccount: make user policy-link available again (T430604)]] [13:32:17] T430604: Account Creation form: Simplify username guidance based on experiment data and user research - https://phabricator.wikimedia.org/T430604 [13:32:43] !log aokoth@cumin1003 END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - T433606 [13:35:03] (03PS1) 10Ayounsi: Fastnetmon bump threshold_pps to 2Mpps and add threshold_udp/icmp [puppet] - 10https://gerrit.wikimedia.org/r/1319483 (https://phabricator.wikimedia.org/T431683) [13:35:15] (03CR) 10Majavah: [C:03+1] Puppet 8: Replace legacy facts in profile wmcs [puppet] - 10https://gerrit.wikimedia.org/r/1315010 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [13:38:26] (03CR) 10Kevin Bazira: [C:03+2] ml-services: update tts-section-generator image with pre-batch hardening (render_id, IPA strip, phoneme guard, NeMo 1.2.0) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319476 (https://phabricator.wikimedia.org/T433594) (owner: 10Kevin Bazira) [13:41:17] (03CR) 10Lucas Werkmeister (WMDE): [C:03+1] wikidata-query-gui: bump to latest version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319451 (https://phabricator.wikimedia.org/T433233) (owner: 10Arthur taylor) [13:41:21] (03Merged) 10jenkins-bot: ml-services: update tts-section-generator image with pre-batch hardening (render_id, IPA strip, phoneme guard, NeMo 1.2.0) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319476 (https://phabricator.wikimedia.org/T433594) (owner: 10Kevin Bazira) [13:41:38] 06SRE, 06ServiceOps new, 10VisualEditor, 10VisualEditor Suggestion Mode, and 3 others: Deploy Headless VE in k8s for technical pilot - https://phabricator.wikimedia.org/T431497#12171378 (10Ottomata) Thank you! > If you mean "should" in the sense of "we believe this is the correct course of action, as a pr... [13:43:24] (03CR) 10FNegri: [C:03+1] "Looks good to me, IMHO this can be merged. Further improvements can be done in future patches when migrating the existing cookbooks to use" [cookbooks] - 10https://gerrit.wikimedia.org/r/1319088 (https://phabricator.wikimedia.org/T433497) (owner: 10CWilliams) [13:44:37] (03PS1) 10Cathal Mooney: LVS IPIP encapsulation: support backend servers using nftables [puppet] - 10https://gerrit.wikimedia.org/r/1319486 (https://phabricator.wikimedia.org/T433601) [13:44:44] (03PS1) 10STran: SI: Add performer to server-side link_click event [extensions/CheckUser] (wmf/1.47.0-wmf.12) - 10https://gerrit.wikimedia.org/r/1319487 (https://phabricator.wikimedia.org/T433257) [13:45:06] !log kevinbazira@deploy1003 helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . [13:45:19] (03CR) 10CI reject: [V:04-1] LVS IPIP encapsulation: support backend servers using nftables [puppet] - 10https://gerrit.wikimedia.org/r/1319486 (https://phabricator.wikimedia.org/T433601) (owner: 10Cathal Mooney) [13:45:31] (03PS1) 10STran: SI: Add performer to server-side link_click event [extensions/CheckUser] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1319488 (https://phabricator.wikimedia.org/T433257) [13:45:48] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Thursday, July 30 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployca" [extensions/CheckUser] (wmf/1.47.0-wmf.12) - 10https://gerrit.wikimedia.org/r/1319487 (https://phabricator.wikimedia.org/T433257) (owner: 10STran) [13:45:57] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Thursday, July 30 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployca" [extensions/CheckUser] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1319488 (https://phabricator.wikimedia.org/T433257) (owner: 10STran) [13:47:17] Msz2001: If there's time after you, can you ping me so I can deploy a follow-up to my earlier code? 😅 My code does what I wanted but it's missing an attribute [13:47:27] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1311426 (https://phabricator.wikimedia.org/T324335) (owner: 10Btullis) [13:47:47] (03PS9) 10Btullis: cirrus: remove the log4j syslog appender log shipping path [puppet] - 10https://gerrit.wikimedia.org/r/1311426 (https://phabricator.wikimedia.org/T324335) [13:48:03] (03CR) 10Bking: "check puppet" [puppet] - 10https://gerrit.wikimedia.org/r/1311426 (https://phabricator.wikimedia.org/T324335) (owner: 10Btullis) [13:48:05] (03CR) 10Majavah: [V:03+1] "PCC SUCCESS (NOOP 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9090/console" [puppet] - 10https://gerrit.wikimedia.org/r/1315004 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [13:48:13] (03CR) 10CWilliams: [C:03+2] sre.mysql: add MySQL-specific base classes [cookbooks] - 10https://gerrit.wikimedia.org/r/1319088 (https://phabricator.wikimedia.org/T433497) (owner: 10CWilliams) [13:48:24] (03CR) 10Majavah: [V:03+1 C:03+1] Puppet 8: Replace legacy facts in module dynamicproxy [puppet] - 10https://gerrit.wikimedia.org/r/1315004 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [13:48:24] !log kevinbazira@deploy1003 helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . [13:48:33] Tran: I can let you first, my patches are not urgent at all (one is cleanup from before my vacation, and the other was asked for by a friend) [13:49:13] FIRING: NELHigh: Elevated Network Error Logging events (tcp.timed_out) #page - https://wikitech.wikimedia.org/wiki/Network_monitoring#NEL_alerts - https://logstash.wikimedia.org/goto/5c8f4ca1413eda33128e5c5a35da7e28 - https://alerts.wikimedia.org/?q=alertname%3DNELHigh [13:49:17] (03CR) 10Majavah: [C:03+1] Puppet 8: Replace legacy facts in profile cloudceph [puppet] - 10https://gerrit.wikimedia.org/r/1315005 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [13:49:46] !log ssastry@deploy1003 helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply [13:49:56] !log lucaswerkmeister-wmde@deploy1003 migr, lucaswerkmeister-wmde: Backport for [[gerrit:1319449|SpecialCreateAccount: make user policy-link available again (T430604)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [13:49:57] (03CR) 10Majavah: [C:03+1] Puppet 8: Replace legacy facts in profile openstack [puppet] - 10https://gerrit.wikimedia.org/r/1315006 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [13:50:02] T430604: Account Creation form: Simplify username guidance based on experiment data and user research - https://phabricator.wikimedia.org/T430604 [13:50:09] oh god, that’s still running, I almost forgot [13:50:12] MichaelG_WMF: please test [13:50:14] FIRING: NELByCountryHigh: Elevated Network Error Logging events (tcp.timed_out from RU) - https://wikitech.wikimedia.org/wiki/Network_monitoring#NEL_alerts - https://logstash.wikimedia.org/goto/5c8f4ca1413eda33128e5c5a35da7e28 - https://alerts.wikimedia.org/?q=alertname%3DNELByCountryHigh [13:50:21] Lucas_WMDE: testing [13:50:30] (03CR) 10Majavah: [C:03+1] Puppet 8: Replace legacy facts in module openstack [puppet] - 10https://gerrit.wikimedia.org/r/1315003 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [13:50:43] @Lucas_WMDE The train blocker has now been resolved: https://phabricator.wikimedia.org/T433457#12168486. Since there is an active window, what is the best way to deploy the fix? [13:50:43] group-0 would need to be deployed first I think, we ensure validation error rate has gone down. And group-1 and 2 can be deployed with the train (later today)? [13:50:43] cc: @otto [13:51:20] cc: @ottomata [13:51:37] (03Merged) 10jenkins-bot: sre.mysql: add MySQL-specific base classes [cookbooks] - 10https://gerrit.wikimedia.org/r/1319088 (https://phabricator.wikimedia.org/T433497) (owner: 10CWilliams) [13:51:42] (03PS1) 10Jforrester: Monolog: Strip diagnostic context from Monolog-based EventBus events [extensions/EventBus] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1319489 (https://phabricator.wikimedia.org/T433457) [13:52:16] akhatun: I've cherry-picked it to wmf.13. Yes, that will need a production deployment and checking, and then if it works the train can roll to group1 (and group2 hopefully). [13:52:17] Lucas_WMDE: Confirmed that it now works as expected! Good to move on from my side 👍 [13:52:21] 06SRE: Various services hardcode api.svc.eqiad.wmnet - https://phabricator.wikimedia.org/T285518#12171426 (10jijiki) 05Open→03Invalid Many things have changed since this task, we can close it. [13:52:29] (03PS2) 10Cathal Mooney: LVS IPIP encapsulation: support backend servers using nftables [puppet] - 10https://gerrit.wikimedia.org/r/1319486 (https://phabricator.wikimedia.org/T433601) [13:52:30] !log lucaswerkmeister-wmde@deploy1003 migr, lucaswerkmeister-wmde: Continuing with deployment [13:52:31] thanks [13:53:06] (03CR) 10CI reject: [V:04-1] LVS IPIP encapsulation: support backend servers using nftables [puppet] - 10https://gerrit.wikimedia.org/r/1319486 (https://phabricator.wikimedia.org/T433601) (owner: 10Cathal Mooney) [13:54:27] (03PS1) 10Ozge: ml-services: Switch jina-embeddings-staging to eager attention Bump the staging image and set ATTN_IMPLEMENTATION=eager for vLLM. Bug: T432717 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319490 (https://phabricator.wikimedia.org/T432717) [13:54:37] (03PS3) 10Cathal Mooney: LVS IPIP encapsulation: support backend servers using nftables [puppet] - 10https://gerrit.wikimedia.org/r/1319486 (https://phabricator.wikimedia.org/T433601) [13:54:47] (03CR) 10Ozge: [V:03+2 C:03+2] "testing." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319490 (https://phabricator.wikimedia.org/T432717) (owner: 10Ozge) [13:55:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.4% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [13:55:26] !log ozge@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . [13:56:14] !log root@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2038: Maintenance [13:56:30] (03CR) 10Bking: [C:03+2] cirrus: remove the log4j syslog appender log shipping path [puppet] - 10https://gerrit.wikimedia.org/r/1311426 (https://phabricator.wikimedia.org/T324335) (owner: 10Btullis) [13:56:30] !log ssastry@deploy1003 helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply [13:56:31] !log ssastry@deploy1003 helmfile [codfw] START helmfile.d/services/mw-parsoid: apply [13:56:35] !log cwilliams@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2040.codfw.wmnet with reason: Maintenance [13:56:44] !log cwilliams@cumin1003 dbctl commit (dc=all): 'Depooling es2040 (T431660)', diff saved to https://phabricator.wikimedia.org/P95806 and previous config saved to /var/cache/conftool/dbconfig/20260730-135643-cwilliams.json [13:57:28] (03PS4) 10Cathal Mooney: LVS IPIP encapsulation: support backend servers using nftables [puppet] - 10https://gerrit.wikimedia.org/r/1319486 (https://phabricator.wikimedia.org/T433601) [13:57:43] @James_F would we deploy to group-0 in the current window? [13:58:59] Access to Wikimedia Commons has been lost from Russia. Is this a government block or SRE actions? [13:59:23] akhatun: The cherry-pick of the patch needs to be applied to the production branch. [13:59:39] 06SRE, 10SRE-Access-Requests, 10LDAP-Access-Requests, 13Patch-For-Review: Grant Access to wmf for Chandler Diggs - https://phabricator.wikimedia.org/T433302#12171465 (10jijiki) >>! In T433302#12168174, @Dzahn wrote: > @jijiki I think even though it's self-service we must still document it in the repo. At l... [13:59:51] filippo@cumin1003 reimage (PID 3814572) is awaiting input [14:00:03] (03PS1) 10Ozge: ml-services: jina-embeddings-staging getting back to the previous image [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319491 (https://phabricator.wikimedia.org/T432717) [14:00:26] (03CR) 10Ozge: [V:03+2 C:03+2] "testing" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319491 (https://phabricator.wikimedia.org/T432717) (owner: 10Ozge) [14:01:37] !log ozge@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . [14:01:53] 06SRE, 06Infrastructure-Foundations, 13Patch-For-Review: Support LVS backend servers using nftables - https://phabricator.wikimedia.org/T433601#12171476 (10ssingh) The primary change that I think we need to make is in `modules/profile/manifests/lvs/realserver/ipip.pp`, where currently we have support for onl... [14:03:02] !log ssastry@deploy1003 helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply [14:03:12] !log root@cumin1003 START - Cookbook sre.mysql.pool pool es2040: Maintenance [14:03:14] @elukey , @arnoldokoth, Access to Wikimedia Commons has been lost from Russia. Is this a government block or SRE actions? See topic on news forum https://w.wiki/StzH [14:03:55] !log lucaswerkmeister-wmde@deploy1003 Finished scap sync-world: Backport for [[gerrit:1319449|SpecialCreateAccount: make user policy-link available again (T430604)]] (duration: 31m 41s) [14:03:59] T430604: Account Creation form: Simplify username guidance based on experiment data and user research - https://phabricator.wikimedia.org/T430604 [14:04:04] jouncebot: now [14:04:04] No deployments scheduled for the next 0 hour(s) and 25 minute(s) [14:04:18] Tran: over to you? or to whoever wants to get the train back on track, maybe [14:04:23] * Lucas_WMDE hasn’t followed the discussion closely [14:04:30] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-07-03 - 2026-07-31): Degraded RAID on an-worker1191 - https://phabricator.wikimedia.org/T431828#12171489 (10VRiley-WMF) [14:04:43] The train is probably most important as it's stalled? But if they're not ready I can deploy and test quickly-ish. [14:05:22] (I'll start if I don't hear back in a minute or two) [14:05:39] 06SRE, 06Infrastructure-Foundations, 06Traffic, 13Patch-For-Review: Support LVS backend servers using nftables - https://phabricator.wikimedia.org/T433601#12171495 (10ssingh) [14:06:26] k, starting [14:06:59] (03CR) 10TrainBranchBot: [C:03+2] "Approved by stran@deploy1003 using scap backport" [extensions/CheckUser] (wmf/1.47.0-wmf.12) - 10https://gerrit.wikimedia.org/r/1319487 (https://phabricator.wikimedia.org/T433257) (owner: 10STran) [14:06:59] (03CR) 10TrainBranchBot: [C:03+2] "Approved by stran@deploy1003 using scap backport" [extensions/CheckUser] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1319488 (https://phabricator.wikimedia.org/T433257) (owner: 10STran) [14:07:37] Iluvatar: hi! Sadly it is not under our control - https://www.wikimediastatus.net/incidents/wfdnbnv00w3r [14:08:49] !log root@cumin1003 END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2040: Maintenance [14:08:54] !log root@cumin1003 START - Cookbook sre.mysql.pool pool es2040: Maintenance [14:09:05] ok, thanks! [14:09:12] (03Merged) 10jenkins-bot: SI: Add performer to server-side link_click event [extensions/CheckUser] (wmf/1.47.0-wmf.12) - 10https://gerrit.wikimedia.org/r/1319487 (https://phabricator.wikimedia.org/T433257) (owner: 10STran) [14:09:19] (03Merged) 10jenkins-bot: SI: Add performer to server-side link_click event [extensions/CheckUser] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1319488 (https://phabricator.wikimedia.org/T433257) (owner: 10STran) [14:09:28] FIRING: [3x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-3/1/2 (Transport: Hurricane Electric (dc4841.dal5)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [14:09:46] !log stran@deploy1003 Started scap sync-world: Backport for [[gerrit:1319487|SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488|SI: Add performer to server-side link_click event (T433257)]] [14:09:49] T433257: Instrument link_click server-side instead of client-side - https://phabricator.wikimedia.org/T433257 [14:10:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 21.52% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [14:11:23] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host ms-be2079.codfw.wmnet with OS trixie [14:11:37] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12171514 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host ms-be2079.codfw.wmnet with OS trixie [14:12:23] 06SRE, 10LDAP-Access-Requests: Grant Access to NDA for jaleman-vdr-wmf - https://phabricator.wikimedia.org/T433417#12171516 (10jijiki) @JAATPH in order to move forward we need to have your NDA on file, as @Dzahn suggested, as well as either of @JArguello-WMF or @HShaikh to comment on this task that they approv... [14:13:30] (03CR) 10Ssingh: "Very nice work, will take some time to review it though since I don't understand everything anyway. Starting by running PCC." [puppet] - 10https://gerrit.wikimedia.org/r/1319486 (https://phabricator.wikimedia.org/T433601) (owner: 10Cathal Mooney) [14:13:37] !log stran@deploy1003 stran: Backport for [[gerrit:1319487|SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488|SI: Add performer to server-side link_click event (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [14:13:38] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for iberker - https://phabricator.wikimedia.org/T433258#12171525 (10jijiki) [14:14:02] testing now [14:14:23] !log root@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1036: Maintenance [14:14:28] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [14:14:30] (03CR) 10Ottomata: [C:03+2] Monolog: Strip diagnostic context from Monolog-based EventBus events [extensions/EventBus] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1319489 (https://phabricator.wikimedia.org/T433457) (owner: 10Jforrester) [14:14:32] !log cwilliams@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1038.eqiad.wmnet with reason: Maintenance [14:14:40] !log cwilliams@cumin1003 dbctl commit (dc=all): 'Depooling es1038 (T431660)', diff saved to https://phabricator.wikimedia.org/P95810 and previous config saved to /var/cache/conftool/dbconfig/20260730-141439-cwilliams.json [14:14:47] lgtm, continuing [14:14:50] (03PS1) 10Bking: cirrus: enable ship_server_json_logs for relforge and cloudelastic [puppet] - 10https://gerrit.wikimedia.org/r/1319497 (https://phabricator.wikimedia.org/T324335) [14:14:53] !log stran@deploy1003 stran: Continuing with deployment [14:15:14] (03CR) 10Bking: "check puppet" [puppet] - 10https://gerrit.wikimedia.org/r/1319497 (https://phabricator.wikimedia.org/T324335) (owner: 10Bking) [14:16:01] ottomata: Are you deploying that after Tran? [14:17:27] (03CR) 10Cathal Mooney: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1319486 (https://phabricator.wikimedia.org/T433601) (owner: 10Cathal Mooney) [14:20:03] !log root@cumin1003 START - Cookbook sre.mysql.pool pool es1038: Maintenance [14:20:41] (03CR) 10Brouberol: [C:03+1] cirrus: enable ship_server_json_logs for relforge and cloudelastic [puppet] - 10https://gerrit.wikimedia.org/r/1319497 (https://phabricator.wikimedia.org/T324335) (owner: 10Bking) [14:21:05] !log stran@deploy1003 Finished scap sync-world: Backport for [[gerrit:1319487|SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488|SI: Add performer to server-side link_click event (T433257)]] (duration: 11m 19s) [14:21:11] T433257: Instrument link_click server-side instead of client-side - https://phabricator.wikimedia.org/T433257 [14:21:36] (03CR) 10Bking: [C:03+2] cirrus: enable ship_server_json_logs for relforge and cloudelastic [puppet] - 10https://gerrit.wikimedia.org/r/1319497 (https://phabricator.wikimedia.org/T324335) (owner: 10Bking) [14:21:39] I'm done. if there's time, Msz2001 also wanted to deploy. Thanks for letting me step ahead of you 🙇 [14:21:55] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for iberker - https://phabricator.wikimedia.org/T433258#12171550 (10jijiki) @IBerker-WMF since that is the case, please follow the instructions in the task, specifically, please sign the L3 Wikimedia Server Access and ask your man... [14:22:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.6% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [14:22:19] @James_F @ottomata and I would like to, yes otto is in a meeting, I would need help from someone with deploying this. Haven't done this myself before. We need a merge for the patch. It has been +2'ed. Are we good to verify+2? [14:22:20] (03CR) 10TrainBranchBot: [C:03+2] "Approved by mszwarc@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319482 (https://phabricator.wikimedia.org/T433541) (owner: 10Mszwarc) [14:22:21] (03CR) 10TrainBranchBot: [C:03+2] "Approved by mszwarc@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309552 (https://phabricator.wikimedia.org/T430512) (owner: 10Mszwarc) [14:22:40] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in profile wmcs [puppet] - 10https://gerrit.wikimedia.org/r/1315010 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [14:22:43] akhatun: Do not V+2 ever. That bypasses CI. [14:22:55] Should I abort my deployment? I started it, as nobody responded to James for a few minutes [14:22:57] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in module dynamicproxy [puppet] - 10https://gerrit.wikimedia.org/r/1315004 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [14:23:08] But it's still before merged, so I can abort [14:23:10] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in profile cloudceph [puppet] - 10https://gerrit.wikimedia.org/r/1315005 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [14:23:23] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in profile openstack [puppet] - 10https://gerrit.wikimedia.org/r/1315006 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [14:23:25] (03Merged) 10jenkins-bot: plwikiquote: Set AutoConfirmCount to 25 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319482 (https://phabricator.wikimedia.org/T433541) (owner: 10Mszwarc) [14:23:29] (03Merged) 10jenkins-bot: Revert "Temporarily change plwiki tagline for 1.7M articles" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309552 (https://phabricator.wikimedia.org/T430512) (owner: 10Mszwarc) [14:23:35] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in module openstack [puppet] - 10https://gerrit.wikimedia.org/r/1315003 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [14:23:44] Ack. [14:23:46] !log mszwarc@deploy1003 Started scap sync-world: Backport for [[gerrit:1319482|plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552|Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] [14:23:53] T433541: Set $wgAutoconfirmCount to 25 on plwikiquote - https://phabricator.wikimedia.org/T433541 [14:23:53] T430512: Temporarily change plwiki tagline to reflect 1.7M articles - https://phabricator.wikimedia.org/T430512 [14:25:36] (03CR) 10Klausman: [C:03+1] ml-serve: enlarge kubelet LV on GPU nodes [puppet] - 10https://gerrit.wikimedia.org/r/1307368 (https://phabricator.wikimedia.org/T431017) (owner: 10Dpogorzelski) [14:25:39] !log root@cumin1003 END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1038: Maintenance [14:25:43] !log root@cumin1003 START - Cookbook sre.mysql.pool pool es1038: Maintenance [14:25:47] !log mszwarc@deploy1003 mszwarc: Backport for [[gerrit:1319482|plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552|Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [14:25:56] (03PS1) 10Ssingh: conftool: switch realservers for urldownloader [puppet] - 10https://gerrit.wikimedia.org/r/1319502 (https://phabricator.wikimedia.org/T429175) [14:26:13] !log mszwarc@deploy1003 mszwarc: Continuing with deployment [14:27:10] sorry James_F akhatun yes I am in a meeting. [14:27:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.6% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [14:27:20] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie [14:27:22] I can help in 30 mins, or anyone with powers can deploy now [14:27:31] ottomata: But you merged into the production branch now? [14:27:35] 10ops-codfw, 06SRE, 10SRE-swift-storage, 06Data-Persistence, 06DC-Ops: Q4:rack/setup/install ms-be209[7,8] - https://phabricator.wikimedia.org/T424892#12171592 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-be2097.codfw.wmnet with OS trixie [14:28:22] (03CR) 10JHathaway: [C:03+1] profile::spicerack: remove unnecessary filter for empty values [puppet] - 10https://gerrit.wikimedia.org/r/1307080 (https://phabricator.wikimedia.org/T429699) (owner: 10Elukey) [14:28:42] James_F: I +2ed , akhatun asked me in slack, i haven't been following IRC closely. Sorry I thought i was helping! [14:28:45] (03Merged) 10jenkins-bot: Monolog: Strip diagnostic context from Monolog-based EventBus events [extensions/EventBus] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1319489 (https://phabricator.wikimedia.org/T433457) (owner: 10Jforrester) [14:28:51] ^ welp it is merged. [14:30:05] Deploy window Test Kitchen Experiment Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260730T1430) [14:30:17] !log mszwarc@deploy1003 Finished scap sync-world: Backport for [[gerrit:1319482|plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552|Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] (duration: 06m 31s) [14:30:23] T433541: Set $wgAutoconfirmCount to 25 on plwikiquote - https://phabricator.wikimedia.org/T433541 [14:30:23] T430512: Temporarily change plwiki tagline to reflect 1.7M articles - https://phabricator.wikimedia.org/T430512 [14:30:29] I finished my deployment [14:31:01] apologies! I messed up the timeline it seems. I will try to find someone... now. [14:32:02] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage [14:32:33] (03CR) 10Fabfur: [C:03+1] conftool: switch realservers for urldownloader [puppet] - 10https://gerrit.wikimedia.org/r/1319502 (https://phabricator.wikimedia.org/T429175) (owner: 10Ssingh) [14:32:58] (03CR) 10Fabfur: [C:03+2] Add key to X-Analytics if a thumbnail was generated during the request [puppet] - 10https://gerrit.wikimedia.org/r/1318088 (https://phabricator.wikimedia.org/T433075) (owner: 10Cparle) [14:36:12] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS trixie [14:36:26] 10ops-codfw, 06SRE, 10SRE-swift-storage, 06Data-Persistence, 06DC-Ops: Q4:rack/setup/install ms-be209[7,8] - https://phabricator.wikimedia.org/T424892#12171655 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host ms-be2098.codfw.wmnet with OS trixie [14:38:17] (03CR) 10Ssingh: [C:03+2] conftool: switch realservers for urldownloader [puppet] - 10https://gerrit.wikimedia.org/r/1319502 (https://phabricator.wikimedia.org/T429175) (owner: 10Ssingh) [14:39:21] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage [14:39:31] (03CR) 10Majavah: [V:03+1] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9092/co" [puppet] - 10https://gerrit.wikimedia.org/r/1315011 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [14:39:51] (03CR) 10Hnowlan: "Just one new note, apologies" [puppet] - 10https://gerrit.wikimedia.org/r/1306792 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [14:39:53] !log sukhe@lvs1020:~$ sudo systemctl restart pybal.service [14:39:55] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:40:03] (03CR) 10Majavah: [V:03+1 C:03+1] Puppet 8: Replace legacy facts in profile terraform [puppet] - 10https://gerrit.wikimedia.org/r/1315011 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [14:40:23] (03CR) 10Majavah: [V:03+1] "PCC SUCCESS (NOOP 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9093/console" [puppet] - 10https://gerrit.wikimedia.org/r/1315013 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [14:40:31] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [14:41:13] (03CR) 10Scott French: "Nice! I'd missed that this was already fixed :)" [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1319453 (https://phabricator.wikimedia.org/T433312) (owner: 10Kamila Součková) [14:41:27] PROBLEM - PyBal IPVS diff check on lvs2014 is CRITICAL: (CRITICAL: Mismatch between IPVS and PyBal https://wikitech.wikimedia.org/wiki/PyBal [14:42:18] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-be1081.eqiad.wmnet with OS trixie [14:42:28] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12171693 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-be1081.eqiad.wmnet with OS trixie [14:42:46] !log sukhe@puppetserver1001 conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid [14:42:50] !log sukhe@puppetserver1001 conftool action : set/weight=1; selector: cluster=urldownloader,service=squid [14:43:07] * tchin Ok so just jumping in, I should just backport the eventbus change now I guess? [14:43:34] whoops my habit of pressing ctrl before sending a message bites me again [14:44:46] PROBLEM - PyBal IPVS diff check on lvs2013 is CRITICAL: (CRITICAL: Mismatch between IPVS and PyBal https://wikitech.wikimedia.org/wiki/PyBal [14:45:00] !log tchin@deploy1003 Started scap sync-world: Backport for [[gerrit:1319489|Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] [14:45:04] T433457: `mediawiki.api-request` validation errors from Extension:OAuth `context.oauth_consumer_*` logging context - https://phabricator.wikimedia.org/T433457 [14:45:19] backporting now [14:45:22] @tchin yes, thank you! @James_F this is being deployed. [14:45:36] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - urldownloader_8080: Servers urldownloader1003.wikimedia.org are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [14:46:10] PROBLEM - PyBal IPVS diff check on lvs1019 is CRITICAL: (CRITICAL: Mismatch between IPVS and PyBal https://wikitech.wikimedia.org/wiki/PyBal [14:46:59] !log tchin@deploy1003 jforrester, tchin: Backport for [[gerrit:1319489|Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [14:47:04] !log mvernon@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage [14:47:22] PROBLEM - SSH on crm2001 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/SSH/monitoring [14:47:44] !log tchin@deploy1003 jforrester, tchin: Continuing with deployment [14:48:18] RECOVERY - SSH on crm2001 is OK: SSH OK - OpenSSH_9.2p1 Debian-2+deb12u10 (protocol 2.0) https://wikitech.wikimedia.org/wiki/SSH/monitoring [14:49:05] FIRING: HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-rollback - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs-next - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [14:49:44] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [14:49:57] (03CR) 10Cwhite: [C:04-2] "Adding search folks for collab. They manipulate the plugins rather than disable the security plugin via config, but per comment, this gat" [puppet] - 10https://gerrit.wikimedia.org/r/1318792 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [14:50:02] akhatun: Aha, awesome, thank you! Sorry, am in meetings too. [14:50:19] 06SRE, 10LDAP-Access-Requests: Grant Access to NDA for jaleman-vdr-wmf - https://phabricator.wikimedia.org/T433417#12171713 (10JArguello-WMF) Hi @Dzahn and @jijiki thanks for the reply! @JAATPH is a contractor, not WMF staff. I approve this request. Where can we get the NDA file to be filled out? [14:51:25] PROBLEM - Postgres Replication Lag on puppetdb2003 is CRITICAL: POSTGRES_HOT_STANDBY_DELAY CRITICAL: DB puppetdb (host:localhost) 415738640 and 22 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [14:51:31] 06SRE, 10LDAP-Access-Requests: Grant Access to ciadmin for Vaughn Walters - https://phabricator.wikimedia.org/T433615 (10hashar) 03NEW [14:51:48] !log tchin@deploy1003 Finished scap sync-world: Backport for [[gerrit:1319489|Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] (duration: 06m 48s) [14:51:52] T433457: `mediawiki.api-request` validation errors from Extension:OAuth `context.oauth_consumer_*` logging context - https://phabricator.wikimedia.org/T433457 [14:52:35] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - urldownloader_8080: Servers urldownloader1003.wikimedia.org are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [14:54:05] yeah [14:54:07] I know [14:54:30] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage [14:54:41] !log root@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2040: Maintenance [14:55:03] !log cwilliams@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2048.codfw.wmnet with reason: Maintenance [14:55:11] !log cwilliams@cumin1003 dbctl commit (dc=all): 'Depooling es2048 (T431660)', diff saved to https://phabricator.wikimedia.org/P95816 and previous config saved to /var/cache/conftool/dbconfig/20260730-145510-cwilliams.json [14:55:25] RECOVERY - Postgres Replication Lag on puppetdb2003 is OK: POSTGRES_HOT_STANDBY_DELAY OK: DB puppetdb (host:localhost) 255368 and 0 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [14:56:08] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage [14:56:20] 10ops-codfw, 06SRE, 06DC-Ops: move wikikube-worker servers in same rack - https://phabricator.wikimedia.org/T431585#12171730 (10Jhancock.wm) [14:58:09] 10ops-codfw, 06SRE, 06DC-Ops: move wikikube-worker servers in same rack - https://phabricator.wikimedia.org/T431585#12171736 (10Jhancock.wm) @Clement_Goubert I wanted to move these servers for fixing some 10G availability in the racks. Is there a good time in the next two coming weeks to take care of this? N... [14:58:44] RESOLVED: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to esams RIPE Atlas anchor: failures over threshold for measurement 59940414 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [14:59:41] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2079.codfw.wmnet with OS trixie [14:59:49] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12171742 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-be2079.codfw.wmnet with OS trixie completed: - ms-be2079 (**PASS*... [15:00:09] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage [15:00:14] !log root@cumin1003 START - Cookbook sre.mysql.pool pool es2048: Maintenance [15:00:30] !log mvernon@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage [15:00:50] thank you tchin akhatun James_F ! [15:00:54] and tgr_ ! :) [15:02:38] (03CR) 10Aleksandar Mastilovic: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1319190 (https://phabricator.wikimedia.org/T432312) (owner: 10Aleksandar Mastilovic) [15:02:51] (03CR) 10Phuedx: [C:03+2] Test Kitchen UI: Deploy v1.5.0 release to staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319178 (https://phabricator.wikimedia.org/T432304) (owner: 10Clare Ming) [15:04:22] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage [15:04:40] !log elukey@cumin1003 START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts [15:04:56] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts [15:05:19] (03Merged) 10jenkins-bot: Test Kitchen UI: Deploy v1.5.0 release to staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319178 (https://phabricator.wikimedia.org/T432304) (owner: 10Clare Ming) [15:05:44] 10ops-codfw, 06SRE, 06DBA, 06DC-Ops: db2209 firmware upgrade after a crash - https://phabricator.wikimedia.org/T431952#12171774 (10Jhancock.wm) scheduling a zoom meeting with Dell for some reason. suggests they have no idea why it's doing this. [15:06:26] PROBLEM - Postgres Replication Lag on puppetdb2003 is CRITICAL: POSTGRES_HOT_STANDBY_DELAY CRITICAL: DB puppetdb (host:localhost) 43419960 and 1 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [15:07:26] RECOVERY - Postgres Replication Lag on puppetdb2003 is OK: POSTGRES_HOT_STANDBY_DELAY OK: DB puppetdb (host:localhost) 85520 and 0 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [15:11:32] !log root@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: Maintenance [15:11:52] !log cwilliams@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1047.eqiad.wmnet with reason: Maintenance [15:12:00] !log cwilliams@cumin1003 dbctl commit (dc=all): 'Depooling es1047 (T431660)', diff saved to https://phabricator.wikimedia.org/P95820 and previous config saved to /var/cache/conftool/dbconfig/20260730-151200-cwilliams.json [15:13:49] !log mvernon@cumin1003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie [15:13:59] 10ops-codfw, 06SRE, 10SRE-swift-storage, 06Data-Persistence, 06DC-Ops: Q4:rack/setup/install ms-be209[7,8] - https://phabricator.wikimedia.org/T424892#12171793 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-be2097.codfw.wmnet with OS trixie executed w... [15:15:43] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host ms-be2080.codfw.wmnet with OS trixie [15:15:56] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12171802 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host ms-be2080.codfw.wmnet with OS trixie [15:16:54] (03PS1) 10Bking: relforge: Allow, but disable the security plugin. [puppet] - 10https://gerrit.wikimedia.org/r/1319506 (https://phabricator.wikimedia.org/T350516) [15:17:01] !log root@cumin1003 START - Cookbook sre.mysql.pool pool es1047: Maintenance [15:17:12] (03CR) 10Bking: "check puppet" [puppet] - 10https://gerrit.wikimedia.org/r/1319506 (https://phabricator.wikimedia.org/T350516) (owner: 10Bking) [15:17:32] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie [15:17:46] 10ops-codfw, 06SRE, 10SRE-swift-storage, 06Data-Persistence, 06DC-Ops: Q4:rack/setup/install ms-be209[7,8] - https://phabricator.wikimedia.org/T424892#12171807 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-be2097.codfw.wmnet with OS trixie [15:18:24] !log mvernon@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" [15:18:33] !log mvernon@cumin1003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie [15:18:43] 10ops-codfw, 06SRE, 10SRE-swift-storage, 06Data-Persistence, 06DC-Ops: Q4:rack/setup/install ms-be209[7,8] - https://phabricator.wikimedia.org/T424892#12171811 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-be2097.codfw.wmnet with OS trixie executed w... [15:19:23] (03CR) 10Bking: [C:03+2] relforge: Allow, but disable the security plugin. [puppet] - 10https://gerrit.wikimedia.org/r/1319506 (https://phabricator.wikimedia.org/T350516) (owner: 10Bking) [15:19:34] !log mvernon@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003" [15:19:35] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2098.codfw.wmnet with OS trixie [15:19:43] 10ops-codfw, 06SRE, 10SRE-swift-storage, 06Data-Persistence, 06DC-Ops: Q4:rack/setup/install ms-be209[7,8] - https://phabricator.wikimedia.org/T424892#12171819 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-be2098.codfw.wmnet with OS trixie completed:... [15:20:18] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie [15:20:25] RESOLVED: SystemdUnitFailed: generate_vrts_aliases.service on mx-in2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [15:20:27] 10ops-codfw, 06SRE, 10SRE-swift-storage, 06Data-Persistence, 06DC-Ops: Q4:rack/setup/install ms-be209[7,8] - https://phabricator.wikimedia.org/T424892#12171821 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-be2097.codfw.wmnet with OS trixie [15:23:25] FIRING: SystemdUnitFailed: generate_vrts_aliases.service on mx-in2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [15:23:50] FIRING: [3x] HelmReleaseBadStatus: Helm release wdqs/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [15:23:56] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1081.eqiad.wmnet with OS trixie [15:24:05] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12171846 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-be1081.eqiad.wmnet with OS trixie completed: - ms-be1081 (**PASS*... [15:30:12] !log mvernon@cumin1003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS trixie [15:30:20] 10ops-codfw, 06SRE, 10SRE-swift-storage, 06Data-Persistence, 06DC-Ops: Q4:rack/setup/install ms-be209[7,8] - https://phabricator.wikimedia.org/T424892#12171864 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-be2097.codfw.wmnet with OS trixie executed w... [15:30:41] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie [15:30:49] 10ops-codfw, 06SRE, 10SRE-swift-storage, 06Data-Persistence, 06DC-Ops: Q4:rack/setup/install ms-be209[7,8] - https://phabricator.wikimedia.org/T424892#12171866 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-be2097.codfw.wmnet with OS trixie [15:32:29] (03CR) 10Bking: "It looks like Search will be able to work around this by allowlisting the security plugin and disabling it with `disable_security_plugin: " [puppet] - 10https://gerrit.wikimedia.org/r/1318792 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [15:34:29] 06SRE, 10LDAP-Access-Requests: Grant Access to NDA for jaleman-vdr-wmf - https://phabricator.wikimedia.org/T433417#12171883 (10Dzahn) >>! In T433417#12171713, @JArguello-WMF wrote: > Where can we get the NDA file to be filled out? By contacting Katie Francis of Legal (https://meta.wikimedia.org/wiki/User:KF... [15:36:18] 06SRE, 10SRE-Access-Requests, 10LDAP-Access-Requests, 13Patch-For-Review: Grant Access to wmf for Chandler Diggs - https://phabricator.wikimedia.org/T433302#12171892 (10Dzahn) @SLyngshede-WMF Should we add contractors with a -ctr wikimedia.org address and an expiry_date to the data.yaml? Or is it not done... [15:36:28] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage [15:38:08] !log bking@cumin2003 START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - T324335 [15:38:12] T324335: Index the OpenSearch server application logs in the central observability cluster, using ECS format - https://phabricator.wikimedia.org/T324335 [15:38:27] PROBLEM - Postgres Replication Lag on puppetdb2003 is CRITICAL: POSTGRES_HOT_STANDBY_DELAY CRITICAL: DB puppetdb (host:localhost) 445702800 and 15 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [15:40:26] RECOVERY - Postgres Replication Lag on puppetdb2003 is OK: POSTGRES_HOT_STANDBY_DELAY OK: DB puppetdb (host:localhost) 2588088 and 0 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [15:41:52] 10ops-eqsin, 06SRE, 06Infrastructure-Foundations, 10netops, 06Traffic: eqsin downtime scheduling for switch upgrade - https://phabricator.wikimedia.org/T433097#12171900 (10RobH) 05Open→03Resolved Resolution Notes: * Work was completed without any unexpected downtime * The fibers/optics needed ha... [15:42:15] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage [15:43:49] !log bking@deploy1003 Started deploy [wdqs/wdqs@e8fb00c]: T430880 [15:43:54] T430880: Migrate WDQS hosts to Bookworm or later - https://phabricator.wikimedia.org/T430880 [15:43:59] !log bking@deploy1003 Finished deploy [wdqs/wdqs@e8fb00c]: T430880 (duration: 00m 24s) [15:44:52] 10ops-codfw, 10ops-eqiad, 06SRE, 06DC-Ops: Eqiad Pdu tripped breaker on ps1-c3-eqiad no automated alerts - https://phabricator.wikimedia.org/T414134#12171922 (10Jhancock.wm) was checking up on this. there were two separate alerts in July but no tickets for them. [15:45:28] (03CR) 10JHathaway: [C:03+1] data.yaml: record LDAP access for cdiggs-ctr [puppet] - 10https://gerrit.wikimedia.org/r/1319385 (owner: 10Slyngshede) [15:46:26] (03PS5) 10Cathal Mooney: LVS IPIP encapsulation: support backend servers using nftables [puppet] - 10https://gerrit.wikimedia.org/r/1319486 (https://phabricator.wikimedia.org/T433601) [15:47:22] PROBLEM - SSH on crm2001 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/SSH/monitoring [15:48:16] !log root@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2048: Maintenance [15:48:48] (03CR) 10Cathal Mooney: "Updated patchset as I'd omitted to create separate MSS clamp rules for every interface. Can't imagine why that's needed (at least on lo w" [puppet] - 10https://gerrit.wikimedia.org/r/1319486 (https://phabricator.wikimedia.org/T433601) (owner: 10Cathal Mooney) [15:49:12] RECOVERY - SSH on crm2001 is OK: SSH OK - OpenSSH_9.2p1 Debian-2+deb12u10 (protocol 2.0) https://wikitech.wikimedia.org/wiki/SSH/monitoring [15:49:40] !log mvernon@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage [15:50:17] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-be1082.eqiad.wmnet with OS trixie [15:50:25] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12171952 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-be1082.eqiad.wmnet with OS trixie [15:50:26] FIRING: [2x] GanetiBGPNoInPrefixes: ganeti2033 is not sending any prefix to lsw1-b7-codfw - group Ganeti4 - https://wikitech.wikimedia.org/wiki/Ganeti#GanetiBGPNoInPrefixes - https://alerts.wikimedia.org/?q=alertname%3DGanetiBGPNoInPrefixes [15:53:17] !log cwilliams@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1040.eqiad.wmnet with reason: Maintenance [15:53:25] !log cwilliams@cumin1003 dbctl commit (dc=all): 'Depooling es1040 (T431660)', diff saved to https://phabricator.wikimedia.org/P95827 and previous config saved to /var/cache/conftool/dbconfig/20260730-155324-cwilliams.json [15:53:42] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage [15:54:19] (03CR) 10Dzahn: "duplicate of https://gerrit.wikimedia.org/r/c/operations/puppet/+/1319123" [puppet] - 10https://gerrit.wikimedia.org/r/1319385 (owner: 10Slyngshede) [15:54:48] (03CR) 10Dzahn: "could you also clarify over at https://phabricator.wikimedia.org/T433302#12171465 ?" [puppet] - 10https://gerrit.wikimedia.org/r/1319385 (owner: 10Slyngshede) [15:55:22] (03CR) 10Dzahn: "duplicate of https://gerrit.wikimedia.org/r/c/operations/puppet/+/1319385 - feel free to either merge or abandon either of them :)" [puppet] - 10https://gerrit.wikimedia.org/r/1319123 (https://phabricator.wikimedia.org/T433302) (owner: 10Dzahn) [15:58:48] !log root@cumin1003 START - Cookbook sre.mysql.pool pool es1040: Maintenance [16:00:48] (03PS1) 10Kevin Bazira: ml-services: update tts-section-generator image to strip phonos ⓘ, fix year-slash reading (2026.07.31) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319513 (https://phabricator.wikimedia.org/T433594) [16:03:42] !log bking@cumin2003 END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - T324335 [16:03:48] T324335: Index the OpenSearch server application logs in the central observability cluster, using ECS format - https://phabricator.wikimedia.org/T324335 [16:04:09] (03CR) 10Kevin Bazira: [C:03+2] ml-services: update tts-section-generator image to strip phonos ⓘ, fix year-slash reading (2026.07.31) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319513 (https://phabricator.wikimedia.org/T433594) (owner: 10Kevin Bazira) [16:04:23] !log root@cumin1003 END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1040: Maintenance [16:04:28] !log root@cumin1003 START - Cookbook sre.mysql.pool pool es1040: Maintenance [16:04:45] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2080.codfw.wmnet with OS trixie [16:04:57] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12172009 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-be2080.codfw.wmnet with OS trixie completed: - ms-be2080 (**PASS*... [16:05:04] !log root@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1047: Maintenance [16:06:25] (03Merged) 10jenkins-bot: ml-services: update tts-section-generator image to strip phonos ⓘ, fix year-slash reading (2026.07.31) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319513 (https://phabricator.wikimedia.org/T433594) (owner: 10Kevin Bazira) [16:08:17] !log kevinbazira@deploy1003 helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . [16:08:40] !log mvernon@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage [16:13:10] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage [16:13:27] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2097.codfw.wmnet with OS trixie [16:13:40] 10ops-codfw, 06SRE, 10SRE-swift-storage, 06Data-Persistence, 06DC-Ops: Q4:rack/setup/install ms-be209[7,8] - https://phabricator.wikimedia.org/T424892#12172018 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-be2097.codfw.wmnet with OS trixie completed:... [16:20:52] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in profile terraform [puppet] - 10https://gerrit.wikimedia.org/r/1315011 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [16:25:02] (03PS2) 10Dzahn: varnish: new policy to allow websockets and caching, apply to phab [puppet] - 10https://gerrit.wikimedia.org/r/1171263 (https://phabricator.wikimedia.org/T274228) [16:25:42] (03CR) 10Dzahn: [C:03+1] "@Slyngshede are tickets required for this or can we just merge it? lgtm" [puppet] - 10https://gerrit.wikimedia.org/r/1318214 (owner: 10Ammarpad) [16:26:55] (03CR) 10Dzahn: [V:03+1] varnish: new policy to allow websockets and caching, apply to phab (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1171263 (https://phabricator.wikimedia.org/T274228) (owner: 10Dzahn) [16:27:21] (03PS7) 10Dzahn: site: add zuul1004/2004 [puppet] - 10https://gerrit.wikimedia.org/r/1310656 (https://phabricator.wikimedia.org/T427353) [16:27:21] (03CR) 10Dzahn: [C:04-1] "no codfw (at least so far)" [puppet] - 10https://gerrit.wikimedia.org/r/1310656 (https://phabricator.wikimedia.org/T427353) (owner: 10Dzahn) [16:27:27] 10ops-eqiad, 06SRE, 06DBA, 06DC-Ops: db1218 crashed - https://phabricator.wikimedia.org/T433565#12172065 (10VRiley-WMF) BIOS firmware installing now [16:30:53] 10ops-codfw, 06SRE, 10SRE-swift-storage, 06Data-Persistence, 06DC-Ops: Q4:rack/setup/install ms-be209[7,8] - https://phabricator.wikimedia.org/T424892#12172074 (10MatthewVernon) @Papaul I've fixed it. In both cases, this is because these systems have arrived with non-blank spinning disks. ms-be2097 had... [16:31:58] (03PS8) 10Dzahn: site: add zuul1004 [puppet] - 10https://gerrit.wikimedia.org/r/1310656 (https://phabricator.wikimedia.org/T427353) [16:32:31] (03CR) 10Majavah: [V:03+1 C:03+1] Puppet 8: Replace legacy facts in role idp_clouddev [puppet] - 10https://gerrit.wikimedia.org/r/1315013 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [16:32:34] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host ms-be2081.codfw.wmnet with OS trixie [16:32:44] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12172087 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host ms-be2081.codfw.wmnet with OS trixie [16:33:21] (03PS6) 10Cathal Mooney: LVS IPIP encapsulation: support backend servers using nftables [puppet] - 10https://gerrit.wikimedia.org/r/1319486 (https://phabricator.wikimedia.org/T433601) [16:34:37] (03CR) 10Majavah: [C:03+1] Puppet 8: Replace legacy facts in misc hiera [puppet] - 10https://gerrit.wikimedia.org/r/1315019 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [16:36:00] 10ops-codfw, 06SRE, 06DC-Ops, 06ServiceOps new: Rerack mc-gp2006 and mc2046 - https://phabricator.wikimedia.org/T432716#12172126 (10Jhancock.wm) @jijiki i can get this tomorrow some time. I will ping when i can start it. [16:36:11] (03PS7) 10Cathal Mooney: LVS IPIP encapsulation: support backend servers using nftables [puppet] - 10https://gerrit.wikimedia.org/r/1319486 (https://phabricator.wikimedia.org/T433601) [16:36:12] (03CR) 10Majavah: [V:03+1] "PCC SUCCESS (NOOP 6): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9094/console" [puppet] - 10https://gerrit.wikimedia.org/r/1315012 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [16:36:29] (03CR) 10Majavah: [V:03+1 C:03+1] Puppet 8: Replace legacy facts in role wmcs [puppet] - 10https://gerrit.wikimedia.org/r/1315012 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [16:36:59] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1082.eqiad.wmnet with OS trixie [16:37:06] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12172135 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-be1082.eqiad.wmnet with OS trixie completed: - ms-be1082 (**PASS*... [16:37:17] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in misc hiera [puppet] - 10https://gerrit.wikimedia.org/r/1315019 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [16:37:27] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in role idp_clouddev [puppet] - 10https://gerrit.wikimedia.org/r/1315013 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [16:37:43] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in role wmcs [puppet] - 10https://gerrit.wikimedia.org/r/1315012 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [16:37:53] 10ops-eqiad, 06SRE, 06DBA, 06DC-Ops: db1218 crashed - https://phabricator.wikimedia.org/T433565#12172144 (10VRiley-WMF) BIOS is now at 1.21 proceeding with iDRAC [16:38:19] 10ops-eqiad, 06SRE, 06DBA, 06DC-Ops: db1218 crashed - https://phabricator.wikimedia.org/T433565#12172149 (10VRiley-WMF) 05Open→03In progress [16:41:25] 10ops-codfw, 06SRE, 10SRE-swift-storage, 06Data-Persistence, 06DC-Ops: Q4:rack/setup/install ms-be209[7,8] - https://phabricator.wikimedia.org/T424892#12172160 (10Papaul) @MatthewVernon thank you for the update and fix. Yes i recalled you brought this up once in the pass . I will follow up with @wiki_wil... [16:43:25] FIRING: [2x] SystemdUnitFailed: generate_vrts_aliases.service on mx-in2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [16:44:02] (03CR) 10Cathal Mooney: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1319486 (https://phabricator.wikimedia.org/T433601) (owner: 10Cathal Mooney) [16:45:36] (03PS1) 10Dzahn: zuul: replace puppet legacy facts [puppet] - 10https://gerrit.wikimedia.org/r/1319519 (https://phabricator.wikimedia.org/T372666) [16:46:56] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage [16:47:45] (03PS2) 10Dzahn: zuul: replace puppet legacy facts [puppet] - 10https://gerrit.wikimedia.org/r/1319519 (https://phabricator.wikimedia.org/T372666) [16:48:25] FIRING: [2x] SystemdUnitFailed: generate_vrts_aliases.service on mx-in2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [16:48:53] (03PS1) 10OSleger: Bump Parsoid image limit to 5000 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319521 (https://phabricator.wikimedia.org/T430854) [16:49:15] (03CR) 10Dzahn: [C:03+2] site: add zuul1004 [puppet] - 10https://gerrit.wikimedia.org/r/1310656 (https://phabricator.wikimedia.org/T427353) (owner: 10Dzahn) [16:50:15] !log root@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1040: Maintenance [16:50:45] !log cwilliams@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1048.eqiad.wmnet with reason: Maintenance [16:50:53] !log cwilliams@cumin1003 dbctl commit (dc=all): 'Depooling es1048 (T431660)', diff saved to https://phabricator.wikimedia.org/P95833 and previous config saved to /var/cache/conftool/dbconfig/20260730-165053-cwilliams.json [16:52:35] (03CR) 10Dzahn: [C:03+2] "Unable to find zookeeper ID for zuul1004.eqiad.wmnet in {zuul1001.eqiad.wmnet => 1001}" [puppet] - 10https://gerrit.wikimedia.org/r/1310656 (https://phabricator.wikimedia.org/T427353) (owner: 10Dzahn) [16:52:43] 10ops-eqiad, 06SRE, 06collaboration-services, 06DC-Ops, 13Patch-For-Review: Repurpose ganeti102[3456] for Zuul migration - https://phabricator.wikimedia.org/T427353#12172198 (10VRiley-WMF) Hey @Dzahn I may need some help with zuul1006. I believe all these servers will be using a 1 gig cable, but this uni... [16:54:11] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage [16:54:36] (03PS1) 10Bking: cloudelastic: Allow, but disable the security plugin. [puppet] - 10https://gerrit.wikimedia.org/r/1319522 (https://phabricator.wikimedia.org/T350516) [16:54:46] 10ops-eqiad, 06SRE, 06collaboration-services, 06DC-Ops, 13Patch-For-Review: Repurpose ganeti102[3456] for Zuul migration - https://phabricator.wikimedia.org/T427353#12172202 (10Dzahn) @VRiley-WMF Ok, thanks for the update. Take your time. I will be busy setting up zuul1004 at first regardless. It's ok if... [16:55:47] !log root@cumin1003 START - Cookbook sre.mysql.pool pool es1048: Maintenance [16:58:43] 06SRE, 06ServiceOps new, 10VisualEditor, 10VisualEditor Suggestion Mode, and 3 others: Deploy Headless VE in k8s for technical pilot - https://phabricator.wikimedia.org/T431497#12172209 (10GGoncalves-WMF) We do need to have the conversations about harmonizing Prep Pantry and SCROLL in the general case. @ML... [17:00:54] (03PS1) 10Bking: cirrussearch CODFW: Allow, but disable the security plugin. [puppet] - 10https://gerrit.wikimedia.org/r/1319524 (https://phabricator.wikimedia.org/T350516) [17:01:07] 10ops-codfw, 06SRE, 10SRE-swift-storage, 06Data-Persistence, 06DC-Ops: Q4:rack/setup/install ms-be209[7,8] - https://phabricator.wikimedia.org/T424892#12172234 (10wiki_willy) Hey @MatthewVernon & @Papaul - I mentioned this in the previous email "Installation issues with new config-J nodes", but Dell ende... [17:01:16] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1319524 (https://phabricator.wikimedia.org/T350516) (owner: 10Bking) [17:01:20] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1319522 (https://phabricator.wikimedia.org/T350516) (owner: 10Bking) [17:03:14] (03PS1) 10Bking: cirrussearch EQIAD: Allow, but disable the security plugin. [puppet] - 10https://gerrit.wikimedia.org/r/1319525 (https://phabricator.wikimedia.org/T350516) [17:03:24] 10ops-eqiad, 06SRE, 06DBA, 06DC-Ops: db1218 crashed - https://phabricator.wikimedia.org/T433565#12172238 (10VRiley-WMF) 05In progress→03Open iDRAC is at 7.30 now If possible, please take a look and see if that helps. @Marostegui thank you! [17:05:30] (03CR) 10Bking: [C:03+2] cloudelastic: Allow, but disable the security plugin. [puppet] - 10https://gerrit.wikimedia.org/r/1319522 (https://phabricator.wikimedia.org/T350516) (owner: 10Bking) [17:10:21] (03PS1) 10Dzahn: zuul: add zuul1004 to zookeeper and list of main hosts [puppet] - 10https://gerrit.wikimedia.org/r/1319528 (https://phabricator.wikimedia.org/T427353) [17:10:53] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-be1083.eqiad.wmnet with OS trixie [17:11:01] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12172264 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-be1083.eqiad.wmnet with OS trixie [17:11:04] !log bking@cumin2003 START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - T350516 [17:11:08] T350516: Enable OpenSearch security plugin - Beta Logs - https://phabricator.wikimedia.org/T350516 [17:11:49] (03PS1) 10Kevin Bazira: ml-services: update tts-section-generator image to add .31 corpus rescan (seed 43, 100-article sample) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319529 (https://phabricator.wikimedia.org/T433594) [17:12:32] (03CR) 10Dzahn: "@dduvall @hashar FYI - we are now getting the first -physical- machine for new zuul" [puppet] - 10https://gerrit.wikimedia.org/r/1319528 (https://phabricator.wikimedia.org/T427353) (owner: 10Dzahn) [17:13:50] FIRING: [3x] HelmReleaseBadStatus: Helm release wdqs/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [17:14:01] (03CR) 10Dzahn: [C:03+2] zuul: add zuul1004 to zookeeper and list of main hosts [puppet] - 10https://gerrit.wikimedia.org/r/1319528 (https://phabricator.wikimedia.org/T427353) (owner: 10Dzahn) [17:15:25] (03CR) 10Kevin Bazira: [C:03+2] ml-services: update tts-section-generator image to add .31 corpus rescan (seed 43, 100-article sample) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319529 (https://phabricator.wikimedia.org/T433594) (owner: 10Kevin Bazira) [17:16:55] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2081.codfw.wmnet with OS trixie [17:17:08] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12172301 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-be2081.codfw.wmnet with OS trixie completed: - ms-be2081 (**WARN*... [17:18:03] (03Merged) 10jenkins-bot: ml-services: update tts-section-generator image to add .31 corpus rescan (seed 43, 100-article sample) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319529 (https://phabricator.wikimedia.org/T433594) (owner: 10Kevin Bazira) [17:20:07] !log kevinbazira@deploy1003 helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . [17:23:00] (03CR) 10Subramanya Sastry: Bump Parsoid image limit to 5000 (031 comment) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319521 (https://phabricator.wikimedia.org/T430854) (owner: 10OSleger) [17:23:11] !log sukhe@lvs2014:~$ sudo systemctl restart pybal.service [17:23:13] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [17:24:15] !log mvernon@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage [17:26:05] !log bking@apt1002 `reprepro --noskipold --component thirdparty/opensearch3 update trixie-wikimedia` T433624 [17:26:09] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [17:26:09] T433624: Upload opensearch 3.7.0 to apt.wikimedia.org (trixie/thirdparty) - https://phabricator.wikimedia.org/T433624 [17:26:27] RECOVERY - PyBal IPVS diff check on lvs2014 is OK: OK: no difference between hosts in IPVS/PyBal https://wikitech.wikimedia.org/wiki/PyBal [17:29:00] RECOVERY - PyBal backends health check on lvs2014 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [17:29:20] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage [17:29:46] PROBLEM - Check unit status of push_cross_cluster_settings_9600 on cloudelastic1009 is CRITICAL: CRITICAL: Status of the systemd unit push_cross_cluster_settings_9600 https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [17:29:48] PROBLEM - Check unit status of push_cross_cluster_settings_9200 on cloudelastic1009 is CRITICAL: CRITICAL: Status of the systemd unit push_cross_cluster_settings_9200 https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [17:30:21] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12172333 (10MatthewVernon) [17:30:51] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12172335 (10MatthewVernon) [counting 2 extra codfw backends which are already trixie, after T424892] [17:32:50] PROBLEM - PyBal backends health check on lvs2014 is CRITICAL: PYBAL CRITICAL - CRITICAL - urldownloader_8080: Servers urldownloader2003.wikimedia.org are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [17:33:16] yeah, it's known unfortunately but I have no good reasons as to why [17:33:25] so going through that now [17:34:07] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12172371 (10MatthewVernon) [17:34:25] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-07-03 - 2026-07-31): Degraded RAID on an-presto1013 - https://phabricator.wikimedia.org/T433030#12172373 (10RobH) [17:34:34] (03CR) 10Bking: [C:03+2] Use a separate severity value to route to data engineering Slack [puppet] - 10https://gerrit.wikimedia.org/r/1319190 (https://phabricator.wikimedia.org/T432312) (owner: 10Aleksandar Mastilovic) [17:34:42] PROBLEM - Check unit status of push_cross_cluster_settings_9400 on cloudelastic1009 is CRITICAL: CRITICAL: Status of the systemd unit push_cross_cluster_settings_9400 https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [17:35:25] 10ops-eqiad, 06SRE, 06DBA, 06DC-Ops: db1218 crashed - https://phabricator.wikimedia.org/T433565#12172375 (10CWilliams-WMF) Thanks for the updates @VRiley-WMF! I will get the host replicating again and see if it stays up. [17:36:47] !log bking@cumin2003 END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - T350516 [17:36:51] T350516: Enable OpenSearch security plugin - Beta Logs - https://phabricator.wikimedia.org/T350516 [17:37:15] 06SRE, 10LDAP-Access-Requests: Grant Access to ciadmin for Vaughn Walters - https://phabricator.wikimedia.org/T433615#12172386 (10Dzahn) Welcome @vaughnwalters Can you please create a new SSH key pair for production access, that is not already used anywhere, and then paste the public part here on the ticket?... [17:37:59] RESOLVED: HelmReleaseBadStatus: Helm release wdqs/scholarly-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [17:38:00] RECOVERY - PyBal backends health check on lvs2014 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [17:38:50] FIRING: HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-rollback - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs-next - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [17:40:09] (03CR) 10Bking: [C:04-1] "You'll need to write unit tests for the alerts by creating team-data-engineering/airflow-k8s_test.yaml and verify them with `promtool tes" [alerts] - 10https://gerrit.wikimedia.org/r/1319189 (https://phabricator.wikimedia.org/T432312) (owner: 10Aleksandar Mastilovic) [17:40:50] PROBLEM - PyBal backends health check on lvs2014 is CRITICAL: PYBAL CRITICAL - CRITICAL - urldownloader_8080: Servers urldownloader2004.wikimedia.org are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [17:41:30] (03CR) 10Aleksandar Mastilovic: "I'm pretty sure this is already covered by unit tests because they failed my first commit (I had to add the `# page` in the summary to make" [alerts] - 10https://gerrit.wikimedia.org/r/1319189 (https://phabricator.wikimedia.org/T432312) (owner: 10Aleksandar Mastilovic) [17:42:59] FIRING: [2x] HelmReleaseBadStatus: Helm release wdqs/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [17:43:50] !log root@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1048: Maintenance [17:44:42] RECOVERY - Check unit status of push_cross_cluster_settings_9400 on cloudelastic1009 is OK: OK: Status of the systemd unit push_cross_cluster_settings_9400 https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [17:49:46] RECOVERY - Check unit status of push_cross_cluster_settings_9600 on cloudelastic1009 is OK: OK: Status of the systemd unit push_cross_cluster_settings_9600 https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [17:49:48] RECOVERY - Check unit status of push_cross_cluster_settings_9200 on cloudelastic1009 is OK: OK: Status of the systemd unit push_cross_cluster_settings_9200 https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [17:50:36] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [17:53:04] FIRING: NELByCountryHigh: Elevated Network Error Logging events (tcp.timed_out from RU) - https://wikitech.wikimedia.org/wiki/Network_monitoring#NEL_alerts - https://logstash.wikimedia.org/goto/5c8f4ca1413eda33128e5c5a35da7e28 - https://alerts.wikimedia.org/?q=alertname%3DNELByCountryHigh [17:54:56] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1083.eqiad.wmnet with OS trixie [17:55:02] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12172426 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-be1083.eqiad.wmnet with OS trixie completed: - ms-be1083 (**WARN*... [17:59:10] (03PS2) 10OSleger: Bump Parsoid image limit to 5000 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319521 (https://phabricator.wikimedia.org/T430854) [17:59:54] (03PS3) 10OSleger: Increase Parsoid image limit to 5000 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319521 (https://phabricator.wikimedia.org/T430854) [17:59:57] (03CR) 10Bking: [C:04-1] "Yes, I expect CI to always run `promtool`, but it doesn't seem to be doing it here. If you look any other alerts file, they always come wi" [alerts] - 10https://gerrit.wikimedia.org/r/1319189 (https://phabricator.wikimedia.org/T432312) (owner: 10Aleksandar Mastilovic) [18:00:05] dduvall and dancy: #bothumor Q:How do functions break up? A:They stop calling each other. Rise for MediaWiki train - Utc-7 Version deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260730T1800). [18:00:17] (03CR) 10OSleger: Increase Parsoid image limit to 5000 (031 comment) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319521 (https://phabricator.wikimedia.org/T430854) (owner: 10OSleger) [18:04:32] (03PS1) 10Ssingh: service/urldownloader: switch to TCP check (IdleConnection) [puppet] - 10https://gerrit.wikimedia.org/r/1319533 (https://phabricator.wikimedia.org/T429175) [18:04:56] fyi, we're rolling to group1 and all wikis today if possible [18:07:12] (03PS1) 10TrainBranchBot: group1 to 1.47.0-wmf.13 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319534 (https://phabricator.wikimedia.org/T430832) [18:07:15] (03CR) 10TrainBranchBot: [C:03+2] "Initiated by dduvall@deploy1003" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319534 (https://phabricator.wikimedia.org/T430832) (owner: 10TrainBranchBot) [18:07:59] FIRING: [2x] SystemdUnitFailed: generate_vrts_aliases.service on mx-in2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [18:09:13] (03Merged) 10jenkins-bot: group1 to 1.47.0-wmf.13 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319534 (https://phabricator.wikimedia.org/T430832) (owner: 10TrainBranchBot) [18:09:51] (03PS6) 10Aaron Schulz: Remove some API Gateway specific template files [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319114 (https://phabricator.wikimedia.org/T428625) [18:12:59] FIRING: [3x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-3/1/2 (Transport: Hurricane Electric (dc4841.dal5)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [18:15:20] !log dduvall@deploy1003 rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs T430832 [18:15:25] T430832: 1.47.0-wmf.13 deployment blockers - https://phabricator.wikimedia.org/T430832 [18:17:06] (03PS2) 10Ssingh: service/urldownloader: switch to TCP check (IdleConnection) [puppet] - 10https://gerrit.wikimedia.org/r/1319533 (https://phabricator.wikimedia.org/T429175) [18:17:59] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [18:18:12] (03PS1) 10TrainBranchBot: group2 to 1.47.0-wmf.13 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319535 (https://phabricator.wikimedia.org/T430832) [18:18:15] (03CR) 10TrainBranchBot: [C:03+2] "Initiated by dduvall@deploy1003" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319535 (https://phabricator.wikimedia.org/T430832) (owner: 10TrainBranchBot) [18:18:39] 10ops-eqiad, 06SRE, 06DBA, 06DC-Ops: db1218 crashed - https://phabricator.wikimedia.org/T433565#12172459 (10CWilliams-WMF) The host has caught up on replication: `sql MariaDB [(none)]> select @@hostname, timestampdiff(SECOND, ts, NOW()) from heartbeat.heartbeat; +------------+------------------------------... [18:19:17] (03Merged) 10jenkins-bot: group2 to 1.47.0-wmf.13 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319535 (https://phabricator.wikimedia.org/T430832) (owner: 10TrainBranchBot) [18:25:16] !log dduvall@deploy1003 rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.13 refs T430832 [18:25:21] T430832: 1.47.0-wmf.13 deployment blockers - https://phabricator.wikimedia.org/T430832 [18:35:14] group2 looks alright. train {{done}} for now [18:36:36] (03CR) 10Fabfur: [C:03+1] "LGTM!" [puppet] - 10https://gerrit.wikimedia.org/r/1319533 (https://phabricator.wikimedia.org/T429175) (owner: 10Ssingh) [18:38:51] (03PS8) 10Cathal Mooney: LVS IPIP encapsulation: support backend servers using nftables [puppet] - 10https://gerrit.wikimedia.org/r/1319486 (https://phabricator.wikimedia.org/T433601) [18:39:20] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-be1084.eqiad.wmnet with OS trixie [18:39:29] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12172504 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-be1084.eqiad.wmnet with OS trixie [18:40:07] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12172505 (10MatthewVernon) [18:40:51] !log cjming@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply [18:41:20] (03CR) 10CI reject: [V:04-1] LVS IPIP encapsulation: support backend servers using nftables [puppet] - 10https://gerrit.wikimedia.org/r/1319486 (https://phabricator.wikimedia.org/T433601) (owner: 10Cathal Mooney) [18:41:29] !log cjming@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply [18:47:03] (03CR) 10BCornwall: [C:03+1] service/urldownloader: switch to TCP check (IdleConnection) [puppet] - 10https://gerrit.wikimedia.org/r/1319533 (https://phabricator.wikimedia.org/T429175) (owner: 10Ssingh) [18:51:12] (03PS9) 10Cathal Mooney: LVS IPIP encapsulation: support backend servers using nftables [puppet] - 10https://gerrit.wikimedia.org/r/1319486 (https://phabricator.wikimedia.org/T433601) [18:52:53] !log mvernon@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage [18:54:51] (03CR) 10Cathal Mooney: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1319486 (https://phabricator.wikimedia.org/T433601) (owner: 10Cathal Mooney) [18:56:52] PROBLEM - Check if Pybal has been restarted after pybal.conf was changed on lvs1020 is CRITICAL: CRITICAL: Service pybal.service has not been restarted after /etc/pybal/pybal.conf was changed (gt 1h). https://wikitech.wikimedia.org/wiki/PyBal%23Pybal_service_has_not_been_restarted [18:57:43] yeah I will fix that in a bit [18:58:49] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage [19:06:29] (03PS10) 10Cathal Mooney: LVS IPIP encapsulation: support backend servers using nftables [puppet] - 10https://gerrit.wikimedia.org/r/1319486 (https://phabricator.wikimedia.org/T433601) [19:09:04] (03PS11) 10Cathal Mooney: LVS IPIP encapsulation: support backend servers using nftables [puppet] - 10https://gerrit.wikimedia.org/r/1319486 (https://phabricator.wikimedia.org/T433601) [19:19:39] (03CR) 10Arlolra: [C:03+1] Increase Parsoid image limit to 5000 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319521 (https://phabricator.wikimedia.org/T430854) (owner: 10OSleger) [19:19:46] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Thursday, July 30 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-ite" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319521 (https://phabricator.wikimedia.org/T430854) (owner: 10OSleger) [19:19:55] (03CR) 10Ssingh: [C:03+2] service/urldownloader: switch to TCP check (IdleConnection) [puppet] - 10https://gerrit.wikimedia.org/r/1319533 (https://phabricator.wikimedia.org/T429175) (owner: 10Ssingh) [19:20:53] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1084.eqiad.wmnet with OS trixie [19:21:02] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12172666 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-be1084.eqiad.wmnet with OS trixie completed: - ms-be1084 (**PASS*... [19:22:43] (03CR) 10Lerickson: [C:03+1] "Sorry for the delay on this! Thanks for adding back the affinity, I think it makes sense (at least for now)." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1318733 (https://phabricator.wikimedia.org/T431395) (owner: 10Trueg) [19:23:55] !log sukhe@lvs1019:~$ sudo systemctl restart pybal.service [19:23:57] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [19:24:04] !log sukhe@lvs1020:~$ sudo systemctl restart pybal.service [19:24:04] !log mvernon@cumin2003 START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART [19:24:06] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [19:24:17] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [19:24:17] !log mvernon@cumin2003 END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART [19:24:23] 06SRE, 06ServiceOps new, 10VisualEditor, 10VisualEditor Suggestion Mode, and 3 others: Deploy Headless VE in k8s for technical pilot - https://phabricator.wikimedia.org/T431497#12172668 (10Ottomata) > If that sounds right, this ticket can be closed as there's no need to engage ServiceOps via a ticket to se... [19:24:59] !log sukhe@lvs2014:~$ sudo systemctl restart pybal.service [19:25:00] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [19:25:40] !log sukhe@lvs2013:~$ sudo systemctl restart pybal.service [19:25:42] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [19:25:51] RECOVERY - PyBal backends health check on lvs2014 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [19:25:56] cool [19:26:11] RECOVERY - PyBal IPVS diff check on lvs1019 is OK: OK: no difference between hosts in IPVS/PyBal https://wikitech.wikimedia.org/wiki/PyBal [19:26:37] RECOVERY - PyBal backends health check on lvs2013 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [19:26:39] RECOVERY - Check if Pybal has been restarted after pybal.conf was changed on lvs1020 is OK: OK: pybal.service was restarted after /etc/pybal/pybal.conf was changed. https://wikitech.wikimedia.org/wiki/PyBal%23Pybal_service_has_not_been_restarted [19:26:43] RECOVERY - PyBal IPVS diff check on lvs2013 is OK: OK: no difference between hosts in IPVS/PyBal https://wikitech.wikimedia.org/wiki/PyBal [19:29:35] !log mvernon@cumin2003 START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART [19:29:47] !log mvernon@cumin2003 END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART [19:32:09] 06SRE, 10SRE-swift-storage, 06Infrastructure-Foundations: Unable to reimage or reprovision ms-be2082 due to redfish connection errors - https://phabricator.wikimedia.org/T433635 (10MatthewVernon) 03NEW [19:32:18] 06SRE, 10SRE-swift-storage, 06Infrastructure-Foundations: Unable to reimage or reprovision ms-be2082 due to redfish connection errors - https://phabricator.wikimedia.org/T433635#12172706 (10MatthewVernon) p:05Triage→03High [19:32:57] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1319525 (https://phabricator.wikimedia.org/T350516) (owner: 10Bking) [19:33:53] (03PS2) 10Kamila Součková: php: rebuild to include Timo's APCU metrics fix [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1319453 (https://phabricator.wikimedia.org/T433312) [19:35:06] (03CR) 10Kamila Součková: php: rebuild to include Timo's APCU metrics fix (031 comment) [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1319453 (https://phabricator.wikimedia.org/T433312) (owner: 10Kamila Součková) [19:35:51] 06SRE, 06Infrastructure-Foundations, 06ServiceOps new, 06Traffic, 13Patch-For-Review: Scaling urldownloaders by adding redundancy and load balancing - https://phabricator.wikimedia.org/T429175#12172713 (10ssingh) `urldownloader[12]00[34].wikimedia.org` are now behind LVS as a low-traffic IPIP service in... [19:52:59] FIRING: [2x] GanetiBGPNoInPrefixes: ganeti2033 is not sending any prefix to lsw1-b7-codfw - group Ganeti4 - https://wikitech.wikimedia.org/wiki/Ganeti#GanetiBGPNoInPrefixes - https://alerts.wikimedia.org/?q=alertname%3DGanetiBGPNoInPrefixes [19:59:48] 06SRE, 10SRE-Access-Requests: Requesting access to Jenkins server for PWangai-WMF - https://phabricator.wikimedia.org/T433563#12172752 (10Dzahn) Hi there! This is not a super common type access request, so let me start off with some general comments. Jenkins just recently moved to this new dedicated server ro... [20:00:05] RoanKattouw, urbanecm, TheresNoTime, kindrobot, and cjming: May I have your attention please! UTC late backport window. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260730T2000) [20:00:05] arlolra: A patch you scheduled for UTC late backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [20:00:35] 06SRE, 10SRE-Access-Requests, 06Release-Engineering-Team: Requesting access to Jenkins server for PWangai-WMF - https://phabricator.wikimedia.org/T433563#12172753 (10Dzahn) [20:01:20] (03CR) 10TrainBranchBot: [C:03+2] "Approved by arlolra@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319521 (https://phabricator.wikimedia.org/T430854) (owner: 10OSleger) [20:02:28] 06SRE, 06Infrastructure-Foundations, 06ServiceOps new, 06Traffic, 13Patch-For-Review: Scaling urldownloaders by adding redundancy and load balancing - https://phabricator.wikimedia.org/T429175#12172754 (10Krinkle) [20:03:40] (03Merged) 10jenkins-bot: Increase Parsoid image limit to 5000 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319521 (https://phabricator.wikimedia.org/T430854) (owner: 10OSleger) [20:03:55] !log arlolra@deploy1003 Started scap sync-world: Backport for [[gerrit:1319521|Increase Parsoid image limit to 5000 (T430854)]] [20:03:59] T430854: Improve handling of image limits by fixing the UX on pages where these limits are hit - https://phabricator.wikimedia.org/T430854 [20:05:44] !log arlolra@deploy1003 osleger, arlolra: Backport for [[gerrit:1319521|Increase Parsoid image limit to 5000 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [20:07:02] !log arlolra@deploy1003 osleger, arlolra: Continuing with deployment [20:08:13] 10ops-codfw, 06SRE, 10SRE-swift-storage, 06Data-Persistence, 06DC-Ops: Q4:rack/setup/install ms-be209[7,8] - https://phabricator.wikimedia.org/T424892#12172761 (10Papaul) @wiki_willy @MatthewVernon thanks both for the update I will look and see what we can do in our end when we received those servers aga... [20:13:15] !log arlolra@deploy1003 Finished scap sync-world: Backport for [[gerrit:1319521|Increase Parsoid image limit to 5000 (T430854)]] (duration: 09m 20s) [20:13:20] T430854: Improve handling of image limits by fixing the UX on pages where these limits are hit - https://phabricator.wikimedia.org/T430854 [20:16:51] 06SRE, 10SRE-Access-Requests, 06Release-Engineering-Team: Requesting access to Jenkins server for PWangai-WMF - https://phabricator.wikimedia.org/T433563#12172791 (10Dzahn) Hi releng team, so the actual request here is "collect castor metrics". Given our recent split of the CI server roles let's confirm it... [20:22:42] 06SRE, 10LDAP-Access-Requests: Grant Access to ciadmin for Vaughn Walters - https://phabricator.wikimedia.org/T433615#12172807 (10Dzahn) @dduvall Could you formally approve here? [20:26:44] 06SRE, 10LDAP-Access-Requests: Grant Access to ciadmin for Vaughn Walters - https://phabricator.wikimedia.org/T433615#12172820 (10Dzahn) Other than that I can just do it. Confirmed Vaughn is already in admin.yaml as ldap_only_admin (and also in wmf LDAP group and WMF/NDA group in Phabicator). For the step of... [20:28:02] (03CR) 10Papaul: [C:03+2] Add new eqsin ASN and mr1-ge-0/0/3 to prod zone [homer/public] - 10https://gerrit.wikimedia.org/r/1319205 (https://phabricator.wikimedia.org/T418439) (owner: 10Papaul) [20:28:47] (03Abandoned) 10Dzahn: ci: add parameter to toggle jenkins package install, remove it from role::ci [puppet] - 10https://gerrit.wikimedia.org/r/1308249 (https://phabricator.wikimedia.org/T418521) (owner: 10Dzahn) [20:32:54] (03CR) 10Dzahn: [C:03+1] "Simon, any thoughts on this one?" [puppet] - 10https://gerrit.wikimedia.org/r/1296495 (https://phabricator.wikimedia.org/T420184) (owner: 10Arnaudb) [20:33:56] (03PS2) 10Dzahn: data.yaml: record LDAP access for cdiggs-ctr [puppet] - 10https://gerrit.wikimedia.org/r/1319385 (https://phabricator.wikimedia.org/T433302) (owner: 10Slyngshede) [20:34:33] (03CR) 10Dzahn: [C:03+2] data.yaml: record LDAP access for cdiggs-ctr [puppet] - 10https://gerrit.wikimedia.org/r/1319385 (https://phabricator.wikimedia.org/T433302) (owner: 10Slyngshede) [20:34:35] (03CR) 10Dzahn: [V:03+2 C:03+2] data.yaml: record LDAP access for cdiggs-ctr [puppet] - 10https://gerrit.wikimedia.org/r/1319385 (https://phabricator.wikimedia.org/T433302) (owner: 10Slyngshede) [20:34:51] (03PS3) 10Dzahn: data.yaml: record LDAP access for cdiggs-ctr [puppet] - 10https://gerrit.wikimedia.org/r/1319385 (https://phabricator.wikimedia.org/T433302) (owner: 10Slyngshede) [20:35:09] (03Abandoned) 10Dzahn: admin: document ldap_only access for Chandler Diggs [puppet] - 10https://gerrit.wikimedia.org/r/1319123 (https://phabricator.wikimedia.org/T433302) (owner: 10Dzahn) [20:35:24] (03CR) 10Dzahn: [V:03+1 C:03+2] data.yaml: record LDAP access for cdiggs-ctr [puppet] - 10https://gerrit.wikimedia.org/r/1319385 (https://phabricator.wikimedia.org/T433302) (owner: 10Slyngshede) [20:39:17] 06SRE, 10SRE-Access-Requests, 10LDAP-Access-Requests, 13Patch-For-Review: Grant Access to wmf for Chandler Diggs - https://phabricator.wikimedia.org/T433302#12172888 (10Dzahn) @jijiki Simon uploaded the same change to document their access in data.yaml. So we can consider that a yes. So I abandoned my chan... [20:41:08] 06SRE, 10SRE-Access-Requests, 10LDAP-Access-Requests, 13Patch-For-Review: Grant Access to wmf for Chandler Diggs - https://phabricator.wikimedia.org/T433302#12172889 (10Dzahn) @CDiggs-WMF The best way to close this ticket would be if you can tell us your access to Matomo works. cheers! [20:42:24] 06SRE, 10SRE-Access-Requests, 10LDAP-Access-Requests, 13Patch-For-Review: Grant Access to wmf for Chandler Diggs - https://phabricator.wikimedia.org/T433302#12172892 (10Dzahn) a:03CDiggs-WMF [20:42:38] 06SRE, 10SRE-Access-Requests, 10LDAP-Access-Requests, 13Patch-For-Review: Grant Access to wmf for Chandler Diggs - https://phabricator.wikimedia.org/T433302#12172893 (10Dzahn) 05Open→03In progress [20:48:29] 06SRE, 10SRE-tools, 10Cumin, 06DC-Ops, and 2 others: add dcops group to run sre.hosts.downtime cookbook - https://phabricator.wikimedia.org/T433409#12172905 (10Dzahn) The change is about editing the permissions of the `datacenter-ops` shell group. And that group has the "approval:"-owner @wpao. So techni... [20:51:42] 06SRE, 10SRE-Access-Requests, 10SRE-tools, 10Cumin, and 3 others: add dcops group to run sre.hosts.downtime cookbook - https://phabricator.wikimedia.org/T433409#12172933 (10Dzahn) [20:53:02] 06SRE, 10LDAP-Access-Requests: Grant Access to ciadmin for Vaughn Walters - https://phabricator.wikimedia.org/T433615#12172939 (10Dzahn) 05Open→03In progress [20:53:24] 06SRE, 10SRE-Access-Requests, 06Release-Engineering-Team: Requesting access to Jenkins server for PWangai-WMF - https://phabricator.wikimedia.org/T433563#12172940 (10Dzahn) 05Open→03In progress [20:55:34] (03PS3) 10Aleksandar Mastilovic: Change severity of data eng Airflow alerts to "page" [alerts] - 10https://gerrit.wikimedia.org/r/1319189 (https://phabricator.wikimedia.org/T432312) [21:00:05] Deploy window Readers deployment window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260730T2100) [21:01:50] (03PS1) 10Dzahn: site: add zuul1005 as a zuul executor [puppet] - 10https://gerrit.wikimedia.org/r/1319547 (https://phabricator.wikimedia.org/T427353) [21:02:17] (03CR) 10JHathaway: [C:03+1] zuul: replace puppet legacy facts [puppet] - 10https://gerrit.wikimedia.org/r/1319519 (https://phabricator.wikimedia.org/T372666) (owner: 10Dzahn) [21:02:55] preparing to do a security deploy [21:07:41] preparing to run scap [21:12:06] running scap [21:15:00] !log Deployed security fix for T430601 [21:15:02] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [21:15:14] security deploy finished [21:17:19] maryum: OK for me to deploy a scap update? [21:22:51] !log dancy@deploy1003 Installing scap version "4.276.0" for 3 host(s) [21:24:46] !log dancy@deploy1003 Installation of scap version "4.276.0" completed for 3 hosts [21:30:03] (03CR) 10Andrea Denisse: [C:03+1] "LGTM, thank you!!" [puppet] - 10https://gerrit.wikimedia.org/r/1306792 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [21:32:39] (03CR) 10Dzahn: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1319519 (https://phabricator.wikimedia.org/T372666) (owner: 10Dzahn) [21:34:53] (03CR) 10Dzahn: [V:03+1 C:03+2] "https://puppet-compiler.wmflabs.org/output/1319519/7423/contint1002.wikimedia.org/index.html" [puppet] - 10https://gerrit.wikimedia.org/r/1319519 (https://phabricator.wikimedia.org/T372666) (owner: 10Dzahn) [21:35:17] !log dancy@deploy1003 Installing scap version "4.276.1" for 3 host(s) [21:36:03] (03CR) 10Aleksandar Mastilovic: "Got it - added the tests and made sure they pass locally." [alerts] - 10https://gerrit.wikimedia.org/r/1319189 (https://phabricator.wikimedia.org/T432312) (owner: 10Aleksandar Mastilovic) [21:36:29] 06SRE, 10SRE-Access-Requests, 06Release-Engineering-Team: Requesting access to Jenkins server for PWangai-WMF - https://phabricator.wikimedia.org/T433563#12173008 (10pwangai) I would say this may not be a temporary one-off. Right now as a team we are formulating a KR revolving around castor improvements, and... [21:37:11] !log dancy@deploy1003 Installation of scap version "4.276.1" completed for 3 hosts [21:38:39] (03PS1) 10Ahmon Dancy: scap.cfg.erb: Drop delay_messageblobstore_purge [puppet] - 10https://gerrit.wikimedia.org/r/1319551 (https://phabricator.wikimedia.org/T263872) [21:40:32] (03CR) 10Bking: [C:03+2] Change severity of data eng Airflow alerts to "page" [alerts] - 10https://gerrit.wikimedia.org/r/1319189 (https://phabricator.wikimedia.org/T432312) (owner: 10Aleksandar Mastilovic) [21:40:50] sorry dancy yes I see you worked it out [21:40:57] Yep. Thanks! [21:41:17] (03CR) 10Dzahn: [C:03+2] scap.cfg.erb: Drop delay_messageblobstore_purge [puppet] - 10https://gerrit.wikimedia.org/r/1319551 (https://phabricator.wikimedia.org/T263872) (owner: 10Ahmon Dancy) [21:41:50] there you go. if you want to repeat that [21:42:15] Awesome, Thanks mutante! [21:42:28] (03Merged) 10jenkins-bot: Change severity of data eng Airflow alerts to "page" [alerts] - 10https://gerrit.wikimedia.org/r/1319189 (https://phabricator.wikimedia.org/T432312) (owner: 10Aleksandar Mastilovic) [21:42:31] yw, running pupet on eh.. deploy2003 [21:43:56] FIRING: HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-rollback - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs-next - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [21:50:38] (03CR) 10JHathaway: "@cmooney@wikimedia.org I started looking at this today, but didn't finish, I'll pick it up tomorrow morning." [puppet] - 10https://gerrit.wikimedia.org/r/1319486 (https://phabricator.wikimedia.org/T433601) (owner: 10Cathal Mooney) [21:52:59] FIRING: NELByCountryHigh: Elevated Network Error Logging events (tcp.timed_out from RU) - https://wikitech.wikimedia.org/wiki/Network_monitoring#NEL_alerts - https://logstash.wikimedia.org/goto/5c8f4ca1413eda33128e5c5a35da7e28 - https://alerts.wikimedia.org/?q=alertname%3DNELByCountryHigh [21:58:49] (03PS5) 10Ahmon Dancy: wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1315968 (https://phabricator.wikimedia.org/T433104) [22:03:56] (03PS5) 10DLynch: Add script to get constructive edits for all wikis [puppet] - 10https://gerrit.wikimedia.org/r/1272633 (https://phabricator.wikimedia.org/T428490) (owner: 10Clare Ming) [22:07:59] FIRING: [2x] SystemdUnitFailed: generate_vrts_aliases.service on mx-in2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [22:12:31] (03PS2) 10Ahmon Dancy: wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE [mediawiki-config] (train-dev) - 10https://gerrit.wikimedia.org/r/1316090 (https://phabricator.wikimedia.org/T433104) [22:12:46] (03PS6) 10Ahmon Dancy: wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1315968 (https://phabricator.wikimedia.org/T433104) [22:12:59] FIRING: [3x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-3/1/2 (Transport: Hurricane Electric (dc4841.dal5)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [22:15:08] (03CR) 10Ahmon Dancy: wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (031 comment) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1315968 (https://phabricator.wikimedia.org/T433104) (owner: 10Ahmon Dancy) [22:17:59] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [22:35:13] (03PS11) 10Cwhite: profile: add initial opensearch security plugin config [puppet] - 10https://gerrit.wikimedia.org/r/1306792 (https://phabricator.wikimedia.org/T350516) [22:35:13] (03PS5) 10Cwhite: opensearch: add security-plugin specific configuration to opensearch.yml [puppet] - 10https://gerrit.wikimedia.org/r/1318792 (https://phabricator.wikimedia.org/T350516) [22:43:08] (03CR) 10Cwhite: "PCC success: https://puppet-compiler.wmflabs.org/output/1306003/9095/" [puppet] - 10https://gerrit.wikimedia.org/r/1318792 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [22:43:32] (03CR) 10Cwhite: profile: add initial opensearch security plugin config (033 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1306792 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [22:51:18] (03CR) 10Cwhite: [C:03+1] "Looks good!" [alerts] - 10https://gerrit.wikimedia.org/r/1304769 (https://phabricator.wikimedia.org/T407138) (owner: 10Hnowlan) [23:24:47] (03CR) 10Scott French: [C:03+1] php: rebuild to include Timo's APCU metrics fix [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1319453 (https://phabricator.wikimedia.org/T433312) (owner: 10Kamila Součková) [23:27:59] RESOLVED: NELByCountryHigh: Elevated Network Error Logging events (tcp.timed_out from RU) - https://wikitech.wikimedia.org/wiki/Network_monitoring#NEL_alerts - https://logstash.wikimedia.org/goto/5c8f4ca1413eda33128e5c5a35da7e28 - https://alerts.wikimedia.org/?q=alertname%3DNELByCountryHigh [23:31:59] FIRING: NELByCountryHigh: Elevated Network Error Logging events (tcp.timed_out from RU) - https://wikitech.wikimedia.org/wiki/Network_monitoring#NEL_alerts - https://logstash.wikimedia.org/goto/5c8f4ca1413eda33128e5c5a35da7e28 - https://alerts.wikimedia.org/?q=alertname%3DNELByCountryHigh [23:33:18] 06SRE, 10SRE-Access-Requests, 10SRE-tools, 10Cumin, and 3 others: add dcops group to run sre.hosts.downtime cookbook - https://phabricator.wikimedia.org/T433409#12173180 (10RobH) a:03wiki_willy Willy, Can you give your approval as DC Ops manager to increase the rights of the dc ops user group to include... [23:35:40] (03PS7) 10Aaron Schulz: Remove some API Gateway specific template files [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319114 (https://phabricator.wikimedia.org/T428625) [23:43:42] (03PS1) 10TrainBranchBot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1319554 [23:43:42] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1319554 (owner: 10TrainBranchBot) [23:51:59] RESOLVED: NELByCountryHigh: Elevated Network Error Logging events (tcp.timed_out from RU) - https://wikitech.wikimedia.org/wiki/Network_monitoring#NEL_alerts - https://logstash.wikimedia.org/goto/5c8f4ca1413eda33128e5c5a35da7e28 - https://alerts.wikimedia.org/?q=alertname%3DNELByCountryHigh [23:52:59] FIRING: [2x] GanetiBGPNoInPrefixes: ganeti2033 is not sending any prefix to lsw1-b7-codfw - group Ganeti4 - https://wikitech.wikimedia.org/wiki/Ganeti#GanetiBGPNoInPrefixes - https://alerts.wikimedia.org/?q=alertname%3DGanetiBGPNoInPrefixes [23:53:59] FIRING: NELByCountryHigh: Elevated Network Error Logging events (tcp.timed_out from RU) - https://wikitech.wikimedia.org/wiki/Network_monitoring#NEL_alerts - https://logstash.wikimedia.org/goto/5c8f4ca1413eda33128e5c5a35da7e28 - https://alerts.wikimedia.org/?q=alertname%3DNELByCountryHigh [23:56:35] (03Merged) 10jenkins-bot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1319554 (owner: 10TrainBranchBot)