[00:09:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.69% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [00:13:32] !log btullis@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet [00:14:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.9% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [00:20:23] FIRING: CertAlmostExpired: gNMI TLS certificate for lsw1-f6-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [00:20:24] !log btullis@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet [00:20:25] !log btullis@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet [00:20:30] !log btullis@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet [00:29:26] FIRING: [10x] BFDdown: BFD session down between cr1-eqiad and 2620:0:861:107:10:64:48:95 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [00:34:54] (03PS1) 10NguoiDungKhongDinhDanh: Set `noindex,nofollow` for User and User talk [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311141 (https://phabricator.wikimedia.org/T432311) [00:35:46] 06SRE, 10Wikimedia Australia, 10Wikimedia-Mailing-lists: Create new private announce-only mailing list for the ICIP project (Wikimedia Australia) - https://phabricator.wikimedia.org/T432082#12126591 (10AlphaLemur) Thanks for setting this up. We have done some testing and don't seem to be able to subscrib... [00:36:02] 06SRE, 10Wikimedia Australia, 10Wikimedia-Mailing-lists: Create new private announce-only mailing list for the ICIP project (Wikimedia Australia) - https://phabricator.wikimedia.org/T432082#12126593 (10AlphaLemur) 05Resolved→03Open [00:50:34] !log btullis@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet [00:57:09] !log btullis@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet [00:57:10] !log btullis@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet [00:57:16] !log btullis@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet [00:57:48] !log btullis@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet [01:03:37] FIRING: [9x] CoreRouterInterfaceDown: Core router interface down - cr2-eqiad:et-0/0/1 () - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [01:04:15] !log btullis@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet [01:04:16] !log btullis@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet [01:04:21] !log btullis@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet [01:04:54] !log btullis@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet [01:07:59] !log ryankemper@deploy2003 Started deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage [01:08:07] !log ryankemper@deploy2003 Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage (duration: 00m 46s) [01:08:44] !log ryankemper@deploy2003 Started deploy [wdqs/wdqs@e8fb00c] (wcqs): T430879 scap deploy post bookworm reimage [01:08:47] T430879: Migrate WCQS to Bookworm or later - https://phabricator.wikimedia.org/T430879 [01:08:51] !log ryankemper@deploy2003 Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): T430879 scap deploy post bookworm reimage (duration: 00m 23s) [01:11:00] !log btullis@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet [01:11:01] !log btullis@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet [01:11:07] !log btullis@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet [01:11:40] !log btullis@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1012.eqiad.wmnet [01:12:33] (03PS1) 10TrainBranchBot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1311143 [01:12:33] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1311143 (owner: 10TrainBranchBot) [01:14:57] (03PS1) 10Aleksandar Mastilovic: Add a Slack API alerting route for data engineering alerts [puppet] - 10https://gerrit.wikimedia.org/r/1311144 (https://phabricator.wikimedia.org/T432312) [01:15:31] (03CR) 10CI reject: [V:04-1] Add a Slack API alerting route for data engineering alerts [puppet] - 10https://gerrit.wikimedia.org/r/1311144 (https://phabricator.wikimedia.org/T432312) (owner: 10Aleksandar Mastilovic) [01:16:50] (03PS1) 10Aleksandar Mastilovic: Add Prometheus alerts for Airflow's main k8s instance [alerts] - 10https://gerrit.wikimedia.org/r/1311146 (https://phabricator.wikimedia.org/T432312) [01:17:39] !log btullis@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet [01:17:41] !log btullis@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet [01:17:47] !log btullis@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet [01:20:12] !log ryankemper@cumin2003 START - Cookbook sre.wdqs.data-transfer (T430879, restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards [01:20:15] T430879: Migrate WCQS to Bookworm or later - https://phabricator.wikimedia.org/T430879 [01:20:25] !log ryankemper@cumin2003 START - Cookbook sre.wdqs.data-transfer (T430879, restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards [01:21:10] (03Merged) 10jenkins-bot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1311143 (owner: 10TrainBranchBot) [01:23:43] (03PS2) 10Aleksandar Mastilovic: Add a Slack API alerting route for data engineering alerts [puppet] - 10https://gerrit.wikimedia.org/r/1311144 (https://phabricator.wikimedia.org/T432312) [01:23:53] (03CR) 10Aleksandar Mastilovic: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1311144 (https://phabricator.wikimedia.org/T432312) (owner: 10Aleksandar Mastilovic) [01:30:43] FIRING: [3x] JobUnavailable: Reduced availability for job jmx_wcqs_blazegraph in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [01:43:26] PROBLEM - SSH on arclamp2001 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/SSH/monitoring [01:47:31] FIRING: RedisInstanceDown: Redis instance down arclamp2001:9121 redis_arclamp - https://wikitech.wikimedia.org/wiki/Redis#Cluster_redis_arclamp - https://grafana.wikimedia.org/d/000000174/redis?orgId=1&var-site=codfw&var-job=redis_arclamp&var-instance=arclamp2001:9121 - https://alerts.wikimedia.org/?q=alertname%3DRedisInstanceDown [01:47:48] !log btullis@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet [01:50:43] FIRING: [4x] JobUnavailable: Reduced availability for job jmx_wcqs_blazegraph in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [01:54:09] !log btullis@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet [01:54:11] !log btullis@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet [01:54:17] !log btullis@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet [02:00:22] !log mwpresync@deploy2003 Started scap build-images: Publishing wmf/next image [02:03:16] RECOVERY - SSH on arclamp2001 is OK: SSH OK - OpenSSH_8.4p1 Debian-5+deb11u7 (protocol 2.0) https://wikitech.wikimedia.org/wiki/SSH/monitoring [02:05:43] FIRING: [4x] JobUnavailable: Reduced availability for job jmx_wcqs_blazegraph in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:06:59] !log mwpresync@deploy2003 Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) [02:07:31] RESOLVED: RedisInstanceDown: Redis instance down arclamp2001:9121 redis_arclamp - https://wikitech.wikimedia.org/wiki/Redis#Cluster_redis_arclamp - https://grafana.wikimedia.org/d/000000174/redis?orgId=1&var-site=codfw&var-job=redis_arclamp&var-instance=arclamp2001:9121 - https://alerts.wikimedia.org/?q=alertname%3DRedisInstanceDown [02:10:43] FIRING: [6x] JobUnavailable: Reduced availability for job jmx_wcqs_blazegraph in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:14:34] 10ops-eqsin, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 10netops: EQSIN:Switch refresh diagram and wiring - https://phabricator.wikimedia.org/T423724#12126803 (10Papaul) [02:15:43] FIRING: [6x] JobUnavailable: Reduced availability for job jmx_wcqs_blazegraph in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:16:27] 10ops-eqsin, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 10netops: EQSIN:Switch refresh diagram and wiring - https://phabricator.wikimedia.org/T423724#12126805 (10Papaul) [02:17:04] FIRING: HelmReleaseBadStatus: Helm release wdqs/main-external on k8s-dse@eqiad in state failed - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=eqiad&var-cluster=k8s-dse&var-namespace=wdqs - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [02:24:21] !log btullis@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1014.eqiad.wmnet [02:30:18] !log btullis@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet [02:30:19] !log btullis@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet [02:30:25] !log btullis@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet [02:35:29] !log btullis@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet [02:35:43] FIRING: [3x] JobUnavailable: Reduced availability for job jmx_wcqs_blazegraph in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:36:06] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) (T430879, restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards [02:36:09] T430879: Migrate WCQS to Bookworm or later - https://phabricator.wikimedia.org/T430879 [02:36:39] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) (T430879, restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards [02:41:31] !log btullis@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet [02:41:32] !log btullis@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet [02:41:37] !log btullis@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet [02:53:37] FIRING: [10x] CoreRouterInterfaceDown: Core router interface down - cr2-eqiad:et-0/0/1 () - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [02:57:22] FIRING: [10x] CoreRouterInterfaceDown: Core router interface down - cr2-eqiad:et-0/0/1 () - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [03:11:38] 10ops-eqsin, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 10netops: EQSIN:New switch setup/configuration - https://phabricator.wikimedia.org/T418439#12126823 (10Papaul) [03:11:42] !log btullis@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet [03:18:32] !log btullis@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet [03:18:33] !log btullis@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet [03:18:40] !log btullis@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet [03:20:46] !log btullis@cumin1003 END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet [03:25:28] 10ops-eqsin, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 10netops: EQSIN:Switch refresh diagram and wiring - https://phabricator.wikimedia.org/T423724#12126832 (10Papaul) [03:38:03] !log btullis@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet [03:38:04] !log btullis@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet [03:38:11] !log btullis@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet [03:40:59] (03CR) 10Andrea Denisse: [C:03+1] "LGTM, thank you!!" [puppet] - 10https://gerrit.wikimedia.org/r/1311028 (https://phabricator.wikimedia.org/T432226) (owner: 10Hnowlan) [03:57:46] FIRING: Not accepting/receiving prefixes from anycast BGP peer: Alert for device asw1-b4-magru.mgmt.magru.wmnet - Not accepting/receiving prefixes from anycast BGP peer - https://alerts.wikimedia.org/?q=alertname%3DNot+accepting%2Freceiving+prefixes+from+anycast+BGP+peer [04:08:16] !log btullis@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1021.eqiad.wmnet [04:16:51] FIRING: ProbeDown: Service centrallog1002:6514 has failed probes (tcp_rsyslog_receiver_ip4) - https://wikitech.wikimedia.org/wiki/TLS/Runbook#centrallog1002:6514 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [04:19:06] !log btullis@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet [04:19:07] !log btullis@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet [04:19:12] !log btullis@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet [04:20:39] FIRING: CertAlmostExpired: gNMI TLS certificate for lsw1-f6-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [04:21:51] RESOLVED: ProbeDown: Service centrallog1002:6514 has failed probes (tcp_rsyslog_receiver_ip4) - https://wikitech.wikimedia.org/wiki/TLS/Runbook#centrallog1002:6514 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [04:26:54] FIRING: KubernetesAPILatency: High Kubernetes API latency (LIST secrets) on k8s@eqiad - https://wikitech.wikimedia.org/wiki/Kubernetes - https://grafana.wikimedia.org/d/ddNd-sLnk/kubernetes-api-details?var-site=eqiad&var-cluster=k8s&var-latency_percentile=0.95&var-verb=LIST - https://alerts.wikimedia.org/?q=alertname%3DKubernetesAPILatency [04:29:26] FIRING: [10x] BFDdown: BFD session down between cr1-eqiad and 2620:0:861:107:10:64:48:95 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [04:45:28] PROBLEM - Host mr1-eqsin.oob IPv6 is DOWN: PING CRITICAL - Packet loss = 60%, RTA = 3815.63 ms [04:49:18] !log btullis@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1022.eqiad.wmnet [04:50:30] RECOVERY - Host mr1-eqsin.oob IPv6 is UP: PING OK - Packet loss = 0%, RTA = 217.55 ms [04:56:30] !log btullis@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet [04:56:31] !log btullis@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet [04:56:37] !log btullis@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet [05:06:54] RESOLVED: KubernetesAPILatency: High Kubernetes API latency (LIST secrets) on k8s@eqiad - https://wikitech.wikimedia.org/wiki/Kubernetes - https://grafana.wikimedia.org/d/ddNd-sLnk/kubernetes-api-details?var-site=eqiad&var-cluster=k8s&var-latency_percentile=0.95&var-verb=LIST - https://alerts.wikimedia.org/?q=alertname%3DKubernetesAPILatency [05:26:41] !log btullis@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet [05:37:14] !log btullis@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet [05:37:15] !log btullis@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet [05:37:20] !log btullis@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet [05:49:54] FIRING: KubernetesAPILatency: High Kubernetes API latency (LIST secrets) on k8s@eqiad - https://wikitech.wikimedia.org/wiki/Kubernetes - https://grafana.wikimedia.org/d/ddNd-sLnk/kubernetes-api-details?var-site=eqiad&var-cluster=k8s&var-latency_percentile=0.95&var-verb=LIST - https://alerts.wikimedia.org/?q=alertname%3DKubernetesAPILatency [06:00:04] Deploy window MediaWiki infrastructure (UTC early) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260716T0600) [06:00:04] marostegui, Amir1, and federico3: #bothumor Q:How do functions break up? A:They stop calling each other. Rise for Primary database switchover deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260716T0600). [06:04:54] RESOLVED: KubernetesAPILatency: High Kubernetes API latency (LIST secrets) on k8s@eqiad - https://wikitech.wikimedia.org/wiki/Kubernetes - https://grafana.wikimedia.org/d/ddNd-sLnk/kubernetes-api-details?var-site=eqiad&var-cluster=k8s&var-latency_percentile=0.95&var-verb=LIST - https://alerts.wikimedia.org/?q=alertname%3DKubernetesAPILatency [06:06:24] 06SRE, 10Wikimedia Australia, 10Wikimedia-Mailing-lists: Create new private announce-only mailing list for the ICIP project (Wikimedia Australia) - https://phabricator.wikimedia.org/T432082#12126884 (10jcrespo) 05Open→03Resolved Please note those are known/expected limitations, due to abuse/spam suff... [06:07:24] !log btullis@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet [06:10:17] 06SRE, 10Wikimedia Australia, 10Wikimedia-Mailing-lists: Create new private announce-only mailing list for the ICIP project (Wikimedia Australia) - https://phabricator.wikimedia.org/T432082#12126891 (10AlphaLemur) Thanks for the explanation - I appreciate it! I will know this for future list management.... [06:14:39] !log btullis@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet [06:14:40] !log btullis@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet [06:14:45] !log btullis@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet [06:17:04] FIRING: HelmReleaseBadStatus: Helm release wdqs/main-external on k8s-dse@eqiad in state failed - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=eqiad&var-cluster=k8s-dse&var-namespace=wdqs - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [06:19:56] (03PS1) 10Idiakeosemoahu: Revert "presto: Fix the resource-groups configuration" [puppet] - 10https://gerrit.wikimedia.org/r/1311304 [06:23:03] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for mbsantos - https://phabricator.wikimedia.org/T432258#12126897 (10jcrespo) [06:27:12] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for mbsantos - https://phabricator.wikimedia.org/T432258#12126902 (10jcrespo) p:05Triage→03High Hi, @MSantos, you already have production access, so the only requirement needed to deploy this is an OK from your Manager approving... [06:31:39] (03CR) 10Idiakeosemoahu: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1311304 (owner: 10Idiakeosemoahu) [06:31:43] (03CR) 10Idiakeosemoahu: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1311304 (owner: 10Idiakeosemoahu) [06:32:41] (03CR) 10Trueg: [C:03+1] wdqs: codfw: deploy test qlever image (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310533 (https://phabricator.wikimedia.org/T431271) (owner: 10Gmodena) [06:35:46] 06SRE, 10SRE-Access-Requests: Requesting access to Superset for jniren-ctr - https://phabricator.wikimedia.org/T432273#12126907 (10jcrespo) [06:35:58] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [06:41:45] 06SRE, 10SRE-Access-Requests: Requesting access to Superset for jniren-ctr - https://phabricator.wikimedia.org/T432273#12126918 (10jcrespo) Hi, @soworu to add you to the NDA group, we will need 2 things: your currently stated end of contract date and contact for contract (It could be @AAlikhan). And we will n... [06:42:13] (03PS1) 10Idiakeosemoahu: Revert "Use nginx module for protoproxy and disable notify" [puppet] - 10https://gerrit.wikimedia.org/r/1311311 [06:42:16] 06SRE, 10SRE-Access-Requests: Requesting access to Superset for jniren-ctr - https://phabricator.wikimedia.org/T432273#12126923 (10jcrespo) p:05Triage→03High [06:44:30] 10SRE-Access-Requests, 06Infrastructure-Foundations, 10LDAP-Access-Requests, 13Patch-For-Review: Superset data access request for abibendall - https://phabricator.wikimedia.org/T430938#12126933 (10jcrespo) a:05jcrespo→03None [06:44:50] !log btullis@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet [06:47:02] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply [06:47:15] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply [06:51:48] !log btullis@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet [06:51:50] !log btullis@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet [06:51:55] !log btullis@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet [06:53:59] (03CR) 10Idiakeosemoahu: "nice" [puppet] - 10https://gerrit.wikimedia.org/r/1311311 (owner: 10Idiakeosemoahu) [06:58:37] FIRING: [9x] CoreRouterInterfaceDown: Core router interface down - cr2-eqiad:et-0/0/1 () - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [07:00:04] Amir1, urbanecm, and awight: OwO what's this, a deployment window?? UTC morning backport window. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260716T0700). nyaa~ [07:00:05] No Gerrit patches in the queue for this window AFAICS. [07:05:16] (03CR) 10Idiakeosemoahu: [C:03+1] x509-bundle: skip popping first if we have an empty list [puppet] - 10https://gerrit.wikimedia.org/r/922147 (https://phabricator.wikimedia.org/T283001) (owner: 10Jbond) [07:10:30] (03CR) 10Idiakeosemoahu: [C:03+1] "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/922147 (https://phabricator.wikimedia.org/T283001) (owner: 10Jbond) [07:18:55] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie [07:19:02] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12127014 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host ms-be2062.codfw.wmnet with OS trixie [07:19:35] !log elukey@cumin1003 START - Cookbook sre.hosts.reboot-single for host bast2003.wikimedia.org [07:20:22] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-be1066.eqiad.wmnet with OS trixie [07:20:29] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12127045 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-be1066.eqiad.wmnet with OS trixie [07:21:41] (03CR) 10Brouberol: [C:03+1] stat hosts: Set up S3 config for wikidata platform team. [puppet] - 10https://gerrit.wikimedia.org/r/1310643 (https://phabricator.wikimedia.org/T432167) (owner: 10Bking) [07:21:56] !log btullis@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet [07:27:16] (03CR) 10Elukey: [C:03+2] profile::puppetserver::volatile: add metrics for airflow webrequest [puppet] - 10https://gerrit.wikimedia.org/r/1310118 (https://phabricator.wikimedia.org/T402512) (owner: 10Elukey) [07:28:07] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast2003.wikimedia.org [07:28:37] !log btullis@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet [07:28:38] !log btullis@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet [07:28:43] !log btullis@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet [07:29:18] !log btullis@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet [07:35:57] !log btullis@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet [07:35:58] !log btullis@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet [07:36:03] !log btullis@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet [07:36:36] !log btullis@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet [07:38:11] !log mvernon@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage [07:38:41] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage [07:42:54] (03PS3) 10Gerrit maintenance bot: mariadb: Promote db2241 to x3 master [puppet] - 10https://gerrit.wikimedia.org/r/1307084 (https://phabricator.wikimedia.org/T430925) [07:43:37] !log btullis@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet [07:43:38] !log btullis@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet [07:43:38] !log btullis@cumin1003 END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-eqiad [07:43:49] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage [07:44:13] !log cwilliams@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: Primary switchover x3 T430925 [07:44:19] T430925: Switchover x3 master (db2162 -> db2241) - https://phabricator.wikimedia.org/T430925 [07:45:07] !log cwilliams@cumin1003 dbctl commit (dc=all): 'Set db2241 with weight 0 T430925', diff saved to https://phabricator.wikimedia.org/P94868 and previous config saved to /var/cache/conftool/dbconfig/20260716-074507-cwilliams.json [07:47:56] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage [07:50:17] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply [07:50:27] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply [07:50:55] (03CR) 10CWilliams: [C:03+2] mariadb: Promote db2241 to x3 master [puppet] - 10https://gerrit.wikimedia.org/r/1307084 (https://phabricator.wikimedia.org/T430925) (owner: 10Gerrit maintenance bot) [07:52:20] !log Starting x3 codfw failover from db2162 to db2241 - T430925 [07:52:23] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [07:52:24] T430925: Switchover x3 master (db2162 -> db2241) - https://phabricator.wikimedia.org/T430925 [07:52:27] PROBLEM - SSH on arclamp2001 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/SSH/monitoring [07:53:14] !log cwilliams@cumin1003 dbctl commit (dc=all): 'Promote db2241 to x3 primary T430925', diff saved to https://phabricator.wikimedia.org/P94869 and previous config saved to /var/cache/conftool/dbconfig/20260716-075314-cwilliams.json [07:53:17] RECOVERY - SSH on arclamp2001 is OK: SSH OK - OpenSSH_8.4p1 Debian-5+deb11u7 (protocol 2.0) https://wikitech.wikimedia.org/wiki/SSH/monitoring [07:55:31] !log cwilliams@cumin1003 dbctl commit (dc=all): 'Depool db2162 T430925', diff saved to https://phabricator.wikimedia.org/P94870 and previous config saved to /var/cache/conftool/dbconfig/20260716-075530-cwilliams.json [07:56:26] !log cwilliams@cumin1003 START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover [07:57:46] FIRING: Not accepting/receiving prefixes from anycast BGP peer: Alert for device asw1-b4-magru.mgmt.magru.wmnet - Not accepting/receiving prefixes from anycast BGP peer - https://alerts.wikimedia.org/?q=alertname%3DNot+accepting%2Freceiving+prefixes+from+anycast+BGP+peer [07:58:07] (03PS4) 10Klausman: profiles/roles: mount CephFS on /home on ml-lab1002 [puppet] - 10https://gerrit.wikimedia.org/r/1311046 (https://phabricator.wikimedia.org/T380279) [07:58:17] (03CR) 10Klausman: profiles/roles: mount CephFS on /home on ml-lab1002 (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1311046 (https://phabricator.wikimedia.org/T380279) (owner: 10Klausman) [08:00:05] jeena and hashar: Deploy window MediaWiki train - Utc-7+Utc-0 Version (secondary timeslot) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260716T0800) [08:02:00] !log cwilliams@cumin1003 END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2162: Repooling after switchover [08:05:19] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1066.eqiad.wmnet with OS trixie [08:05:26] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12127145 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-be1066.eqiad.wmnet with OS trixie completed: - ms-be1066 (**PASS*... [08:10:46] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie [08:10:58] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12127153 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-be2062.codfw.wmnet with OS trixie completed: - ms-be2062 (**PASS*... [08:15:50] !log cgoubert@cumin1003 START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-eqiad [08:16:12] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet [08:16:23] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-be1067.eqiad.wmnet with OS trixie [08:16:36] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12127159 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-be1067.eqiad.wmnet with OS trixie [08:19:27] (03CR) 10Klausman: [C:03+2] profiles/roles: mount CephFS on /home on ml-lab1002 [puppet] - 10https://gerrit.wikimedia.org/r/1311046 (https://phabricator.wikimedia.org/T380279) (owner: 10Klausman) [08:20:39] FIRING: CertAlmostExpired: gNMI TLS certificate for lsw1-f6-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [08:21:55] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host ms-be2064.codfw.wmnet with OS trixie [08:21:56] !log cwilliams@cumin1003 START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover [08:22:04] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12127173 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host ms-be2064.codfw.wmnet with OS trixie [08:26:11] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet [08:28:30] jouncebot: nowandnext [08:28:30] For the next 1 hour(s) and 31 minute(s): MediaWiki train - Utc-7+Utc-0 Version (secondary timeslot) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260716T0800) [08:28:30] In 1 hour(s) and 31 minute(s): MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260716T1000) [08:29:26] FIRING: [10x] BFDdown: BFD session down between cr1-eqiad and 2620:0:861:107:10:64:48:95 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [08:29:27] (03PS1) 10Jelto: admin_ng: pin calico chart to current version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311394 (https://phabricator.wikimedia.org/T307943) [08:29:58] (03PS2) 10Jelto: admin_ng: pin calico chart to current version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311394 (https://phabricator.wikimedia.org/T427400) [08:30:30] (03CR) 10Jelto: "like in I0ebc5486b540c26787708cd4040b291422a4a5d3 ?" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306307 (https://phabricator.wikimedia.org/T427400) (owner: 10Jelto) [08:33:45] !log mvernon@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage [08:33:49] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet [08:33:52] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet [08:34:12] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1041-1050].eqiad.wmnet [08:34:33] (03PS1) 10Elukey: site.pp: comment change [puppet] - 10https://gerrit.wikimedia.org/r/1311395 [08:35:15] (03CR) 10Klausman: [C:03+1] site.pp: comment change [puppet] - 10https://gerrit.wikimedia.org/r/1311395 (owner: 10Elukey) [08:35:28] (03CR) 10Elukey: [C:03+2] site.pp: comment change [puppet] - 10https://gerrit.wikimedia.org/r/1311395 (owner: 10Elukey) [08:39:25] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage [08:39:31] (03PS1) 10A-pizzata: Daily sqoops for logging and cu_log tables [puppet] - 10https://gerrit.wikimedia.org/r/1311397 (https://phabricator.wikimedia.org/T432208) [08:40:04] (03CR) 10CI reject: [V:04-1] Daily sqoops for logging and cu_log tables [puppet] - 10https://gerrit.wikimedia.org/r/1311397 (https://phabricator.wikimedia.org/T432208) (owner: 10A-pizzata) [08:40:19] PROBLEM - SSH on an-worker1192 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/SSH/monitoring [08:40:21] 10SRE-swift-storage, 06Data-Platform-SRE (2026-07-03 - 2026-07-31): RdfStreamingUpdaterSpaceUsageTooHigh - https://phabricator.wikimedia.org/T431506#12127254 (10dcausse) The growth on `AUTH_search-update-pipeline` is due to a misconfiguration of the search job, it did not properly delete its automatic savepoin... [08:41:01] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1041-1050].eqiad.wmnet [08:41:13] 10ops-codfw, 10ops-drmrs, 10ops-eqdfw, 10ops-eqiad, and 5 others: Fix cables with placeholders names and planned status - https://phabricator.wikimedia.org/T432317 (10ayounsi) 03NEW p:05Triage→03Low [08:41:47] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage [08:45:03] (03CR) 10DCausse: [C:03+2] cirrus: change the restart strategy to exponential-delay [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309210 (https://phabricator.wikimedia.org/T428863) (owner: 10DCausse) [08:47:25] (03Merged) 10jenkins-bot: cirrus: change the restart strategy to exponential-delay [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309210 (https://phabricator.wikimedia.org/T428863) (owner: 10DCausse) [08:48:09] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage [08:48:28] !log elukey@cumin1003 START - Cookbook sre.hosts.reboot-single for host ping1004.eqiad.wmnet [08:49:00] !log dcausse@deploy2003 helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply [08:51:15] (03PS1) 10DCausse: cirrus: fix typo in restart-strategy config option [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311399 (https://phabricator.wikimedia.org/T428863) [08:51:18] !log dcausse@deploy2003 helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply [08:51:50] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1041-1050].eqiad.wmnet [08:51:52] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1041-1050].eqiad.wmnet [08:51:58] (03CR) 10DCausse: [C:03+2] cirrus: fix typo in restart-strategy config option [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311399 (https://phabricator.wikimedia.org/T428863) (owner: 10DCausse) [08:52:11] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping1004.eqiad.wmnet [08:52:12] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet [08:52:51] !log elukey@cumin1003 START - Cookbook sre.hosts.reboot-single for host ping2004.codfw.wmnet [08:54:45] (03Merged) 10jenkins-bot: cirrus: fix typo in restart-strategy config option [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311399 (https://phabricator.wikimedia.org/T428863) (owner: 10DCausse) [08:55:04] (03CR) 10Joal: "Some nits." [puppet] - 10https://gerrit.wikimedia.org/r/1311397 (https://phabricator.wikimedia.org/T432208) (owner: 10A-pizzata) [08:56:40] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping2004.codfw.wmnet [08:57:35] !log bump space for prometheus k8s-dse in eqiad [08:57:36] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [08:57:39] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet [08:58:14] 06SRE, 10SRE-Access-Requests: Requesting access to Superset for jniren-ctr - https://phabricator.wikimedia.org/T432273#12127359 (10Jniren-ctr) HI @jcrespo, as part of the Main Services Agreement I have signed it includes the below section on confidential information but I have not signed a separate NDA. Pleas... [08:59:18] jouncebot: nowandnext [08:59:18] For the next 1 hour(s) and 0 minute(s): MediaWiki train - Utc-7+Utc-0 Version (secondary timeslot) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260716T0800) [08:59:18] In 1 hour(s) and 0 minute(s): MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260716T1000) [08:59:50] !log dcausse@deploy2003 helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply [09:00:08] !log dcausse@deploy2003 helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply [09:01:19] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1067.eqiad.wmnet with OS trixie [09:01:31] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12127389 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-be1067.eqiad.wmnet with OS trixie completed: - ms-be1067 (**PASS*... [09:07:27] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: Repooling after switchover [09:07:43] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet [09:07:46] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet [09:08:05] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1067-1076].eqiad.wmnet [09:09:24] (03PS2) 10A-pizzata: Daily sqoops for logging and cu_log tables [puppet] - 10https://gerrit.wikimedia.org/r/1311397 (https://phabricator.wikimedia.org/T432208) [09:09:57] (03CR) 10CI reject: [V:04-1] Daily sqoops for logging and cu_log tables [puppet] - 10https://gerrit.wikimedia.org/r/1311397 (https://phabricator.wikimedia.org/T432208) (owner: 10A-pizzata) [09:10:29] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2064.codfw.wmnet with OS trixie [09:10:37] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12127407 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-be2064.codfw.wmnet with OS trixie completed: - ms-be2064 (**PASS*... [09:10:47] (03PS3) 10A-pizzata: Daily sqoops for logging and cu_log tables [puppet] - 10https://gerrit.wikimedia.org/r/1311397 (https://phabricator.wikimedia.org/T432208) [09:11:20] (03CR) 10CI reject: [V:04-1] Daily sqoops for logging and cu_log tables [puppet] - 10https://gerrit.wikimedia.org/r/1311397 (https://phabricator.wikimedia.org/T432208) (owner: 10A-pizzata) [09:13:31] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1067-1076].eqiad.wmnet [09:13:40] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-be1068.eqiad.wmnet with OS trixie [09:14:41] (03PS3) 10Blake: wmnet: Add CNAME records for mw-pretrain. [dns] - 10https://gerrit.wikimedia.org/r/1311074 (https://phabricator.wikimedia.org/T427668) [09:15:00] (03PS4) 10A-pizzata: Daily sqoops for logging and cu_log tables [puppet] - 10https://gerrit.wikimedia.org/r/1311397 (https://phabricator.wikimedia.org/T432208) [09:15:33] (03CR) 10CI reject: [V:04-1] Daily sqoops for logging and cu_log tables [puppet] - 10https://gerrit.wikimedia.org/r/1311397 (https://phabricator.wikimedia.org/T432208) (owner: 10A-pizzata) [09:20:17] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host ms-be2065.codfw.wmnet with OS trixie [09:20:25] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12127474 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host ms-be2065.codfw.wmnet with OS trixie [09:21:44] !log gmodena@deploy2003 helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply [09:24:06] !log gmodena@deploy2003 helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply [09:24:09] !log gmodena@deploy2003 helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply [09:24:18] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1067-1076].eqiad.wmnet [09:24:21] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1067-1076].eqiad.wmnet [09:24:40] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet [09:24:49] (03CR) 10Blake: wmnet: Add CNAME records for mw-pretrain. (031 comment) [dns] - 10https://gerrit.wikimedia.org/r/1311074 (https://phabricator.wikimedia.org/T427668) (owner: 10Blake) [09:25:27] !log gmodena@deploy2003 helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply [09:25:30] !log gmodena@deploy2003 helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply [09:25:57] !log elukey@cumin1003 START - Cookbook sre.hosts.reboot-single for host netbox2003.codfw.wmnet [09:26:53] (03CR) 10Blake: [C:03+2] service-catalog: add a new service for mw-pretrain. [puppet] - 10https://gerrit.wikimedia.org/r/1310999 (https://phabricator.wikimedia.org/T427668) (owner: 10Blake) [09:28:36] PROBLEM - OSPF status on cr1-eqiad is CRITICAL: OSPFv2: 7/8 UP : OSPFv3: 7/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [09:29:56] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox2003.codfw.wmnet [09:30:13] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet [09:31:36] RECOVERY - OSPF status on cr1-eqiad is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [09:34:19] FIRING: [3x] HelmReleaseBadStatus: Helm release wdqs/main-external on k8s-dse@codfw in state pending-rollback - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [09:37:10] !log gmodena@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply [09:37:20] mvernon@cumin1003 reimage (PID 289386) is awaiting input [09:37:56] (03PS1) 10Federico Ceratto: sre.mysql: Reorder imports [cookbooks] - 10https://gerrit.wikimedia.org/r/1311404 [09:37:56] (03CR) 10Federico Ceratto: "Just a basic cleanup with ruff" [cookbooks] - 10https://gerrit.wikimedia.org/r/1311404 (owner: 10Federico Ceratto) [09:39:19] FIRING: [3x] HelmReleaseBadStatus: Helm release wdqs/main-external on k8s-dse@codfw in state pending-rollback - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [09:39:30] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet [09:39:33] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet [09:39:48] !log gmodena@deploy2003 helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply [09:39:53] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet [09:40:04] !log gmodena@deploy2003 helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply [09:40:08] (03PS1) 10Federico Ceratto: sre.mysql.pool: support esX sections [cookbooks] - 10https://gerrit.wikimedia.org/r/1311405 (https://phabricator.wikimedia.org/T430769) [09:40:08] (03CR) 10Federico Ceratto: "This is a prototype version - we need to discuss what to do when the DC-master of a RW section is being depooled." [cookbooks] - 10https://gerrit.wikimedia.org/r/1311405 (https://phabricator.wikimedia.org/T430769) (owner: 10Federico Ceratto) [09:40:08] (03CR) 10CI reject: [V:04-1] sre.mysql.pool: support esX sections [cookbooks] - 10https://gerrit.wikimedia.org/r/1311405 (https://phabricator.wikimedia.org/T430769) (owner: 10Federico Ceratto) [09:40:27] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage [09:43:24] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage [09:46:31] !log gmodena@deploy2003 helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply [09:46:40] !log gmodena@deploy2003 helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply [09:47:11] !log gmodena@deploy2003 helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply [09:48:37] 10ops-codfw, 10ops-eqiad, 07sre-alert-triage, 06DC-Ops: Alert in need of triage: NetboxAccounting - https://phabricator.wikimedia.org/T428132#12127583 (10ayounsi) a:05ayounsi→03None Reassigning the task to DCops/procurement as they're alerts about tasks, asset tag or serial numbers missmatch. https://... [09:49:19] FIRING: [4x] HelmReleaseBadStatus: Helm release wdqs/main-external on k8s-dse@codfw in state pending-rollback - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [09:49:54] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet [09:50:49] 06SRE, 10SRE-Access-Requests: Requesting access to Superset for jniren-ctr - https://phabricator.wikimedia.org/T432273#12127603 (10jcrespo) I am not a lawyer, so @KFrancis will confirm. [09:52:14] !log gmodena@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply [09:54:11] (03CR) 10A-pizzata: Daily sqoops for logging and cu_log tables (037 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1311397 (https://phabricator.wikimedia.org/T432208) (owner: 10A-pizzata) [09:54:19] FIRING: [4x] HelmReleaseBadStatus: Helm release wdqs/main-external on k8s-dse@codfw in state pending-rollback - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [09:57:18] !log dcausse@deploy2003 helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply [09:57:25] !log dcausse@deploy2003 helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply [09:58:23] !log mvernon@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage [09:59:19] FIRING: [4x] HelmReleaseBadStatus: Helm release wdqs/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [09:59:45] !log gmodena@deploy2003 helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply [10:00:02] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet [10:00:05] Deploy window MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260716T1000) [10:00:05] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet [10:00:25] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1121-1130].eqiad.wmnet [10:02:04] FIRING: [4x] HelmReleaseBadStatus: Helm release wdqs/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [10:03:30] !log gmodena@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply [10:03:36] !log gmodena@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply [10:03:38] (03PS1) 10Santiago Faci: Test Kitchen UI: Deploy v1.4.8 release to staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311413 (https://phabricator.wikimedia.org/T431015) [10:04:10] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage [10:04:19] FIRING: [5x] HelmReleaseBadStatus: Helm release wdqs/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [10:04:44] (03PS2) 10Santiago Faci: Test Kitchen UI: Deploy v1.4.9 release to staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311413 (https://phabricator.wikimedia.org/T431015) [10:05:40] (03PS1) 10Santiago Faci: Test Kitchen UI: Deploy v1.4.9 release to production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311414 (https://phabricator.wikimedia.org/T431015) [10:06:14] !log dcausse@deploy2003 helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply [10:06:23] !log dcausse@deploy2003 helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply [10:06:32] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1121-1130].eqiad.wmnet [10:07:44] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2065.codfw.wmnet with OS trixie [10:07:53] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12127641 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-be2065.codfw.wmnet with OS trixie completed: - ms-be2065 (**PASS*... [10:13:17] 10ops-codfw, 10ops-drmrs, 10ops-eqdfw, 10ops-eqiad, and 6 others: Fix cables with placeholders names and planned status - https://phabricator.wikimedia.org/T432317#12127646 (10cmooney) [10:15:11] 10ops-codfw, 10ops-drmrs, 10ops-eqdfw, 10ops-eqiad, and 6 others: Fix cables with placeholders names and planned status - https://phabricator.wikimedia.org/T432317#12127659 (10cmooney) I updated the desc to give a bit more info and see which ones I might have been responsible for / insight on. One thing t... [10:15:58] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1121-1130].eqiad.wmnet [10:16:01] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1121-1130].eqiad.wmnet [10:16:24] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1131-1140].eqiad.wmnet [10:19:07] (03PS1) 10MVernon: swift: drain remaining 3 bullseye old/style nodes [puppet] - 10https://gerrit.wikimedia.org/r/1311416 (https://phabricator.wikimedia.org/T429630) [10:19:48] (03CR) 10Aklapper: [V:03+2 C:03+2] Update translations [phabricator/translations] (wmf/stable) - 10https://gerrit.wikimedia.org/r/1311078 (owner: 10Pppery) [10:21:07] !log gmodena@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply [10:21:10] !log gmodena@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply [10:21:18] !log gmodena@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply [10:21:26] !log gmodena@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply [10:22:26] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host ms-be2066.codfw.wmnet with OS trixie [10:22:28] (03PS1) 10JavierMonton: stream: pageview-trending-relative + webrequest-page-view [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311418 (https://phabricator.wikimedia.org/T432204) [10:22:38] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12127668 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host ms-be2066.codfw.wmnet with OS trixie [10:23:00] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1131-1140].eqiad.wmnet [10:24:31] RECOVERY - SSH on an-worker1192 is OK: SSH OK - OpenSSH_8.4p1 Debian-5+deb11u5 (protocol 2.0) https://wikitech.wikimedia.org/wiki/SSH/monitoring [10:28:23] (03CR) 10Aqu: [C:03+1] stream: pageview-trending-relative + webrequest-page-view [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311418 (https://phabricator.wikimedia.org/T432204) (owner: 10JavierMonton) [10:33:29] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1131-1140].eqiad.wmnet [10:33:32] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1131-1140].eqiad.wmnet [10:33:51] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1141-1150].eqiad.wmnet [10:35:58] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [10:36:26] !log gmodena@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply [10:37:30] !log gmodena@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply [10:39:19] FIRING: [2x] HelmReleaseBadStatus: Helm release wdqs/scholarly-external on k8s-dse@codfw in state pending-rollback - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [10:39:30] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1141-1150].eqiad.wmnet [10:40:14] (03CR) 10Phuedx: [C:03+1] Test Kitchen UI: Deploy v1.4.9 release to staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311413 (https://phabricator.wikimedia.org/T431015) (owner: 10Santiago Faci) [10:40:24] (03CR) 10Phuedx: [C:03+1] Test Kitchen UI: Deploy v1.4.9 release to production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311414 (https://phabricator.wikimedia.org/T431015) (owner: 10Santiago Faci) [10:41:54] 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 10netops: Netbox report: alert on cables in 'planned' state for more than two months - https://phabricator.wikimedia.org/T432329 (10cmooney) 03NEW p:05Triage→03Low [10:42:24] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage [10:45:51] (03PS2) 10Blake: kubernetes: Add a debug deployment for mw-pretrain. [puppet] - 10https://gerrit.wikimedia.org/r/1311049 (https://phabricator.wikimedia.org/T427668) [10:45:51] (03CR) 10Blake: "I'm also happy to include the other new deployments in this change if that seems reasonable (jobrunner, jobrunner-canary, and main-canary)" [puppet] - 10https://gerrit.wikimedia.org/r/1311049 (https://phabricator.wikimedia.org/T427668) (owner: 10Blake) [10:47:02] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1141-1150].eqiad.wmnet [10:47:04] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1141-1150].eqiad.wmnet [10:47:25] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1151-1160].eqiad.wmnet [10:49:01] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage [10:49:19] (03PS2) 10Federico Ceratto: sre.mysql.pool: support esX sections [cookbooks] - 10https://gerrit.wikimedia.org/r/1311405 (https://phabricator.wikimedia.org/T430769) [10:49:19] FIRING: [2x] HelmReleaseBadStatus: Helm release wdqs/scholarly-external on k8s-dse@codfw in state pending-rollback - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [10:51:47] (03CR) 10Federico Ceratto: [C:03+2] swift: drain remaining 3 bullseye old/style nodes [puppet] - 10https://gerrit.wikimedia.org/r/1311416 (https://phabricator.wikimedia.org/T429630) (owner: 10MVernon) [10:52:19] (03CR) 10Federico Ceratto: [C:03+1] "LGTM, I ran ls /srv/swift-storage as discussed" [puppet] - 10https://gerrit.wikimedia.org/r/1311416 (https://phabricator.wikimedia.org/T429630) (owner: 10MVernon) [10:52:59] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1151-1160].eqiad.wmnet [10:54:30] !log gmodena@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply [10:55:32] !log cwilliams@cumin1003 START - Cookbook sre.mysql.major-upgrade [10:55:33] !log cwilliams@cumin1003 END (FAIL) - Cookbook sre.mysql.major-upgrade (exit_code=99) [10:56:00] (03CR) 10MVernon: [C:03+2] swift: drain remaining 3 bullseye old/style nodes [puppet] - 10https://gerrit.wikimedia.org/r/1311416 (https://phabricator.wikimedia.org/T429630) (owner: 10MVernon) [10:56:39] jouncebot: nowandnext [10:56:40] For the next 0 hour(s) and 3 minute(s): MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260716T1000) [10:56:40] In 1 hour(s) and 3 minute(s): Mobileapps/RESTBase/Wikifeeds (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260716T1200) [10:56:51] perfection [10:57:04] FIRING: [2x] HelmReleaseBadStatus: Helm release wdqs/scholarly-external on k8s-dse@codfw in state pending-rollback - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [10:57:58] (03PS1) 10Ladsgroup: Upload: Do not throw for failure to save a chunk file [core] (wmf/1.47.0-wmf.11) - 10https://gerrit.wikimedia.org/r/1311420 (https://phabricator.wikimedia.org/T430986) [10:58:12] (03CR) 10Ladsgroup: [C:03+2] Upload: Do not throw for failure to save a chunk file [core] (wmf/1.47.0-wmf.11) - 10https://gerrit.wikimedia.org/r/1311420 (https://phabricator.wikimedia.org/T430986) (owner: 10Ladsgroup) [10:58:37] FIRING: [9x] CoreRouterInterfaceDown: Core router interface down - cr2-eqiad:et-0/0/1 () - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [10:59:49] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1068.eqiad.wmnet with OS trixie [11:00:01] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12127812 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-be1068.eqiad.wmnet with OS trixie completed... [11:03:09] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1151-1160].eqiad.wmnet [11:03:11] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1151-1160].eqiad.wmnet [11:03:35] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet [11:04:30] (03PS11) 10Effie Mouzeli: trafficserver: Remove XWD routing for /w/rest.php mw-debug backend [puppet] - 10https://gerrit.wikimedia.org/r/1310103 (https://phabricator.wikimedia.org/T428909) [11:05:24] !log gmodena@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply [11:05:51] (03CR) 10Effie Mouzeli: trafficserver: Remove XWD routing for /w/rest.php mw-debug backend (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1310103 (https://phabricator.wikimedia.org/T428909) (owner: 10Effie Mouzeli) [11:08:55] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2066.codfw.wmnet with OS trixie [11:09:10] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12127845 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-be2066.codfw.wmnet with OS trixie completed... [11:09:36] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet [11:13:18] (03CR) 10TrainBranchBot: [C:03+2] "Approved by krinkle@deploy2003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310596 (https://phabricator.wikimedia.org/T349582) (owner: 10Krinkle) [11:13:51] !log gmodena@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply [11:13:54] !log gmodena@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply [11:14:24] !log gmodena@deploy2003 helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply [11:15:40] !log gmodena@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply [11:15:52] (03CR) 10Santiago Faci: [C:03+2] Test Kitchen UI: Deploy v1.4.9 release to staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311413 (https://phabricator.wikimedia.org/T431015) (owner: 10Santiago Faci) [11:20:25] (03PS3) 10Effie Mouzeli: hiera: add profile::server_depool keys for redis and memcached. [puppet] - 10https://gerrit.wikimedia.org/r/1310566 (https://phabricator.wikimedia.org/T430930) [11:20:29] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet [11:20:32] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet [11:20:52] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1246-1255].eqiad.wmnet [11:22:28] !log gmodena@deploy2003 helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply [11:23:06] (03PS2) 10Ladsgroup: swift: Migrate storage of results to asyncio [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1310618 (https://phabricator.wikimedia.org/T431767) [11:23:06] (03PS2) 10Ladsgroup: images: Move loading of memcached key to asyncio [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1310623 (https://phabricator.wikimedia.org/T431767) [11:23:18] PROBLEM - Host wikikube-worker1163 is DOWN: PING CRITICAL - Packet loss = 60%, RTA = 4258.25 ms [11:23:58] !log jiji@cumin1003 START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw [11:24:01] (03CR) 10Ladsgroup: "Thanks for catching those. Fixed now." [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1310618 (https://phabricator.wikimedia.org/T431767) (owner: 10Ladsgroup) [11:24:12] RECOVERY - Host wikikube-worker1163 is UP: PING OK - Packet loss = 0%, RTA = 0.31 ms [11:24:24] !log jiji@cumin1003 START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-eqiad [11:31:27] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1246-1255].eqiad.wmnet [11:37:22] (03PS1) 10Btullis: logstash: convert opensearch server logs to ECS [puppet] - 10https://gerrit.wikimedia.org/r/1311422 (https://phabricator.wikimedia.org/T324335) [11:37:24] (03PS1) 10Btullis: rsyslog: support imfile readTimeout in rsyslog::input::file [puppet] - 10https://gerrit.wikimedia.org/r/1311423 (https://phabricator.wikimedia.org/T324335) [11:37:27] (03PS1) 10Btullis: opensearch: ship cirrus server JSON logs via rsyslog imfile [puppet] - 10https://gerrit.wikimedia.org/r/1311424 (https://phabricator.wikimedia.org/T324335) [11:37:29] (03PS1) 10Btullis: cirrus: ship server JSON logs on production clusters [puppet] - 10https://gerrit.wikimedia.org/r/1311425 (https://phabricator.wikimedia.org/T324335) [11:37:34] (03PS1) 10Btullis: cirrus: remove the log4j syslog appender log shipping path [puppet] - 10https://gerrit.wikimedia.org/r/1311426 (https://phabricator.wikimedia.org/T324335) [11:39:06] (03CR) 10Ladsgroup: "It's still in the queue... I'll come back to it later" [core] (wmf/1.47.0-wmf.11) - 10https://gerrit.wikimedia.org/r/1311420 (https://phabricator.wikimedia.org/T430986) (owner: 10Ladsgroup) [11:41:49] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1246-1255].eqiad.wmnet [11:41:51] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1246-1255].eqiad.wmnet [11:42:10] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet [11:43:37] !log jiji@cumin1003 END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw [11:44:22] (03PS2) 10Gmodena: wdqs: version bump qlever [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310533 (https://phabricator.wikimedia.org/T431271) [11:44:28] !log jiji@cumin1003 END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-eqiad [11:45:00] 10SRE-Access-Requests, 06Infrastructure-Foundations, 10LDAP-Access-Requests, 13Patch-For-Review: Superset data access request for abibendall - https://phabricator.wikimedia.org/T430938#12127950 (10ABendall-WMF) Please could I ask if there is anything I should be doing to unblock the LDAP new workflow appro... [11:46:03] (03PS1) 10Federico Ceratto: management.yaml: Add python3-pydantic package to cumin [puppet] - 10https://gerrit.wikimedia.org/r/1311427 (https://phabricator.wikimedia.org/T430769) [11:46:03] (03CR) 10Federico Ceratto: "Related to https://gerrit.wikimedia.org/r/c/operations/cookbooks/+/1311405 , as discussed on IRC" [puppet] - 10https://gerrit.wikimedia.org/r/1311427 (https://phabricator.wikimedia.org/T430769) (owner: 10Federico Ceratto) [11:46:27] (03PS2) 10Federico Ceratto: management.yaml: Add python3-pydantic package to cumin [puppet] - 10https://gerrit.wikimedia.org/r/1311427 (https://phabricator.wikimedia.org/T430769) [11:46:30] (03CR) 10Santiago Faci: [C:03+2] Test Kitchen UI: Deploy v1.4.9 release to production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311414 (https://phabricator.wikimedia.org/T431015) (owner: 10Santiago Faci) [11:47:47] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet [11:48:42] !log sukhe@cumin1003 START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqsin [11:48:58] !log sukhe@cumin1003 START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqsin [11:49:37] !log sukhe@cumin1003 START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_esams [11:49:44] !log sukhe@cumin1003 START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_esams [11:50:00] !log sukhe@cumin1003 START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqiad [11:50:05] !log gmodena@deploy2003 helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply [11:50:18] !log sukhe@cumin1003 START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqiad [11:50:19] !log gmodena@deploy2003 helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply [11:51:23] 06SRE, 06Infrastructure-Foundations, 10netops: Investigate using BGP addpath for unicast IBGP spine/leaf pods - https://phabricator.wikimedia.org/T402640#12127963 (10cmooney) >>! In T402640#11121128, @ayounsi wrote: > If I understand correctly we currently get some "per rack" load balancing, where `E3` might... [11:51:32] !log gmodena@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply [11:51:49] !log gmodena@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply [11:52:47] (03CR) 10Trueg: [C:03+1] wdqs: version bump qlever [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310533 (https://phabricator.wikimedia.org/T431271) (owner: 10Gmodena) [11:53:09] !log jiji@cumin1003 START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-eqiad [11:53:17] !log gmodena@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply [11:53:56] (03CR) 10Effie Mouzeli: [C:03+2] hiera: add profile::server_depool keys for redis and memcached. [puppet] - 10https://gerrit.wikimedia.org/r/1310566 (https://phabricator.wikimedia.org/T430930) (owner: 10Effie Mouzeli) [11:54:18] !log gmodena@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply [11:54:50] 06SRE, 06Infrastructure-Foundations, 10netops: Investigate using BGP addpath for unicast IBGP spine/leaf pods - https://phabricator.wikimedia.org/T402640#12127974 (10ayounsi) Yep, I refreshed my memory since then and I aadd path is a must have for proper balancing! +1 [11:54:58] !log sukhe@cumin1003 START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P{dns5003*} or P{dns7002*}) and (A:dnsbox) [11:54:58] !log sukhe@cumin1003 cookbooks.sre.dns.roll-reboot begin reboot of dns5004.wikimedia.org [11:55:40] (03PS1) 10Gmodena: wdqs: bump init container memory and startup probe timeout [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311429 (https://phabricator.wikimedia.org/T429150) [11:56:11] (03CR) 10Santiago Faci: [C:03+2] "recheck" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311413 (https://phabricator.wikimedia.org/T431015) (owner: 10Santiago Faci) [11:57:46] FIRING: Not accepting/receiving prefixes from anycast BGP peer: Alert for device asw1-b4-magru.mgmt.magru.wmnet - Not accepting/receiving prefixes from anycast BGP peer - https://alerts.wikimedia.org/?q=alertname%3DNot+accepting%2Freceiving+prefixes+from+anycast+BGP+peer [11:58:45] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet [11:58:48] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet [11:59:12] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1267-1276].eqiad.wmnet [12:00:00] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp1101.eqiad.wmnet [12:00:05] Deploy window Mobileapps/RESTBase/Wikifeeds (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260716T1200) [12:00:57] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp5025.eqsin.wmnet [12:01:13] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp5017.eqsin.wmnet [12:01:16] !log jiji@cumin1003 START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw [12:01:37] !log gmodena@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply [12:01:46] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp3074.esams.wmnet [12:01:56] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp3066.esams.wmnet [12:02:06] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp1100.eqiad.wmnet [12:02:44] !log gmodena@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply [12:03:41] !log sukhe@cumin1003 cookbooks.sre.dns.roll-reboot finished rebooting dns5004.wikimedia.org [12:03:49] (03CR) 10Federico Ceratto: "Done" [cookbooks] - 10https://gerrit.wikimedia.org/r/1277076 (https://phabricator.wikimedia.org/T419874) (owner: 10Federico Ceratto) [12:03:53] (03CR) 10Federico Ceratto: [C:03+2] sre.mysql.global-read-only Set all sections as RO/RW [cookbooks] - 10https://gerrit.wikimedia.org/r/1277076 (https://phabricator.wikimedia.org/T419874) (owner: 10Federico Ceratto) [12:04:11] FIRING: [14x] BFDdown: BFD session down between cr1-eqiad and 2620:0:861:107:10:64:48:95 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [12:04:44] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1267-1276].eqiad.wmnet [12:07:44] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12128004 (10MatthewVernon) [12:08:03] (03CR) 10TrainBranchBot: [C:03+2] "Approved by krinkle@deploy2003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310596 (https://phabricator.wikimedia.org/T349582) (owner: 10Krinkle) [12:10:36] (03CR) 10Krinkle: [C:03+2] Set $wgMathInternalRestbaseURL explicitly [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310596 (https://phabricator.wikimedia.org/T349582) (owner: 10Krinkle) [12:10:47] (03CR) 10CI reject: [V:04-1] Set $wgMathInternalRestbaseURL explicitly [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310596 (https://phabricator.wikimedia.org/T349582) (owner: 10Krinkle) [12:11:17] (03CR) 10Krinkle: Set $wgMathInternalRestbaseURL explicitly [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310596 (https://phabricator.wikimedia.org/T349582) (owner: 10Krinkle) [12:11:20] (03CR) 10Krinkle: [C:03+2] Set $wgMathInternalRestbaseURL explicitly [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310596 (https://phabricator.wikimedia.org/T349582) (owner: 10Krinkle) [12:11:31] (03CR) 10CI reject: [V:04-1] Set $wgMathInternalRestbaseURL explicitly [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310596 (https://phabricator.wikimedia.org/T349582) (owner: 10Krinkle) [12:11:36] (03PS1) 10Gmodena: wdqs: roll back wdqs-proxy to v0.3.1 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311438 (https://phabricator.wikimedia.org/T432336) [12:11:45] (03CR) 10CI reject: [V:04-1] wdqs: roll back wdqs-proxy to v0.3.1 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311438 (https://phabricator.wikimedia.org/T432336) (owner: 10Gmodena) [12:11:52] (03PS2) 10Btullis: logstash: convert opensearch server logs to ECS [puppet] - 10https://gerrit.wikimedia.org/r/1311422 (https://phabricator.wikimedia.org/T324335) [12:11:52] (03PS2) 10Btullis: rsyslog: support imfile readTimeout in rsyslog::input::file [puppet] - 10https://gerrit.wikimedia.org/r/1311423 (https://phabricator.wikimedia.org/T324335) [12:11:52] (03PS2) 10Btullis: opensearch: ship cirrus server JSON logs via rsyslog imfile [puppet] - 10https://gerrit.wikimedia.org/r/1311424 (https://phabricator.wikimedia.org/T324335) [12:11:53] (03PS2) 10Btullis: cirrus: remove the log4j syslog appender log shipping path [puppet] - 10https://gerrit.wikimedia.org/r/1311426 (https://phabricator.wikimedia.org/T324335) [12:11:54] (03PS1) 10Btullis: cirrus: enable server JSON log shipping on relforge [puppet] - 10https://gerrit.wikimedia.org/r/1311439 (https://phabricator.wikimedia.org/T324335) [12:11:56] (03PS1) 10Btullis: cirrus: enable server JSON log shipping on cloudelastic [puppet] - 10https://gerrit.wikimedia.org/r/1311440 (https://phabricator.wikimedia.org/T324335) [12:12:04] (03PS1) 10Btullis: cirrus: ship server JSON logs by default [puppet] - 10https://gerrit.wikimedia.org/r/1311441 (https://phabricator.wikimedia.org/T324335) [12:13:45] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1267-1276].eqiad.wmnet [12:13:48] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1267-1276].eqiad.wmnet [12:14:08] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1277-1286].eqiad.wmnet [12:14:11] (03CR) 10Santiago Faci: [C:03+2] "recheck" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311413 (https://phabricator.wikimedia.org/T431015) (owner: 10Santiago Faci) [12:14:12] (03PS2) 10Gmodena: wdqs: roll back wdqs-proxy to v0.3.1 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311438 (https://phabricator.wikimedia.org/T432336) [12:14:27] (03CR) 10CI reject: [V:04-1] logstash: convert opensearch server logs to ECS [puppet] - 10https://gerrit.wikimedia.org/r/1311422 (https://phabricator.wikimedia.org/T324335) (owner: 10Btullis) [12:14:44] (03CR) 10Santiago Faci: [C:03+2] "recheck" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311414 (https://phabricator.wikimedia.org/T431015) (owner: 10Santiago Faci) [12:15:43] (03CR) 10CI reject: [V:04-1] Test Kitchen UI: Deploy v1.4.9 release to staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311413 (https://phabricator.wikimedia.org/T431015) (owner: 10Santiago Faci) [12:15:52] (03CR) 10CI reject: [V:04-1] wdqs: roll back wdqs-proxy to v0.3.1 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311438 (https://phabricator.wikimedia.org/T432336) (owner: 10Gmodena) [12:15:55] (03CR) 10Krinkle: "recheck" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310596 (https://phabricator.wikimedia.org/T349582) (owner: 10Krinkle) [12:15:58] (03CR) 10CI reject: [V:04-1] Test Kitchen UI: Deploy v1.4.9 release to production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311414 (https://phabricator.wikimedia.org/T431015) (owner: 10Santiago Faci) [12:16:22] (03CR) 10Santiago Faci: [C:03+2] "recheck" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311414 (https://phabricator.wikimedia.org/T431015) (owner: 10Santiago Faci) [12:16:37] (03CR) 10CI reject: [V:04-1] Test Kitchen UI: Deploy v1.4.9 release to production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311414 (https://phabricator.wikimedia.org/T431015) (owner: 10Santiago Faci) [12:18:41] !log sukhe@cumin1003 cookbooks.sre.dns.roll-reboot begin reboot of dns6001.wikimedia.org [12:20:09] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1277-1286].eqiad.wmnet [12:20:39] FIRING: CertAlmostExpired: gNMI TLS certificate for lsw1-f6-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [12:22:40] PROBLEM - BFD status on asw1-b12-drmrs.mgmt is CRITICAL: Down: 2 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [12:23:42] RECOVERY - BFD status on asw1-b12-drmrs.mgmt is OK: UP: 7 AdminDown: 0 Down: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [12:24:07] (03CR) 10Trueg: [C:03+1] wdqs: roll back wdqs-proxy to v0.3.1 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311438 (https://phabricator.wikimedia.org/T432336) (owner: 10Gmodena) [12:24:11] FIRING: [12x] BFDdown: BFD session down between asw1-b12-drmrs and 185.15.58.5 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [12:24:15] (03PS2) 10Gmodena: wdqs: bump init container memory and startup probe timeout [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311429 (https://phabricator.wikimedia.org/T429150) [12:24:18] (03PS1) 10Jelto: service::catalog: Set ipip_encapsulation for wikikube codfw [puppet] - 10https://gerrit.wikimedia.org/r/1311444 (https://phabricator.wikimedia.org/T420436) [12:24:37] (03CR) 10CI reject: [V:04-1] wdqs: bump init container memory and startup probe timeout [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311429 (https://phabricator.wikimedia.org/T429150) (owner: 10Gmodena) [12:24:50] (03CR) 10Trueg: [C:03+1] wdqs: bump init container memory and startup probe timeout [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311429 (https://phabricator.wikimedia.org/T429150) (owner: 10Gmodena) [12:26:47] (03CR) 10Jelto: "@jmeybohm@wikimedia.org do we need to separate the rollout between eqiad and codfw? This change contains just codfw. Also does the cookboo" [puppet] - 10https://gerrit.wikimedia.org/r/1311444 (https://phabricator.wikimedia.org/T420436) (owner: 10Jelto) [12:27:07] (03PS2) 10Jelto: service::catalog: Set ipip_encapsulation for wikikube codfw [puppet] - 10https://gerrit.wikimedia.org/r/1311444 (https://phabricator.wikimedia.org/T420436) [12:27:43] 06SRE, 06cloud-services-team, 06Data-Persistence, 06Infrastructure-Foundations, and 6 others: codfw: rack B7 maintenance - https://phabricator.wikimedia.org/T430928#12128073 (10cmooney) [12:28:03] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1277-1286].eqiad.wmnet [12:28:05] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1277-1286].eqiad.wmnet [12:28:25] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet [12:29:08] (03CR) 10Krinkle: "recheck" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310596 (https://phabricator.wikimedia.org/T349582) (owner: 10Krinkle) [12:29:11] FIRING: [12x] BFDdown: BFD session down between asw1-b12-drmrs and 185.15.58.5 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [12:30:09] (03PS1) 10Ayounsi: Add BGP sharding support [homer/public] - 10https://gerrit.wikimedia.org/r/1311448 (https://phabricator.wikimedia.org/T320264) [12:30:59] (03CR) 10CI reject: [V:04-1] Localisation updates from https://translatewiki.net. [phabricator/translations] (wmf/stable) - 10https://gerrit.wikimedia.org/r/1311443 (owner: 10L10n-bot) [12:31:43] (03CR) 10CI reject: [V:04-1] Add BGP sharding support [homer/public] - 10https://gerrit.wikimedia.org/r/1311448 (https://phabricator.wikimedia.org/T320264) (owner: 10Ayounsi) [12:33:50] (03CR) 10Elukey: [C:03+1] management.yaml: Add python3-pydantic package to cumin [puppet] - 10https://gerrit.wikimedia.org/r/1311427 (https://phabricator.wikimedia.org/T430769) (owner: 10Federico Ceratto) [12:34:38] !log sukhe@cumin1003 cookbooks.sre.dns.roll-reboot finished rebooting dns6001.wikimedia.org [12:35:00] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet [12:36:07] (03CR) 10Federico Ceratto: [C:03+2] management.yaml: Add python3-pydantic package to cumin [puppet] - 10https://gerrit.wikimedia.org/r/1311427 (https://phabricator.wikimedia.org/T430769) (owner: 10Federico Ceratto) [12:37:34] (03PS2) 10Ayounsi: Add BGP sharding support [homer/public] - 10https://gerrit.wikimedia.org/r/1311448 (https://phabricator.wikimedia.org/T320264) [12:39:37] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp1103.eqiad.wmnet [12:42:00] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp1102.eqiad.wmnet [12:42:10] kamila@cumin1003 renumber-node (PID 177049) is awaiting input [12:43:10] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp5026.eqsin.wmnet [12:43:28] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp5018.eqsin.wmnet [12:43:44] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp3067.esams.wmnet [12:43:57] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp3075.esams.wmnet [12:45:14] (03CR) 10Cathal Mooney: [C:03+1] "LGTM. I'd suggest we do a brief test in containerlab before merging to make sure the show output looks like we expect it to." [homer/public] - 10https://gerrit.wikimedia.org/r/1311448 (https://phabricator.wikimedia.org/T320264) (owner: 10Ayounsi) [12:46:04] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet [12:46:07] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet [12:46:32] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1298-1307].eqiad.wmnet [12:49:38] !log sukhe@cumin1003 cookbooks.sre.dns.roll-reboot begin reboot of dns6002.wikimedia.org [12:52:18] kamila@cumin1003 renumber-node (PID 177049) is awaiting input [12:52:37] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1298-1307].eqiad.wmnet [12:53:42] PROBLEM - BFD status on asw1-b13-drmrs.mgmt is CRITICAL: Down: 2 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [12:54:11] FIRING: [12x] BFDdown: BFD session down between asw1-b13-drmrs and 185.15.58.37 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [12:54:42] RECOVERY - BFD status on asw1-b13-drmrs.mgmt is OK: UP: 6 AdminDown: 0 Down: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [12:59:11] FIRING: [12x] BFDdown: BFD session down between asw1-b13-drmrs and 185.15.58.37 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [12:59:17] (03CR) 10Bking: [C:03+2] stat hosts: Set up S3 config for wikidata platform team. [puppet] - 10https://gerrit.wikimedia.org/r/1310643 (https://phabricator.wikimedia.org/T432167) (owner: 10Bking) [12:59:42] !log kamila@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1262.eqiad.wmnet [12:59:43] !log kamila@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1262.eqiad.wmnet [12:59:45] !log kamila@cumin1003 END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1262.eqiad.wmnet [13:00:05] urbanecm and TheresNoTime: UTC afternoon backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260716T1300). Please do the needful. [13:00:05] stephanebisson: A patch you scheduled for UTC afternoon backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [13:02:58] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1298-1307].eqiad.wmnet [13:03:01] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1298-1307].eqiad.wmnet [13:03:21] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1308-1317].eqiad.wmnet [13:03:32] !log sukhe@cumin1003 cookbooks.sre.dns.roll-reboot finished rebooting dns6002.wikimedia.org [13:03:59] (03CR) 10TrainBranchBot: [C:03+2] "Approved by sbisson@deploy2003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311054 (https://phabricator.wikimedia.org/T432137) (owner: 10Sbisson) [13:04:59] (03Merged) 10jenkins-bot: Enable Article Guidance on Polish Wikipedia [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311054 (https://phabricator.wikimedia.org/T432137) (owner: 10Sbisson) [13:05:46] !log sbisson@deploy2003 Started scap sync-world: Backport for [[gerrit:1311054|Enable Article Guidance on Polish Wikipedia (T432137)]] [13:05:49] T432137: Enable AG workflow in Polish Wikipedia - https://phabricator.wikimedia.org/T432137 [13:07:45] !log sbisson@deploy2003 sbisson: Backport for [[gerrit:1311054|Enable Article Guidance on Polish Wikipedia (T432137)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [13:08:54] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1308-1317].eqiad.wmnet [13:09:30] !log sbisson@deploy2003 sbisson: Continuing with deployment [13:10:37] !log kamila@cumin1003 START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1263.eqiad.wmnet [13:10:40] !log kamila@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1263.eqiad.wmnet [13:11:12] !log kamila@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1263.eqiad.wmnet [13:11:31] !log cdobbins@cumin2003 conftool action : set/pooled=no; selector: name=dns7002.* [13:11:40] !log kamila@cumin1003 START - Cookbook sre.hosts.reimage for host wikikube-worker1263.eqiad.wmnet with OS trixie [13:12:01] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for mbsantos - https://phabricator.wikimedia.org/T432258#12128214 (10jcrespo) My apologies, I hadn't realized you had marked already your manager as the approver. Please ignore my previous message. @Bmueller Asking politely for tha... [13:12:08] !log kamila@cumin1003 START - Cookbook sre.hosts.move-vlan for host wikikube-worker1263 [13:12:12] !log cdobbins@cumin2003 conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update [13:12:24] !log kamila@cumin1003 START - Cookbook sre.dns.netbox [13:13:06] !log cdobbins@dns1004 START - running authdns-update [13:13:49] !log sbisson@deploy2003 Finished scap sync-world: Backport for [[gerrit:1311054|Enable Article Guidance on Polish Wikipedia (T432137)]] (duration: 08m 03s) [13:13:52] T432137: Enable AG workflow in Polish Wikipedia - https://phabricator.wikimedia.org/T432137 [13:14:57] !log cdobbins@dns1004 END - running authdns-update [13:16:13] !log cdobbins@cumin2003 conftool action : set/pooled=yes; selector: name=dns7002.* [13:17:31] RESOLVED: Not accepting/receiving prefixes from anycast BGP peer: Device asw1-b4-magru.mgmt.magru.wmnet recovered from Not accepting/receiving prefixes from anycast BGP peer - https://alerts.wikimedia.org/?q=alertname%3DNot+accepting%2Freceiving+prefixes+from+anycast+BGP+peer [13:18:25] (03CR) 10Majavah: nftables: add support for the VRRP protocol (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1308177 (https://phabricator.wikimedia.org/T427799) (owner: 10JHathaway) [13:18:26] kamila@cumin1003 renumber-node (PID 348953) is awaiting input [13:18:32] !log sukhe@cumin1003 cookbooks.sre.dns.roll-reboot begin reboot of dns7001.wikimedia.org [13:18:59] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1308-1317].eqiad.wmnet [13:19:02] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1308-1317].eqiad.wmnet [13:19:21] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1318-1327].eqiad.wmnet [13:19:25] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp1105.eqiad.wmnet [13:20:15] !log kamila@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" [13:20:20] !log kamila@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003" [13:20:20] !log kamila@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [13:20:20] !log kamila@cumin1003 START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [13:20:40] !log kamila@cumin1003 END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [13:21:31] !log kamila@cumin1003 START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [13:21:51] !log kamila@cumin1003 END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [13:22:02] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp1104.eqiad.wmnet [13:22:22] PROBLEM - BFD status on asw1-b3-magru.mgmt is CRITICAL: Down: 2 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [13:23:53] !log kamila@cumin1003 START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [13:23:57] !log kamila@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [13:23:57] !log kamila@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1263 [13:24:11] FIRING: [12x] BFDdown: BFD session down between asw1-b3-magru and 195.200.68.4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [13:24:22] RECOVERY - BFD status on asw1-b3-magru.mgmt is OK: UP: 2 AdminDown: 0 Down: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [13:24:24] PROBLEM - check if authdns-update was run after a change was merged to operations/dns.git on dns7002 is CRITICAL: Local zone files are NOT in sync with operations/dns.git (SHA: local is , dns.git is 508411e3df7d829a07600fac3eab6739321f5836) https://wikitech.wikimedia.org/wiki/DNS%23authdns_update_run [13:24:38] !log kamila@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1263 [13:24:38] !log kamila@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1263 [13:24:47] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1318-1327].eqiad.wmnet [13:24:58] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp5027.eqsin.wmnet [13:25:25] !log sukhe@dns1004 START - running authdns-update [13:25:31] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp3068.esams.wmnet [13:25:34] 10SRE-Access-Requests, 06Infrastructure-Foundations, 10LDAP-Access-Requests, 13Patch-For-Review: Superset data access request for abibendall - https://phabricator.wikimedia.org/T430938#12128243 (10jcrespo) I will ping the manger of the team so he may be able to find someone and/or clarify timing, but sadly... [13:25:42] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp3076.esams.wmnet [13:25:51] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp5019.eqsin.wmnet [13:27:12] (03CR) 10Ssingh: [C:03+1] service: add k8s-ingress-dse-postgresql TLS passthrough service [puppet] - 10https://gerrit.wikimedia.org/r/1311036 (https://phabricator.wikimedia.org/T432104) (owner: 10Btullis) [13:27:17] !log sukhe@dns1004 END - running authdns-update [13:29:11] FIRING: [12x] BFDdown: BFD session down between asw1-b3-magru and 195.200.68.4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [13:29:38] uh [13:29:44] oh ok, that's the reboot [13:33:08] !log sukhe@cumin1003 cookbooks.sre.dns.roll-reboot finished rebooting dns7001.wikimedia.org [13:33:08] !log sukhe@cumin1003 END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P{dns5003*} or P{dns7002*}) and (A:dnsbox) [13:33:53] (03PS1) 10Ayounsi: cables report: add test_old_planned_cables [software/netbox-extras] - 10https://gerrit.wikimedia.org/r/1311459 (https://phabricator.wikimedia.org/T432329) [13:34:54] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1318-1327].eqiad.wmnet [13:34:57] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1318-1327].eqiad.wmnet [13:35:16] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1328-1337].eqiad.wmnet [13:35:58] (03CR) 10CI reject: [V:04-1] cables report: add test_old_planned_cables [software/netbox-extras] - 10https://gerrit.wikimedia.org/r/1311459 (https://phabricator.wikimedia.org/T432329) (owner: 10Ayounsi) [13:36:56] (03PS2) 10Ayounsi: cables report: add test_old_planned_cables [software/netbox-extras] - 10https://gerrit.wikimedia.org/r/1311459 (https://phabricator.wikimedia.org/T432329) [13:37:30] PROBLEM - SSH on arclamp2001 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/SSH/monitoring [13:38:21] 06SRE, 10SRE-Access-Requests: Requesting access to ml-lab-users for ebernardson - https://phabricator.wikimedia.org/T432345 (10dcausse) 03NEW [13:39:08] RECOVERY - Check if ntpsec.service has been restarted after /etc/ntpsec/ntp.conf was changed on dns7002 is OK: OK: ntpsec.service was restarted after /etc/ntpsec/ntp.conf was changed. https://wikitech.wikimedia.org/wiki/NTP%23Monitoring [13:39:09] 06SRE, 10SRE-Access-Requests: Requesting access to ml-lab-users for dcausse - https://phabricator.wikimedia.org/T432347 (10dcausse) 03NEW [13:39:15] FIRING: [4x] MediaWikiHighErrorRate: Elevated rate of MediaWiki errors - kube-mw-api-ext - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://grafana.wikimedia.org/d/000000438/mediawiki-exceptions-alerts?panelId=18&fullscreen&orgId=1&var-datasource=codfw%20prometheus/ops - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiHighErrorRate [13:39:22] RECOVERY - SSH on arclamp2001 is OK: SSH OK - OpenSSH_8.4p1 Debian-5+deb11u7 (protocol 2.0) https://wikitech.wikimedia.org/wiki/SSH/monitoring [13:40:12] !log jiji@cumin1003 END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-eqiad [13:40:57] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1328-1337].eqiad.wmnet [13:41:10] 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 10netops, 13Patch-For-Review: Netbox report: alert on cables in 'planned' state for more than two months - https://phabricator.wikimedia.org/T432329#12128355 (10ayounsi) Above patch tested in Netbox-next: {F94004219} [13:41:49] 06SRE, 10SRE-Access-Requests: Requesting access to ml-lab-users for ebernhardson - https://phabricator.wikimedia.org/T432345#12128356 (10dcausse) [13:43:23] (03CR) 10Santiago Faci: [C:03+2] "recheck" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311413 (https://phabricator.wikimedia.org/T431015) (owner: 10Santiago Faci) [13:44:15] RESOLVED: [8x] MediaWikiHighErrorRate: Elevated rate of MediaWiki errors - kube-mw-api-ext - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiHighErrorRate [13:45:04] RECOVERY - NTP peers and stratum check on dns7002 is OK: NTP OK: Offset -0.00020385 secs, stratum=2 https://wikitech.wikimedia.org/wiki/NTP [13:45:44] (03Merged) 10jenkins-bot: Test Kitchen UI: Deploy v1.4.9 release to staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311413 (https://phabricator.wikimedia.org/T431015) (owner: 10Santiago Faci) [13:45:51] !log kamila@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage [13:48:44] (03CR) 10Santiago Faci: [C:03+2] "recheck" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311414 (https://phabricator.wikimedia.org/T431015) (owner: 10Santiago Faci) [13:49:20] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1328-1337].eqiad.wmnet [13:49:23] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1328-1337].eqiad.wmnet [13:49:33] !log sfaci@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply [13:49:38] !log jiji@cumin1003 START - Cookbook sre.hosts.reboot-single for host mc-misc1001.eqiad.wmnet [13:49:43] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1338-1347].eqiad.wmnet [13:50:14] !log sfaci@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply [13:50:25] !log kamila@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage [13:50:49] (03Merged) 10jenkins-bot: Test Kitchen UI: Deploy v1.4.9 release to production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311414 (https://phabricator.wikimedia.org/T431015) (owner: 10Santiago Faci) [13:51:44] (03PS3) 10CDobbins: varnish: update CSP report-only header [puppet] - 10https://gerrit.wikimedia.org/r/1311069 [13:52:38] (03CR) 10CDobbins: "Done" [puppet] - 10https://gerrit.wikimedia.org/r/1311069 (owner: 10CDobbins) [13:54:45] !log sfaci@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply [13:54:54] !log jiji@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1001.eqiad.wmnet [13:54:58] !log jiji@cumin1003 START - Cookbook sre.hosts.reboot-single for host mc-misc1002.eqiad.wmnet [13:55:48] !log sfaci@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply [13:56:33] !log jiji@cumin1003 END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw [13:59:17] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp1107.eqiad.wmnet [13:59:34] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1338-1347].eqiad.wmnet [13:59:43] (03PS1) 10Lerickson: Update default Eventgate URL in the wdqs-proxy config. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311460 (https://phabricator.wikimedia.org/T432336) [14:00:02] !log jiji@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1002.eqiad.wmnet [14:01:29] (03PS21) 10Federico Ceratto: mysql: update replication source [cookbooks] - 10https://gerrit.wikimedia.org/r/1238368 (https://phabricator.wikimedia.org/T373436) [14:02:06] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp1106.eqiad.wmnet [14:02:07] !log beginning depools for lsw1-b6-codfw maintenance T430922 [14:02:11] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:02:11] T430922: codfw: rack B6 maintenance - https://phabricator.wikimedia.org/T430922 [14:05:34] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp3077.esams.wmnet [14:06:54] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1338-1347].eqiad.wmnet [14:06:57] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1338-1347].eqiad.wmnet [14:07:10] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp5028.eqsin.wmnet [14:07:17] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1348-1357].eqiad.wmnet [14:07:18] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp3069.esams.wmnet [14:07:25] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for mbsantos - https://phabricator.wikimedia.org/T432258#12128435 (10Bmueller) Approved! [14:07:32] !log btullis@cumin1003 START - Cookbook sre.hadoop.reboot-workers for Hadoop analytics cluster [14:07:56] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp5020.eqsin.wmnet [14:08:03] !log cmooney@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet [14:08:42] btullis@cumin1003 roll-restart-zookeeper (PID 393335) is awaiting input [14:10:57] (03CR) 10Federico Ceratto: mysql: update replication source (031 comment) [cookbooks] - 10https://gerrit.wikimedia.org/r/1238368 (https://phabricator.wikimedia.org/T373436) (owner: 10Federico Ceratto) [14:11:37] !log cmooney@cumin1003 START - Cookbook sre.mysql.depool depool db2161: switch maintenance codfw rack b6 [14:11:58] !log cmooney@cumin1003 END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2161: switch maintenance codfw rack b6 [14:12:03] !log cmooney@cumin1003 START - Cookbook sre.mysql.depool depool db2162: switch maintenance codfw rack b6 [14:12:18] !log kamila@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1263.eqiad.wmnet with OS trixie [14:12:33] !log cmooney@cumin1003 END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2162: switch maintenance codfw rack b6 [14:12:39] !log cmooney@cumin1003 START - Cookbook sre.mysql.depool depool db2251: switch maintenance codfw rack b6 [14:12:40] !log cmooney@cumin1003 START - Cookbook sre.mysql.parsercache [14:12:48] !log cmooney@cumin1003 END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) [14:12:48] !log cmooney@cumin1003 END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: switch maintenance codfw rack b6 [14:12:54] !log cmooney@cumin1003 START - Cookbook sre.mysql.depool depool pc2022: switch maintenance codfw rack b6 [14:12:54] !log cmooney@cumin1003 START - Cookbook sre.mysql.parsercache [14:12:55] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1348-1357].eqiad.wmnet [14:13:04] !log cmooney@cumin1003 END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) [14:13:04] !log cmooney@cumin1003 END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2022: switch maintenance codfw rack b6 [14:13:09] !log cmooney@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet [14:13:11] !log cmooney@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet [14:13:24] 06SRE, 10SRE-Access-Requests, 06Data-Platform-SRE (2026-07-03 - 2026-07-31): Grant Access to analytics-privatedata-users for zsinger - https://phabricator.wikimedia.org/T426458#12128464 (10BTullis) 05Resolved→03Open a:05SLyngshede-WMF→03BTullis Reopening this task, since the membership of `analytics-... [14:14:16] (03PS1) 10Fabfur: admin: add mbsantos to analytics-privatedata-users [puppet] - 10https://gerrit.wikimedia.org/r/1311462 (https://phabricator.wikimedia.org/T432258) [14:14:43] !log cmooney@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet [14:14:50] (03CR) 10Scott French: [C:03+1] "Thanks, Blake!" [dns] - 10https://gerrit.wikimedia.org/r/1311074 (https://phabricator.wikimedia.org/T427668) (owner: 10Blake) [14:15:44] (03CR) 10Scott French: [C:03+1] trafficserver: Remove XWD routing for /w/rest.php mw-debug backend [puppet] - 10https://gerrit.wikimedia.org/r/1310103 (https://phabricator.wikimedia.org/T428909) (owner: 10Effie Mouzeli) [14:16:30] I'd like to steal the deployment servers for a quick test if y'all don't mind [14:16:42] I'll give them back in ~15min [14:16:57] (03PS1) 10Btullis: admin_ng: add the postgresql-superset-metrics namespace on dse-k8s-eqiad [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311463 (https://phabricator.wikimedia.org/T432104) [14:18:32] (03PS1) 10Effie Mouzeli: ProductionServices: reboot poolcounter1006 (#1/4) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311464 (https://phabricator.wikimedia.org/T431705) [14:18:33] (03PS1) 10Effie Mouzeli: ProductionServices: reboot poolcounter1007 (#2/4) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311465 (https://phabricator.wikimedia.org/T431705) [14:18:34] (03PS1) 10Effie Mouzeli: ProductionServices: reboot poolcounter2005 (#3/4) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311466 (https://phabricator.wikimedia.org/T431705) [14:18:35] (03PS1) 10Effie Mouzeli: ProductionServices: reboot poolcounter2006 (#4/4) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311467 (https://phabricator.wikimedia.org/T431705) [14:18:35] (03PS1) 10Effie Mouzeli: ProductionServices: repool poolcounter2006 after reboot [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311468 (https://phabricator.wikimedia.org/T431705) [14:19:09] !log cmooney@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: lsw1-b6-codfw JunOS upgrade [14:19:42] !log jiji@cumin1003 START - Cookbook sre.hosts.reboot-single for host mc-wf1001.eqiad.wmnet [14:20:09] !log jiji@cumin1003 START - Cookbook sre.hosts.reboot-single for host mc-wf2001.codfw.wmnet [14:20:21] (03PS1) 10Btullis: cloudnative-pg-cluster: allow configuring a node maintenance window [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311469 (https://phabricator.wikimedia.org/T432104) [14:20:28] !log cmooney@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b6-codfw,lsw1-b6-codfw IPv6,lsw1-b6-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b6-codfw JunOS upgrade [14:21:08] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host wcqs1001.eqiad.wmnet with OS bookworm [14:21:08] !log kamila@deploy2003 Started scap sync-world: Test deployment to check rsync is working - T432108 [14:21:14] T432108: stunnel4 keeps not exiting cleanly on Bookworm(?) - https://phabricator.wikimedia.org/T432108 [14:21:23] !log reboot lsw1-b6-codfw to upgrade JunOS T430922 [14:21:26] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:21:26] T430922: codfw: rack B6 maintenance - https://phabricator.wikimedia.org/T430922 [14:21:44] !log kamila@deploy2003 sync-world aborted: Test deployment to check rsync is working - T432108 (duration: 00m 36s) [14:22:30] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1348-1357].eqiad.wmnet [14:22:33] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1348-1357].eqiad.wmnet [14:22:52] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1358-1366].eqiad.wmnet [14:24:28] (03PS1) 10Majavah: hieradata: Add new VXLAN metricsinfra scraping addresses [puppet] - 10https://gerrit.wikimedia.org/r/1311470 (https://phabricator.wikimedia.org/T401813) [14:24:31] (03PS1) 10Majavah: openstack: wmcs-securitygroup-backfill: Add support for IPv6 [puppet] - 10https://gerrit.wikimedia.org/r/1311471 (https://phabricator.wikimedia.org/T401813) [14:24:58] !log kamila@deploy2003 Started scap sync-world: Test deployment to check rsync is working - T432108 [14:25:34] !log jiji@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1001.eqiad.wmnet [14:25:38] !log jiji@cumin1003 START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet [14:25:42] 10ops-eqiad, 06SRE, 06Data-Platform-SRE, 06DC-Ops, 06Machine-Learning-Team: hw troubleshooting: failing disk in ml-serve1001.eqiad.wmnet - https://phabricator.wikimedia.org/T432105#12128533 (10VRiley-WMF) CODFW is sending replacment disks. Thanks @Jhancock.wm [14:26:02] !log jiji@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2001.codfw.wmnet [14:26:06] !log jiji@cumin1003 START - Cookbook sre.hosts.reboot-single for host mc-wf2002.codfw.wmnet [14:26:49] (03PS1) 10Aude: Preserve menus after toolbox (e.g. print/export) in page tools [skins/Vector] (wmf/1.47.0-wmf.11) - 10https://gerrit.wikimedia.org/r/1311472 (https://phabricator.wikimedia.org/T432316) [14:27:19] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Thursday, July 16 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-ite" [skins/Vector] (wmf/1.47.0-wmf.11) - 10https://gerrit.wikimedia.org/r/1311472 (https://phabricator.wikimedia.org/T432316) (owner: 10Aude) [14:27:25] (03CR) 10Jcrespo: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1311462 (https://phabricator.wikimedia.org/T432258) (owner: 10Fabfur) [14:27:44] !log btullis@cumin1003 START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. [14:27:55] !log kamila@deploy2003 Finished scap sync-world: Test deployment to check rsync is working - T432108 (duration: 02m 57s) [14:27:58] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1358-1366].eqiad.wmnet [14:27:59] T432108: stunnel4 keeps not exiting cleanly on Bookworm(?) - https://phabricator.wikimedia.org/T432108 [14:29:16] !log cmooney@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet [14:29:51] (03CR) 10Scott French: [C:03+1] kubernetes: Add a debug deployment for mw-pretrain. (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1311049 (https://phabricator.wikimedia.org/T427668) (owner: 10Blake) [14:30:04] Deploy window Test Kitchen Experiment Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260716T1430) [14:30:07] (03CR) 10Btullis: "recheck" [puppet] - 10https://gerrit.wikimedia.org/r/1311422 (https://phabricator.wikimedia.org/T324335) (owner: 10Btullis) [14:30:22] (03CR) 10Jcrespo: "Actually, I am not sure kerberos was asked, given it is for web requests only." [puppet] - 10https://gerrit.wikimedia.org/r/1311462 (https://phabricator.wikimedia.org/T432258) (owner: 10Fabfur) [14:31:23] !log jiji@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet [14:32:03] !log jiji@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2002.codfw.wmnet [14:33:00] (03CR) 10Jcrespo: "Request says: "Superset dashboard", which I understand as web access only. I understand that as level 1 here: https://wikitech.wikimedia.o" [puppet] - 10https://gerrit.wikimedia.org/r/1311462 (https://phabricator.wikimedia.org/T432258) (owner: 10Fabfur) [14:33:33] (03PS5) 10Kamila Součková: modules/rsync: add pidfile to stunnel.conf [puppet] - 10https://gerrit.wikimedia.org/r/1310589 (https://phabricator.wikimedia.org/T432108) [14:33:52] (03CR) 10Jcrespo: "https://wikitech.wikimedia.org/wiki/Data_Platform/Data_access#Access_Levels" [puppet] - 10https://gerrit.wikimedia.org/r/1311462 (https://phabricator.wikimedia.org/T432258) (owner: 10Fabfur) [14:34:05] !log btullis@cumin1003 END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons. [14:34:24] !log cmooney@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet [14:34:53] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1358-1366].eqiad.wmnet [14:34:56] (03CR) 10Kamila Součková: "Very good idea, as it turns out that the config file needed a slight change to actually work 😄 Thank you!" [puppet] - 10https://gerrit.wikimedia.org/r/1310589 (https://phabricator.wikimedia.org/T432108) (owner: 10Kamila Součková) [14:34:56] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1358-1366].eqiad.wmnet [14:35:02] (03CR) 10JavierMonton: "recheck" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311418 (https://phabricator.wikimedia.org/T432204) (owner: 10JavierMonton) [14:35:13] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1367-1375].eqiad.wmnet [14:35:17] (03PS4) 10Aqu: airflow: add gcs-token-search-console secret [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309204 (https://phabricator.wikimedia.org/T427457) [14:35:44] PROBLEM - BFD status on ssw1-a8-codfw.mgmt is CRITICAL: Down: 1 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [14:35:52] PROBLEM - BFD status on ssw1-a1-codfw.mgmt is CRITICAL: Down: 1 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [14:35:58] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [14:36:25] FIRING: [16x] ProbeDown: Service aqs2005-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [14:36:26] (03CR) 10Kamila Součková: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1310589 (https://phabricator.wikimedia.org/T432108) (owner: 10Kamila Součková) [14:36:30] PROBLEM - MariaDB Replica IO: ms1 on db1152 is CRITICAL: CRITICAL slave_io_state Slave_IO_Running: No, Errno: 2003, Errmsg: error reconnecting to master repl2024@db2251.codfw.wmnet:3306 - retry-time: 60 maximum-retries: 100000 message: Cant connect to server on db2251.codfw.wmnet (110 Connection timed out) https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [14:36:44] (03CR) 10Majavah: [V:03+1] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9023/co" [puppet] - 10https://gerrit.wikimedia.org/r/1311470 (https://phabricator.wikimedia.org/T401813) (owner: 10Majavah) [14:38:21] (03CR) 10Trueg: [C:03+1] Update default Eventgate URL in the wdqs-proxy config. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311460 (https://phabricator.wikimedia.org/T432336) (owner: 10Lerickson) [14:39:14] (03PS1) 10Btullis: Add zsinger to the analytics-privatedata-users group [puppet] - 10https://gerrit.wikimedia.org/r/1311476 (https://phabricator.wikimedia.org/T426458) [14:40:04] (03CR) 10CI reject: [V:04-1] Add zsinger to the analytics-privatedata-users group [puppet] - 10https://gerrit.wikimedia.org/r/1311476 (https://phabricator.wikimedia.org/T426458) (owner: 10Btullis) [14:40:20] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1367-1375].eqiad.wmnet [14:40:41] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp1109.eqiad.wmnet [14:41:25] FIRING: [22x] ProbeDown: Service aqs2005-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [14:41:27] (03PS2) 10Fabfur: admin: add mbsantos to analytics-privatedata-users [puppet] - 10https://gerrit.wikimedia.org/r/1311462 (https://phabricator.wikimedia.org/T432258) [14:41:45] FIRING: CirrusStreamingUpdaterRateTooLow: CirrusSearch update rate from flink-app-consumer-cloudelastic is critically low - https://wikitech.wikimedia.org/wiki/Search#Streaming_Updater - https://grafana.wikimedia.org/d/jKqki4MSk/cirrus-streaming-updater - https://alerts.wikimedia.org/?q=alertname%3DCirrusStreamingUpdaterRateTooLow [14:42:01] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp1108.eqiad.wmnet [14:42:08] (03CR) 10Fabfur: "Thanks, amended the patch for this, eventually other permissions can be granted later if needed" [puppet] - 10https://gerrit.wikimedia.org/r/1311462 (https://phabricator.wikimedia.org/T432258) (owner: 10Fabfur) [14:42:42] (03CR) 10CWilliams: mysql: update replication source (031 comment) [cookbooks] - 10https://gerrit.wikimedia.org/r/1238368 (https://phabricator.wikimedia.org/T373436) (owner: 10Federico Ceratto) [14:43:30] PROBLEM - MariaDB Replica Lag: ms1 on db1152 is CRITICAL: CRITICAL slave_sql_lag Replication lag: 605.77 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [14:44:09] (03CR) 10Jcrespo: [C:03+1] admin: add mbsantos to analytics-privatedata-users [puppet] - 10https://gerrit.wikimedia.org/r/1311462 (https://phabricator.wikimedia.org/T432258) (owner: 10Fabfur) [14:44:27] (03CR) 10Fabfur: [C:03+2] admin: add mbsantos to analytics-privatedata-users [puppet] - 10https://gerrit.wikimedia.org/r/1311462 (https://phabricator.wikimedia.org/T432258) (owner: 10Fabfur) [14:44:49] !log kamila@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1263.eqiad.wmnet [14:44:50] !log kamila@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1263.eqiad.wmnet [14:44:53] !log kamila@cumin1003 END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1263.eqiad.wmnet [14:46:30] RECOVERY - MariaDB Replica IO: ms1 on db1152 is OK: OK slave_io_state Slave_IO_Running: Yes https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [14:46:30] RECOVERY - MariaDB Replica Lag: ms1 on db1152 is OK: OK slave_sql_lag Replication lag: 0.00 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [14:46:45] FIRING: [2x] CirrusStreamingUpdaterRateTooLow: CirrusSearch update rate from flink-app-consumer-cloudelastic is critically low - https://wikitech.wikimedia.org/wiki/Search#Streaming_Updater - https://grafana.wikimedia.org/d/jKqki4MSk/cirrus-streaming-updater - https://alerts.wikimedia.org/?q=alertname%3DCirrusStreamingUpdaterRateTooLow [14:46:46] RECOVERY - BFD status on ssw1-a8-codfw.mgmt is OK: UP: 17 AdminDown: 0 Down: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [14:46:48] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to analytics-privatedata-users for mbsantos - https://phabricator.wikimedia.org/T432258#12128623 (10Fabfur) 05Open→03Resolved a:03Fabfur Hi @Bmueller , @MSantos , the change has been deployed, user should have access soon to the se... [14:46:52] RECOVERY - BFD status on ssw1-a1-codfw.mgmt is OK: UP: 17 AdminDown: 0 Down: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [14:47:04] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to analytics-privatedata-users for mbsantos - https://phabricator.wikimedia.org/T432258#12128627 (10Fabfur) [14:47:22] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp3078.esams.wmnet [14:47:34] (03CR) 10JavierMonton: [C:03+2] stream: pageview-trending-relative + webrequest-page-view [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311418 (https://phabricator.wikimedia.org/T432204) (owner: 10JavierMonton) [14:47:39] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1367-1375].eqiad.wmnet [14:47:42] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1367-1375].eqiad.wmnet [14:47:45] FIRING: [2x] CirrusStreamingUpdaterClearWeightedTagsTooLow: CirrusSearch consumer-cloudelastic@eqiad is clearing too few weighted tags - https://wikitech.wikimedia.org/wiki/Search#Streaming_Updater - https://alerts.wikimedia.org/?q=alertname%3DCirrusStreamingUpdaterClearWeightedTagsTooLow [14:47:48] !log kamila@cumin1003 START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet [14:47:52] !log kamila@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet [14:47:56] !log kamila@cumin1003 START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1265.eqiad.wmnet [14:48:00] !log kamila@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1265.eqiad.wmnet [14:48:02] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1376-1384].eqiad.wmnet [14:48:19] !log bking@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage [14:48:24] !log kamila@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet [14:49:06] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp3070.esams.wmnet [14:49:07] !log kamila@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1265.eqiad.wmnet [14:49:29] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp5029.eqsin.wmnet [14:49:39] !log cmooney@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet [14:49:41] !log kamila@cumin1003 START - Cookbook sre.hosts.reimage for host wikikube-worker1265.eqiad.wmnet with OS trixie [14:49:47] !log kamila@cumin1003 START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie [14:49:56] !log cmooney@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet [14:50:11] !log kamila@cumin1003 START - Cookbook sre.hosts.move-vlan for host wikikube-worker1265 [14:50:18] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp5021.eqsin.wmnet [14:50:32] !log cmooney@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet [14:50:37] !log cmooney@cumin1003 START - Cookbook sre.mysql.pool pool db2161: switch maintenance completed codfw rack b6 [14:50:40] !log cmooney@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet [14:50:44] PROBLEM - Host an-worker1191 is DOWN: PING CRITICAL - Packet loss = 100% [14:50:52] !log cmooney@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet [14:50:54] !log cmooney@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet [14:51:12] (03CR) 10Kamila Součková: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1310589 (https://phabricator.wikimedia.org/T432108) (owner: 10Kamila Součková) [14:51:25] RESOLVED: [22x] ProbeDown: Service aqs2005-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [14:51:33] !log kamila@cumin1003 START - Cookbook sre.dns.netbox [14:51:45] RESOLVED: [3x] CirrusStreamingUpdaterRateTooLow: CirrusSearch update rate from flink-app-consumer-cloudelastic is critically low - https://wikitech.wikimedia.org/wiki/Search#Streaming_Updater - https://grafana.wikimedia.org/d/jKqki4MSk/cirrus-streaming-updater - https://alerts.wikimedia.org/?q=alertname%3DCirrusStreamingUpdaterRateTooLow [14:51:49] (03PS2) 10Btullis: Add zsinger to the analytics-privatedata-users group [puppet] - 10https://gerrit.wikimedia.org/r/1311476 (https://phabricator.wikimedia.org/T426458) [14:52:45] RESOLVED: [2x] CirrusStreamingUpdaterClearWeightedTagsTooLow: CirrusSearch consumer-cloudelastic@eqiad is clearing too few weighted tags - https://wikitech.wikimedia.org/wiki/Search#Streaming_Updater - https://alerts.wikimedia.org/?q=alertname%3DCirrusStreamingUpdaterClearWeightedTagsTooLow [14:53:24] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage [14:53:33] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1376-1384].eqiad.wmnet [14:53:42] !log jiji@cumin1003 START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet [14:55:47] (03PS1) 10Ahmon Dancy: jenkins: enable overlay/overlayfs kernel modules for Docker [puppet] - 10https://gerrit.wikimedia.org/r/1311477 (https://phabricator.wikimedia.org/T432326) [14:55:56] (03CR) 10TrainBranchBot: [C:03+2] "Approved by krinkle@deploy2003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310596 (https://phabricator.wikimedia.org/T349582) (owner: 10Krinkle) [14:56:44] PROBLEM - OSPF status on cr2-drmrs is CRITICAL: OSPFv2: 2/4 UP : OSPFv3: 2/4 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [14:56:51] !log kamila@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" [14:56:52] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 5/7 UP : OSPFv3: 5/7 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [14:56:54] (03Merged) 10jenkins-bot: Set $wgMathInternalRestbaseURL explicitly [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310596 (https://phabricator.wikimedia.org/T349582) (owner: 10Krinkle) [14:56:56] !log kamila@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003" [14:56:56] !log kamila@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [14:56:56] !log kamila@cumin1003 START - Cookbook sre.dns.wipe-cache wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [14:57:00] !log kamila@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [14:57:00] !log kamila@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1265 [14:57:04] FIRING: HelmReleaseBadStatus: Helm release wdqs/scholarly-external on k8s-dse@codfw in state pending-rollback - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [14:57:11] !log krinkle@deploy2003 Started scap sync-world: Backport for [[gerrit:1310596|Set $wgMathInternalRestbaseURL explicitly (T349582)]] [14:57:14] T349582: Set $wgMathInternalRestbaseURL explictly in production - https://phabricator.wikimedia.org/T349582 [14:57:21] PROBLEM - Host cr1-drmrs is DOWN: PING CRITICAL - Packet loss = 100% [14:57:27] !log kamila@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1265 [14:57:28] !log kamila@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1265 [14:57:51] !log kamila@cumin1003 START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264 [14:58:07] !log kamila@cumin1003 START - Cookbook sre.dns.netbox [14:58:10] PROBLEM - Host cr1-drmrs IPv6 is DOWN: PING CRITICAL - Packet loss = 100% [14:58:10] PROBLEM - Host cr1-drmrs.mgmt is DOWN: PING CRITICAL - Packet loss = 100% [14:58:21] (03PS2) 10Ahmon Dancy: jenkins: enable overlay/overlayfs kernel modules for Docker [puppet] - 10https://gerrit.wikimedia.org/r/1311477 (https://phabricator.wikimedia.org/T432326) [14:58:25] !incidents [14:58:25] 8201 (UNACKED) Host cr1-drmrs [14:58:25] 8200 (RESOLVED) PHPFPMTooBusy sre (mw-web main eqiad) [14:58:28] !log kamila@cumin1003 START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1266.eqiad.wmnet [14:58:32] !log kamila@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1266.eqiad.wmnet [14:58:33] !ack [14:58:34] 8201 (ACKED) Host cr1-drmrs [14:58:36] FIRING: NetworkDeviceAlarmActive: Alarm active on cr1-drmrs - https://wikitech.wikimedia.org/wiki/Network_monitoring#Juniper_alarm - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-drmrs:9804 - https://alerts.wikimedia.org/?q=alertname%3DNetworkDeviceAlarmActive [14:58:37] FIRING: [9x] CoreRouterInterfaceDown: Core router interface down - cr2-eqiad:et-0/0/1 () - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [14:59:04] !log kamila@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1266.eqiad.wmnet [14:59:09] FIRING: [6x] CoreBGPDown: Core BGP session down between asw1-b12-drmrs and cr1-drmrs (185.15.58.142) - group core - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [14:59:10] ^^ topranks I assume is the b6 ongoing activity [14:59:11] FIRING: [16x] BFDdown: BFD session down between cr1-eqiad and 2620:0:861:107:10:64:48:95 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [14:59:37] Thanks for that ops/puppet merge rzl. A nice thing to see as I start my day today. [14:59:49] !log jiji@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet [15:00:00] (03PS1) 10Krinkle: Revert "Set $wgMathInternalRestbaseURL explicitly" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311478 [15:00:04] jeena and hashar: Deploy window Train log triage (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260716T1500) [15:00:10] (03CR) 10Krinkle: [C:03+2] Revert "Set $wgMathInternalRestbaseURL explicitly" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311478 (owner: 10Krinkle) [15:00:14] bd808: any time :) [15:00:27] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1376-1384].eqiad.wmnet [15:00:30] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1376-1384].eqiad.wmnet [15:00:33] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp1001.eqiad.wmnet [15:01:05] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp1001.eqiad.wmnet [15:01:07] (03Merged) 10jenkins-bot: Revert "Set $wgMathInternalRestbaseURL explicitly" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311478 (owner: 10Krinkle) [15:01:09] (03CR) 10Joal: "Almost! I think my comment are correct here, but someone with more puppet experience should confirm" [puppet] - 10https://gerrit.wikimedia.org/r/1311397 (https://phabricator.wikimedia.org/T432208) (owner: 10A-pizzata) [15:01:54] fabfur: no the alerts in Marseille are unrealted to the (completed) work in Dallas [15:02:01] !log cgoubert@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp1001.eqiad.wmnet [15:02:02] (03PS3) 10Ahmon Dancy: jenkins: enable overlay/overlayfs kernel modules for Docker [puppet] - 10https://gerrit.wikimedia.org/r/1311477 (https://phabricator.wikimedia.org/T432326) [15:02:02] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp1001.eqiad.wmnet [15:02:02] !log cgoubert@cumin1003 END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-eqiad [15:02:05] kamila@cumin1003 renumber-node (PID 427915) is awaiting input [15:02:10] pfe failed on cr1-drmrs unexpectedly (discussion in _security) [15:02:22] RESOLVED: [9x] CoreRouterInterfaceDown: Core router interface down - cr2-eqiad:et-0/0/1 () - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [15:02:23] RECOVERY - Host cr1-drmrs is UP: PING OK - Packet loss = 0%, RTA = 88.37 ms [15:02:43] (03CR) 10Ahmon Dancy: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1311477 (https://phabricator.wikimedia.org/T432326) (owner: 10Ahmon Dancy) [15:02:44] RECOVERY - OSPF status on cr2-drmrs is OK: OSPFv2: 4/4 UP : OSPFv3: 4/4 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [15:02:50] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 7/7 UP : OSPFv3: 7/7 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [15:03:02] topranks: ack, following on _security [15:03:12] RECOVERY - Host cr1-drmrs IPv6 is UP: PING OK - Packet loss = 0%, RTA = 87.46 ms [15:03:12] RECOVERY - Host cr1-drmrs.mgmt is UP: PING OK - Packet loss = 0%, RTA = 87.61 ms [15:03:13] !incidents [15:03:13] 8201 (RESOLVED) Host cr1-drmrs [15:03:14] 8200 (RESOLVED) PHPFPMTooBusy sre (mw-web main eqiad) [15:03:21] FIRING: [2x] SwitchCoreInterfaceDown: Switch core interface down - asw1-b12-drmrs:et-0/0/48 (Core: cr1-drmrs:et-0/0/1 {#D0100}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Switch_interface_down - https://alerts.wikimedia.org/?q=alertname%3DSwitchCoreInterfaceDown [15:03:36] RESOLVED: [2x] SwitchCoreInterfaceDown: Switch core interface down - asw1-b12-drmrs:et-0/0/48 (Core: cr1-drmrs:et-0/0/1 {#D0100}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Switch_interface_down - https://alerts.wikimedia.org/?q=alertname%3DSwitchCoreInterfaceDown [15:03:36] RESOLVED: NetworkDeviceAlarmActive: Alarm active on cr1-drmrs - https://wikitech.wikimedia.org/wiki/Network_monitoring#Juniper_alarm - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-drmrs:9804 - https://alerts.wikimedia.org/?q=alertname%3DNetworkDeviceAlarmActive [15:03:43] !log kamila@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" [15:03:48] !log kamila@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003" [15:03:48] !log kamila@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [15:03:48] !log kamila@cumin1003 START - Cookbook sre.dns.wipe-cache wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [15:03:52] !log kamila@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [15:03:52] !log kamila@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1264 [15:04:09] FIRING: [6x] CoreBGPDown: Core BGP session down between asw1-b12-drmrs and cr1-drmrs (185.15.58.142) - group core - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [15:04:09] !log kamila@cumin1003 START - Cookbook sre.hosts.reimage for host wikikube-worker1266.eqiad.wmnet with OS trixie [15:04:11] FIRING: [16x] BFDdown: BFD session down between cr1-eqiad and 2620:0:861:107:10:64:48:95 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [15:04:13] !log kamila@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1264 [15:04:13] !log kamila@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264 [15:04:38] !log kamila@cumin1003 START - Cookbook sre.hosts.move-vlan for host wikikube-worker1266 [15:05:24] FIRING: [2x] CertAlmostExpired: gNMI TLS certificate for cr1-drmrs.wikimedia.org is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [15:06:22] (03CR) 10Ahmon Dancy: "PCC results: https://puppet-compiler.wmflabs.org/output/1311477/7270/" [puppet] - 10https://gerrit.wikimedia.org/r/1311477 (https://phabricator.wikimedia.org/T432326) (owner: 10Ahmon Dancy) [15:07:41] kamila@cumin1003 renumber-node (PID 427915) is awaiting input [15:07:56] (03CR) 10Effie Mouzeli: [C:03+2] trafficserver: Remove XWD routing for /w/rest.php mw-debug backend [puppet] - 10https://gerrit.wikimedia.org/r/1310103 (https://phabricator.wikimedia.org/T428909) (owner: 10Effie Mouzeli) [15:07:59] !log kamila@cumin1003 START - Cookbook sre.dns.netbox [15:08:25] (03PS3) 10Blake: kubernetes: Add a debug deployment for mw-pretrain. [puppet] - 10https://gerrit.wikimedia.org/r/1311049 (https://phabricator.wikimedia.org/T427668) [15:08:36] FIRING: NetworkDeviceAlarmActive: Alarm active on cr1-drmrs - https://wikitech.wikimedia.org/wiki/Network_monitoring#Juniper_alarm - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-drmrs:9804 - https://alerts.wikimedia.org/?q=alertname%3DNetworkDeviceAlarmActive [15:08:43] (03CR) 10Blake: kubernetes: Add a debug deployment for mw-pretrain. (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1311049 (https://phabricator.wikimedia.org/T427668) (owner: 10Blake) [15:08:58] FIRING: NELHigh: Elevated Network Error Logging events (tcp.timed_out) #page - https://wikitech.wikimedia.org/wiki/Network_monitoring#NEL_alerts - https://logstash.wikimedia.org/goto/5c8f4ca1413eda33128e5c5a35da7e28 - https://alerts.wikimedia.org/?q=alertname%3DNELHigh [15:09:09] RESOLVED: [6x] CoreBGPDown: Core BGP session down between asw1-b12-drmrs and cr1-drmrs (185.15.58.142) - group core - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [15:09:11] FIRING: [16x] BFDdown: BFD session down between cr1-eqiad and 2620:0:861:107:10:64:48:95 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [15:09:21] !incidents [15:09:22] 8202 (UNACKED) NELHigh sre (thanos-rule@main tcp.timed_out) [15:09:22] 8201 (RESOLVED) Host cr1-drmrs [15:09:22] 8200 (RESOLVED) PHPFPMTooBusy sre (mw-web main eqiad) [15:09:23] !ack [15:09:23] (03CR) 10Btullis: [C:03+2] admin_ng: add the postgresql-superset-metrics namespace on dse-k8s-eqiad [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311463 (https://phabricator.wikimedia.org/T432104) (owner: 10Btullis) [15:09:24] 8202 (ACKED) NELHigh sre (thanos-rule@main tcp.timed_out) [15:13:29] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 - https://phabricator.wikimedia.org/T431682#12128739 (10VRiley-WMF) Hey @fgiunchedi thanks for your patience. As it turns out there is a lot going on here at eqiad. I know you're currently out, but I'm planning on hitti... [15:13:58] RESOLVED: NELHigh: Elevated Network Error Logging events (tcp.timed_out) #page - https://wikitech.wikimedia.org/wiki/Network_monitoring#NEL_alerts - https://logstash.wikimedia.org/goto/5c8f4ca1413eda33128e5c5a35da7e28 - https://alerts.wikimedia.org/?q=alertname%3DNELHigh [15:15:06] !log kamila@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" [15:15:10] !log kamila@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003" [15:15:10] !log kamila@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [15:15:10] !log kamila@cumin1003 START - Cookbook sre.dns.wipe-cache wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [15:15:14] !log kamila@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [15:15:15] !log kamila@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1266 [15:15:34] (03PS1) 10Btullis: kubernetes: add deploy tokens for the postgresql-superset-metrics namespace [puppet] - 10https://gerrit.wikimedia.org/r/1311480 (https://phabricator.wikimedia.org/T432104) [15:16:08] jouncebot: now [15:16:08] For the next 0 hour(s) and 43 minute(s): Train log triage (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260716T1500) [15:16:23] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1001.eqiad.wmnet with OS bookworm [15:16:47] !log kamila@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage [15:16:52] jeena: are you folks using this window? [15:16:57] 10ops-eqiad, 06SRE, 06DC-Ops: Degraded RAID on an-worker1191 - https://phabricator.wikimedia.org/T431828#12128762 (10VRiley-WMF) 05Open→03Resolved I have updated the firmware on what Dell has recommended. It seems as though this has corrected the hard drive issue, and it seems to be healthy. I'm goin... [15:17:48] ah timo is using it [15:18:30] Krinkle: please ping me when you are done, I need to do some backports [15:18:55] 10ops-eqiad, 06SRE, 06cloud-services-team, 10Cloud-VPS, and 2 others: New thermal paste for cloudvirt1071.eqiad.wmnet - https://phabricator.wikimedia.org/T431429#12128768 (10VRiley-WMF) 05Open→03Resolved Hey @Andrew I'm going to go ahead and close this ticket for now. Please feel free to open it if... [15:19:37] effie: Krinkle's backport didn't work out so I think you're good to go. [15:19:57] thank you! [15:20:02] effie: And no apparent train activities at this time [15:20:10] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp1111.eqiad.wmnet [15:20:26] dancy: cheers! [15:21:41] (03CR) 10Blake: [C:03+1] ProductionServices: reboot poolcounter1006 (#1/4) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311464 (https://phabricator.wikimedia.org/T431705) (owner: 10Effie Mouzeli) [15:21:55] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp1110.eqiad.wmnet [15:23:59] !log kamila@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1266 [15:23:59] !log kamila@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1266 [15:24:12] (03CR) 10TrainBranchBot: [C:03+2] "Approved by jiji@deploy2003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311464 (https://phabricator.wikimedia.org/T431705) (owner: 10Effie Mouzeli) [15:25:05] !log kamila@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage [15:25:38] (03CR) 10Btullis: "recheck" [puppet] - 10https://gerrit.wikimedia.org/r/1311476 (https://phabricator.wikimedia.org/T426458) (owner: 10Btullis) [15:25:41] (03Merged) 10jenkins-bot: ProductionServices: reboot poolcounter1006 (#1/4) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311464 (https://phabricator.wikimedia.org/T431705) (owner: 10Effie Mouzeli) [15:27:07] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp3079.esams.wmnet [15:27:51] (03PS3) 10Jforrester: abstractwiki-rust: Ship the wasm32-wasip1 Rust std, built by the image's own rustc [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1307232 (https://phabricator.wikimedia.org/T430145) [15:31:06] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp3071.esams.wmnet [15:31:43] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp5030.eqsin.wmnet [15:32:15] !log jiji@deploy2003 Started scap sync-world: Backport for [[gerrit:1311464|ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] [15:32:52] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp5023.eqsin.wmnet [15:34:12] !log jiji@deploy2003 jiji: Backport for [[gerrit:1311464|ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [15:36:12] !log cmooney@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: switch maintenance completed codfw rack b6 [15:36:18] !log cmooney@cumin1003 START - Cookbook sre.mysql.pool pool db2162: switch maintenance completed codfw rack b6 [15:37:42] !log jiji@deploy2003 jiji: Continuing with deployment [15:39:36] 10ops-eqiad, 06SRE, 06DC-Ops: Degraded RAID on an-worker1191 - https://phabricator.wikimedia.org/T431828#12128862 (10VRiley-WMF) 05Resolved→03Open [15:42:01] !log jiji@deploy2003 Finished scap sync-world: Backport for [[gerrit:1311464|ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] (duration: 09m 46s) [15:43:56] (03PS2) 10SBassett: Use non-sampled authentication log channel instead of authevents [core] (wmf/1.47.0-wmf.11) - 10https://gerrit.wikimedia.org/r/1311482 (https://phabricator.wikimedia.org/T432042) [15:43:59] !log jiji@cumin1003 START - Cookbook sre.hosts.reboot-single for host poolcounter1006.eqiad.wmnet [15:44:04] (03PS2) 10SBassett: Cleanup: remove reauth indicator from log message [core] (wmf/1.47.0-wmf.11) - 10https://gerrit.wikimedia.org/r/1311481 (https://phabricator.wikimedia.org/T432042) [15:44:52] !log kamila@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage [15:45:33] !log kamila@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1265.eqiad.wmnet with OS trixie [15:45:36] (03Abandoned) 10Btullis: cloudnative-pg-cluster: allow configuring a node maintenance window [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311469 (https://phabricator.wikimedia.org/T432104) (owner: 10Btullis) [15:46:35] 10ops-eqiad, 06SRE, 06DC-Ops: Degraded RAID on an-worker1191 - https://phabricator.wikimedia.org/T431828#12128914 (10VRiley-WMF) Hey @BTullis I apologize on this ticket. I just noticed this unit is giving me this error after the firmware update. {F94025107} I was working with Jenn on this and she mentioned... [15:47:44] !log jiji@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1006.eqiad.wmnet [15:49:02] (03CR) 10CI reject: [V:04-1] Cleanup: remove reauth indicator from log message [core] (wmf/1.47.0-wmf.11) - 10https://gerrit.wikimedia.org/r/1311481 (https://phabricator.wikimedia.org/T432042) (owner: 10SBassett) [15:49:14] !log kamila@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage [15:49:19] 10ops-codfw, 10ops-eqiad, 06SRE, 07sre-alert-triage, 06DC-Ops: Alert in need of triage: NetboxAccounting - https://phabricator.wikimedia.org/T428132#12128935 (10Jhancock.wm) a:03Jhancock.wm i'm on clinic duty this week. i'll see if i can get these fixed [15:50:00] (03CR) 10SBassett: "recheck" [core] (wmf/1.47.0-wmf.11) - 10https://gerrit.wikimedia.org/r/1311481 (https://phabricator.wikimedia.org/T432042) (owner: 10SBassett) [15:50:06] 10ops-eqiad, 06SRE, 10Data-Persistence-Backup, 06DC-Ops: db1245 crashed - https://phabricator.wikimedia.org/T431115#12128938 (10VRiley-WMF) They recommened updating the firmware on the BIOS, but that's where I'm currently getting stuck at now, since it isn't powering back on. Reached back out to Dell to se... [15:52:23] jouncebot: now [15:52:23] For the next 0 hour(s) and 7 minute(s): Train log triage (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260716T1500) [15:52:28] jouncebot: next [15:52:28] In 0 hour(s) and 7 minute(s): Puppet request window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260716T1600) [15:53:05] (03PS1) 10Scott French: shellbox: Pick up new images [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311135 (https://phabricator.wikimedia.org/T431099) [15:53:31] (03CR) 10Btullis: [C:03+2] service: add k8s-ingress-dse-postgresql TLS passthrough service [puppet] - 10https://gerrit.wikimedia.org/r/1311036 (https://phabricator.wikimedia.org/T432104) (owner: 10Btullis) [15:56:58] (03PS2) 10Effie Mouzeli: ProductionServices: reboot poolcounter1007 (#2/4) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311465 (https://phabricator.wikimedia.org/T431705) [15:59:28] (03CR) 10Blake: [C:03+1] ProductionServices: reboot poolcounter1007 (#2/4) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311465 (https://phabricator.wikimedia.org/T431705) (owner: 10Effie Mouzeli) [15:59:44] effie: nothing for the puppet window today, you're welcome to keep going [15:59:54] cheers reuven, thanks! [15:59:55] (03CR) 10Blake: [C:03+1] ProductionServices: reboot poolcounter2005 (#3/4) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311466 (https://phabricator.wikimedia.org/T431705) (owner: 10Effie Mouzeli) [16:00:05] jhathaway and rzl: Time to do the Puppet request window deploy. Don't look at me like that. You signed up for it. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260716T1600). [16:00:05] No Gerrit patches in the queue for this window AFAICS. [16:00:24] (03CR) 10TrainBranchBot: [C:03+2] "Approved by jiji@deploy2003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311465 (https://phabricator.wikimedia.org/T431705) (owner: 10Effie Mouzeli) [16:00:43] (03CR) 10Blake: [C:03+1] ProductionServices: reboot poolcounter2006 (#4/4) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311467 (https://phabricator.wikimedia.org/T431705) (owner: 10Effie Mouzeli) [16:00:53] (03CR) 10Blake: [C:03+1] ProductionServices: repool poolcounter2006 after reboot [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311468 (https://phabricator.wikimedia.org/T431705) (owner: 10Effie Mouzeli) [16:01:50] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp1112.eqiad.wmnet [16:01:53] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp1113.eqiad.wmnet [16:01:53] !log swfrench@cumin1003 START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1270.eqiad.wmnet [16:01:54] (03Merged) 10jenkins-bot: ProductionServices: reboot poolcounter1007 (#2/4) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311465 (https://phabricator.wikimedia.org/T431705) (owner: 10Effie Mouzeli) [16:01:56] !log swfrench@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1270.eqiad.wmnet [16:02:12] !log jiji@deploy2003 Started scap sync-world: Backport for [[gerrit:1311465|ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] [16:02:28] !log swfrench@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1270.eqiad.wmnet [16:03:11] !log swfrench@cumin1003 START - Cookbook sre.hosts.reimage for host wikikube-worker1270.eqiad.wmnet with OS trixie [16:03:41] !log swfrench@cumin1003 START - Cookbook sre.hosts.move-vlan for host wikikube-worker1270 [16:03:58] !log kamila@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1265.eqiad.wmnet [16:03:59] !log kamila@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1265.eqiad.wmnet [16:04:00] !log kamila@cumin1003 END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1265.eqiad.wmnet [16:04:22] !log swfrench@cumin1003 START - Cookbook sre.dns.netbox [16:04:24] (03PS1) 10Btullis: kubernetes: add the vanilla postgresql-17 image to common_images [puppet] - 10https://gerrit.wikimedia.org/r/1311486 (https://phabricator.wikimedia.org/T432104) [16:06:20] !log jiji@deploy2003 jiji: Backport for [[gerrit:1311465|ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [16:07:43] !log jiji@deploy2003 jiji: Continuing with deployment [16:08:39] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp3080.esams.wmnet [16:09:22] !log swfrench@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" [16:09:27] !log swfrench@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003" [16:09:27] !log swfrench@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [16:09:27] !log swfrench@cumin1003 START - Cookbook sre.dns.wipe-cache wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [16:09:31] !log swfrench@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [16:09:32] !log swfrench@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1270 [16:09:45] (03CR) 10SBassett: "recheck" [core] (wmf/1.47.0-wmf.11) - 10https://gerrit.wikimedia.org/r/1311481 (https://phabricator.wikimedia.org/T432042) (owner: 10SBassett) [16:10:43] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:10:56] !log swfrench@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1270 [16:10:56] !log swfrench@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1270 [16:10:58] !log kamila@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1266.eqiad.wmnet with OS trixie [16:11:51] (03CR) 10Cathal Mooney: [C:03+1] "FWIW I tested in containerlab and this config appears to work as expected. So let's give it a whirl on one router and see how it behaves." [homer/public] - 10https://gerrit.wikimedia.org/r/1311448 (https://phabricator.wikimedia.org/T320264) (owner: 10Ayounsi) [16:11:59] !log jiji@deploy2003 Finished scap sync-world: Backport for [[gerrit:1311465|ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] (duration: 09m 47s) [16:12:29] !log jiji@cumin1003 START - Cookbook sre.hosts.reboot-single for host poolcounter1007.eqiad.wmnet [16:13:04] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp3072.esams.wmnet [16:13:57] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp5031.eqsin.wmnet [16:14:59] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp5024.eqsin.wmnet [16:14:59] !log sukhe@cumin1003 END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqsin [16:15:43] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:16:15] !log jiji@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1007.eqiad.wmnet [16:17:57] (03CR) 10TrainBranchBot: [C:03+2] "Approved by jiji@deploy2003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311466 (https://phabricator.wikimedia.org/T431705) (owner: 10Effie Mouzeli) [16:18:49] kamila@cumin1003 renumber-node (PID 427915) is awaiting input [16:19:06] !log cwilliams@cumin1003 START - Cookbook sre.mysql.major-upgrade [16:19:41] (03CR) 10Dzahn: [C:03+2] jenkins: enable overlay/overlayfs kernel modules for Docker [puppet] - 10https://gerrit.wikimedia.org/r/1311477 (https://phabricator.wikimedia.org/T432326) (owner: 10Ahmon Dancy) [16:19:42] (03CR) 10Dzahn: [V:03+2 C:03+2] jenkins: enable overlay/overlayfs kernel modules for Docker [puppet] - 10https://gerrit.wikimedia.org/r/1311477 (https://phabricator.wikimedia.org/T432326) (owner: 10Ahmon Dancy) [16:20:01] !log cwilliams@cumin1003 START - Cookbook sre.hosts.reimage for host db2185.codfw.wmnet with OS trixie [16:20:40] btullis@cumin1003 roll-restart-zookeeper (PID 442079) is awaiting input [16:21:40] (03PS1) 10Sbisson: Article Guidance: migrate wikidata config [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311489 (https://phabricator.wikimedia.org/T421250) [16:21:50] !log cmooney@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: switch maintenance completed codfw rack b6 [16:23:32] !log btullis@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. [16:24:29] !log btullis@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. [16:24:54] !log kamila@cumin1003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie [16:24:56] !log kamila@cumin1003 END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1264.eqiad.wmnet [16:27:24] (03CR) 10Dzahn: [C:03+1] modules/rsync: add pidfile to stunnel.conf [puppet] - 10https://gerrit.wikimedia.org/r/1310589 (https://phabricator.wikimedia.org/T432108) (owner: 10Kamila Součková) [16:28:22] I have been waiting for 10+ for my changes to be merged in spiderpig :) [16:28:31] (03PS4) 10Dzahn: gerrit: ssh_host_keys: use openssh format and add ed25519 key [puppet] - 10https://gerrit.wikimedia.org/r/1282395 (https://phabricator.wikimedia.org/T240266) [16:28:53] any ideas of what to do here? [16:29:10] (03CR) 10Scott French: [C:03+1] kubernetes: Add a debug deployment for mw-pretrain. [puppet] - 10https://gerrit.wikimedia.org/r/1311049 (https://phabricator.wikimedia.org/T427668) (owner: 10Blake) [16:29:40] effie: Being discussed in -releng. Mostly this is colleagues pushing way too many patches in one go. [16:30:19] (03CR) 10Dzahn: "can we move ahead with this now?" [puppet] - 10https://gerrit.wikimedia.org/r/1308249 (https://phabricator.wikimedia.org/T418521) (owner: 10Dzahn) [16:30:20] effie: Options a (1) wait or (2) lossily drop all of CI and restart the service without the current in-flight requests. [16:30:25] tx James_F :) [16:31:28] 10ops-drmrs, 06SRE, 06Infrastructure-Foundations, 10netops: cr1-drmrs: No power on PEM 0 - https://phabricator.wikimedia.org/T432370 (10cmooney) 03NEW p:05Triage→03High [16:31:43] !log swfrench@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage [16:33:18] 06SRE, 06cloud-services-team, 06Data-Persistence, 06Infrastructure-Foundations, and 6 others: codfw: rack B7 maintenance - Tuesday July 21st 14:00 UTC - https://phabricator.wikimedia.org/T430928#12129249 (10cmooney) [16:34:23] 10ops-codfw, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 10netops: codfw: pod AB switches upgrade (2026) - https://phabricator.wikimedia.org/T426197#12129252 (10cmooney) [16:35:25] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [16:35:45] (03CR) 10SBassett: "recheck" [core] (wmf/1.47.0-wmf.11) - 10https://gerrit.wikimedia.org/r/1311481 (https://phabricator.wikimedia.org/T432042) (owner: 10SBassett) [16:37:41] !log kamila@cumin1003 START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1267.eqiad.wmnet [16:37:45] !log kamila@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1267.eqiad.wmnet [16:38:03] !log swfrench@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage [16:38:16] !log kamila@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1267.eqiad.wmnet [16:38:56] !log kamila@cumin1003 START - Cookbook sre.hosts.reimage for host wikikube-worker1267.eqiad.wmnet with OS trixie [16:39:09] !log cwilliams@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on db2185.codfw.wmnet with reason: host reimage [16:39:24] !log kamila@cumin1003 START - Cookbook sre.hosts.move-vlan for host wikikube-worker1267 [16:39:36] !log kamila@cumin1003 START - Cookbook sre.dns.netbox [16:40:11] (03PS2) 10Scott French: P:kubernetes::deployment_server::mediawiki: mediawiki-deployments v2 [puppet] - 10https://gerrit.wikimedia.org/r/1311081 (https://phabricator.wikimedia.org/T428971) [16:40:14] 06SRE, 10SRE-Access-Requests: Requesting access to ml-lab-users for ebernhardson - https://phabricator.wikimedia.org/T432345#12129274 (10Dzahn) The approvers for additions to this group are: ` approval: - Chris Albon - Ilias Sarantopoulos ` Please get one of them to approve on ticket. [16:40:19] 06SRE, 10SRE-Access-Requests: Requesting access to ml-lab-users for dcausse - https://phabricator.wikimedia.org/T432347#12129277 (10Dzahn) The approvers for additions to this group are: ` approval: - Chris Albon - Ilias Sarantopoulos ` Please get one of them to approve on ticket. [16:41:50] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp1115.eqiad.wmnet [16:41:50] !log sukhe@cumin1003 END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqiad [16:41:50] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp1114.eqiad.wmnet [16:41:51] !log sukhe@cumin1003 END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqiad [16:41:57] 10SRE-Access-Requests: Grant shell access (fr-tech-devs) for laurabarluzzi - https://phabricator.wikimedia.org/T431338#12129285 (10Dzahn) 05In progress→03Stalled stalled on manager approval [16:43:10] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2185.codfw.wmnet with reason: host reimage [16:43:35] 10SRE-Access-Requests, 06Infrastructure-Foundations, 10LDAP-Access-Requests, 13Patch-For-Review: Superset data access request for abibendall - https://phabricator.wikimedia.org/T430938#12129294 (10Dzahn) a:03LSobanski [16:43:49] 10ops-drmrs, 06SRE, 06Infrastructure-Foundations, 10netops: cr1-drmrs: No power on PEM 0 - https://phabricator.wikimedia.org/T432370#12129297 (10RobH) 05Open→03In progress [16:44:11] RESOLVED: [10x] BFDdown: BFD session down between cr1-eqiad and 2620:0:861:107:10:64:48:95 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [16:44:40] (03CR) 10Dzahn: [C:03+2] gerrit: add other gerrit server names to ssh host aliases [puppet] - 10https://gerrit.wikimedia.org/r/1311100 (https://phabricator.wikimedia.org/T398401) (owner: 10Dzahn) [16:45:48] !log kamila@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1266.eqiad.wmnet [16:45:49] !log kamila@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1266.eqiad.wmnet [16:45:50] !log kamila@cumin1003 END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1266.eqiad.wmnet [16:45:57] kamila@cumin1003 renumber-node (PID 445401) is awaiting input [16:46:46] (03CR) 10Effie Mouzeli: "After backporting this patch:" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311467 (https://phabricator.wikimedia.org/T431705) (owner: 10Effie Mouzeli) [16:47:17] (03PS2) 10Effie Mouzeli: ProductionServices: repool poolcounter2006 after reboot [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311468 (https://phabricator.wikimedia.org/T431705) [16:49:33] (03CR) 10Effie Mouzeli: "sudo cookbook sre.hosts.reboot-single -r reboots poolcounter2006.codfw.wmnet" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311467 (https://phabricator.wikimedia.org/T431705) (owner: 10Effie Mouzeli) [16:50:16] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Thursday, July 16 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-ite" [core] (wmf/1.47.0-wmf.11) - 10https://gerrit.wikimedia.org/r/1311481 (https://phabricator.wikimedia.org/T432042) (owner: 10SBassett) [16:50:24] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Thursday, July 16 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-ite" [core] (wmf/1.47.0-wmf.11) - 10https://gerrit.wikimedia.org/r/1311482 (https://phabricator.wikimedia.org/T432042) (owner: 10SBassett) [16:50:39] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp3081.esams.wmnet [16:50:39] !log sukhe@cumin1003 END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_esams [16:52:24] (03CR) 10CWilliams: sre.mysql: Reorder imports (031 comment) [cookbooks] - 10https://gerrit.wikimedia.org/r/1311404 (owner: 10Federico Ceratto) [16:52:52] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp3073.esams.wmnet [16:52:52] !log sukhe@cumin1003 END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_esams [16:53:59] (03CR) 10Jforrester: [C:03+2] "…" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311466 (https://phabricator.wikimedia.org/T431705) (owner: 10Effie Mouzeli) [16:54:04] (03CR) 10Brennen Bearnes: [C:03+2] ProductionServices: reboot poolcounter2005 (#3/4) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311466 (https://phabricator.wikimedia.org/T431705) (owner: 10Effie Mouzeli) [16:55:11] heh, jinx James_F [16:55:14] (03CR) 10Majavah: [C:03+2] ProductionServices: reboot poolcounter2005 (#3/4) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311466 (https://phabricator.wikimedia.org/T431705) (owner: 10Effie Mouzeli) [16:55:27] 10ops-drmrs, 06SRE, 06Infrastructure-Foundations, 10netops: cr1-drmrs: No power on PEM 0 - https://phabricator.wikimedia.org/T432370#12129329 (10RobH) Filed urgent work request CS4924823, this includes out of hours rates. Listed myself, Papaul, Cathal, and Arzhel for ticket updates. > Support, > > We... [16:55:32] C+2x4 = C…8? [16:56:04] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp5032.eqsin.wmnet [16:56:04] !log sukhe@cumin1003 END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqsin [16:56:59] !log kamila@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" [16:57:03] !log kamila@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003" [16:57:04] !log kamila@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [16:57:04] !log kamila@cumin1003 START - Cookbook sre.dns.wipe-cache wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [16:57:07] !log kamila@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [16:57:08] !log kamila@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1267 [16:57:45] !log kamila@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1267 [16:57:45] !log kamila@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1267 [16:58:09] !log swfrench@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1270.eqiad.wmnet with OS trixie [16:59:15] (03PS2) 10Effie Mouzeli: ProductionServices: reboot poolcounter2005 (#3/4) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311466 (https://phabricator.wikimedia.org/T431705) [16:59:21] (03CR) 10Jforrester: "…" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311466 (https://phabricator.wikimedia.org/T431705) (owner: 10Effie Mouzeli) [16:59:49] !log cmooney@cumin1003 START - Cookbook sre.dns.netbox [17:00:03] effie: 1311466 is finally picked up by Zuul; I think it looked at it and saw it wasn't based on master, so just decided to do nothing. (Yay Zuul v2.) [17:00:05] bd808: That opportune time for a Cloud Services/Technical Documentation weekly deploy (Toolhub, Developer portal, Striker) deploy is upon us again. Don't be afraid. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260716T1700). [17:00:05] swfrench-wmf and dancy: #bothumor I � Unicode. All rise for MediaWiki infrastructure (UTC late) deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260716T1700). [17:00:12] o/ [17:00:20] (03CR) 10Btullis: [C:04-1] "I think that we can do something better than this. Re-working this patch." [puppet] - 10https://gerrit.wikimedia.org/r/1311486 (https://phabricator.wikimedia.org/T432104) (owner: 10Btullis) [17:00:25] (03CR) 10TrainBranchBot: [C:03+2] "Approved by jiji@deploy2003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311466 (https://phabricator.wikimedia.org/T431705) (owner: 10Effie Mouzeli) [17:00:33] (03Merged) 10jenkins-bot: ProductionServices: reboot poolcounter2005 (#3/4) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311466 (https://phabricator.wikimedia.org/T431705) (owner: 10Effie Mouzeli) [17:00:44] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2185.codfw.wmnet with OS trixie [17:00:51] !log jiji@deploy2003 Started scap sync-world: Backport for [[gerrit:1311466|ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] [17:01:09] swfrench@cumin1003 renumber-node (PID 439058) is awaiting input [17:01:40] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) [17:02:07] (03CR) 10JHathaway: nftables: add support for the VRRP protocol (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1308177 (https://phabricator.wikimedia.org/T427799) (owner: 10JHathaway) [17:03:10] !log jiji@deploy2003 jiji: Backport for [[gerrit:1311466|ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [17:03:16] Finally. [17:03:33] effie: glad to see your backport is rolling :) let me know when you've made it do the 2005 reboot and I should take over the remainder [17:04:10] !log jiji@deploy2003 jiji: Continuing with deployment [17:04:15] s/made it do/made it to/ [17:04:42] swfrench-wmf: thank you scott! [17:04:52] (03CR) 10Bking: "Per IRC conversation with @tfogli@wikimedia.org, we'll wait for a review from @cwhite@wikimedia.org before moving forward." [puppet] - 10https://gerrit.wikimedia.org/r/1311422 (https://phabricator.wikimedia.org/T324335) (owner: 10Btullis) [17:04:56] (03CR) 10Majavah: nftables: add support for the VRRP protocol (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1308177 (https://phabricator.wikimedia.org/T427799) (owner: 10JHathaway) [17:05:55] (03PS2) 10Effie Mouzeli: ProductionServices: reboot poolcounter2006 (#4/4) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311467 (https://phabricator.wikimedia.org/T431705) [17:06:02] cmooney@cumin1003 netbox (PID 450706) is awaiting input [17:08:25] !log jiji@deploy2003 Finished scap sync-world: Backport for [[gerrit:1311466|ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] (duration: 07m 34s) [17:08:39] !log cmooney@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" [17:08:44] !log cmooney@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003" [17:08:44] !log cmooney@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [17:09:08] !log jiji@cumin1003 START - Cookbook sre.hosts.reboot-single for host poolcounter2005.codfw.wmnet [17:11:55] !log swfrench@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1270.eqiad.wmnet [17:11:56] !log swfrench@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1270.eqiad.wmnet [17:11:58] !log swfrench@cumin1003 END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1270.eqiad.wmnet [17:12:02] (03CR) 10JHathaway: nftables: add support for the VRRP protocol (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1308177 (https://phabricator.wikimedia.org/T427799) (owner: 10JHathaway) [17:12:56] !log jiji@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2005.codfw.wmnet [17:16:19] (03PS1) 10BryanDavis: developer-portal: Bump container to 2026-07-16-140503-production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311498 [17:17:49] !log kamila@cumin1003 START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1268.eqiad.wmnet [17:17:53] !log kamila@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1268.eqiad.wmnet [17:18:25] !log kamila@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1268.eqiad.wmnet [17:18:29] !log kamila@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage [17:19:28] (03PS3) 10Effie Mouzeli: ProductionServices: reboot poolcounter2006 (#4/4) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311467 (https://phabricator.wikimedia.org/T431705) [17:20:41] (03CR) 10BryanDavis: [C:03+2] developer-portal: Bump container to 2026-07-16-140503-production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311498 (owner: 10BryanDavis) [17:21:00] swfrench-wmf: over to you, poolcounter2005 is good to be repooled [17:21:26] kamila@cumin1003 renumber-node (PID 451981) is awaiting input [17:21:40] effie: great, thanks! I'll continue with https://gerrit.wikimedia.org/r/c/operations/mediawiki-config/+/1311467, reboot 2006, and then continue with https://gerrit.wikimedia.org/r/c/operations/mediawiki-config/+/1311468 [17:22:08] 👍 [17:22:51] 10SRE-Access-Requests, 06Infrastructure-Foundations, 10LDAP-Access-Requests, 13Patch-For-Review: Superset data access request for abibendall - https://phabricator.wikimedia.org/T430938#12129476 (10LSobanski) I haven't seen a request notification come in, @jhathaway could you take a look if it's in the appr... [17:23:08] (03CR) 10Scott French: [C:03+1] "Thanks, Effie!" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311467 (https://phabricator.wikimedia.org/T431705) (owner: 10Effie Mouzeli) [17:23:08] (03CR) 10BryanDavis: "recheck" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311418 (https://phabricator.wikimedia.org/T432204) (owner: 10JavierMonton) [17:23:50] !log kamila@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage [17:24:56] * swfrench-wmf is proceeding with poolcounter reboots, then planned infra-window stuff [17:25:42] (03CR) 10TrainBranchBot: [C:03+2] "Approved by swfrench@deploy2003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311467 (https://phabricator.wikimedia.org/T431705) (owner: 10Effie Mouzeli) [17:26:30] 06SRE, 10Cloud-Services: Wikimania 2026: Request to whitelist venue IP for all Wikimedia projects, Toolforge, and Cloud VPS etc... - https://phabricator.wikimedia.org/T432377 (10Michael_Barbereau_WMFr) 03NEW The #Cloud-Services project tag is not intended to have any tasks. Please check the list on https://p... [17:27:36] (03Merged) 10jenkins-bot: ProductionServices: reboot poolcounter2006 (#4/4) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311467 (https://phabricator.wikimedia.org/T431705) (owner: 10Effie Mouzeli) [17:27:51] !log swfrench@deploy2003 Started scap sync-world: Backport for [[gerrit:1311467|ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] [17:28:10] !log kamila@cumin1003 START - Cookbook sre.hosts.reimage for host wikikube-worker1268.eqiad.wmnet with OS trixie [17:28:12] (03PS1) 10Dzahn: gerrit: add SSH config for replication between servers [puppet] - 10https://gerrit.wikimedia.org/r/1311500 (https://phabricator.wikimedia.org/T398401) [17:28:38] !log kamila@cumin1003 START - Cookbook sre.hosts.move-vlan for host wikikube-worker1268 [17:28:45] (03PS2) 10Dzahn: gerrit: add SSH config for replication between servers [puppet] - 10https://gerrit.wikimedia.org/r/1311500 (https://phabricator.wikimedia.org/T398401) [17:28:46] !log kamila@cumin1003 START - Cookbook sre.dns.netbox [17:28:47] (03CR) 10BryanDavis: [C:03+2] "This change seems to have a bad state in zuul and thus is stuck blocking the gate-ans-submit queue. Trying another +2 to unstick it." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311418 (https://phabricator.wikimedia.org/T432204) (owner: 10JavierMonton) [17:29:38] (03PS3) 10Dzahn: gerrit: add SSH config for replication between servers [puppet] - 10https://gerrit.wikimedia.org/r/1311500 (https://phabricator.wikimedia.org/T398401) [17:29:50] !log swfrench@deploy2003 jiji, swfrench: Backport for [[gerrit:1311467|ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [17:31:04] !log swfrench@deploy2003 jiji, swfrench: Continuing with deployment [17:33:05] 06SRE, 06Infrastructure-Foundations, 10netops: Don't announce OSPF routes in unicast BGP on Nokia SR-Linux - https://phabricator.wikimedia.org/T423430#12129544 (10cmooney) 05Open→03Resolved Re-resolving after updating the alert for filtered prefixes to also monitor peers in the 'ibgp' group we use on... [17:34:10] !log kamila@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" [17:34:14] !log kamila@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003" [17:34:14] !log kamila@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [17:34:15] !log kamila@cumin1003 START - Cookbook sre.dns.wipe-cache wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [17:34:17] 10ops-drmrs, 06SRE, 06Infrastructure-Foundations, 10netops: cr1-drmrs: No power on PEM 0 - https://phabricator.wikimedia.org/T432370#12129548 (10Papaul) ` Jul 16 14:52:36 cr1-drmrs chassisd[18580]: CHASSISD_SNMP_TRAP6: SNMP trap generated: Power Supply failed (jnxContentsContainerIndex 2, jnxContentsL1Ind... [17:34:18] !log kamila@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [17:34:19] !log kamila@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1268 [17:35:19] !log swfrench@deploy2003 Finished scap sync-world: Backport for [[gerrit:1311467|ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] (duration: 07m 27s) [17:37:02] 06SRE, 06Infrastructure-Foundations, 10netops: cr1-drmrs: unexpected reboot Jul 16 2026 - https://phabricator.wikimedia.org/T432380 (10cmooney) 03NEW p:05Triage→03High [17:37:15] 06SRE, 06Infrastructure-Foundations, 10netops: cr1-drmrs: unexpected reboot Jul 16 2026 - https://phabricator.wikimedia.org/T432380#12129571 (10cmooney) [17:37:17] 10ops-drmrs, 06SRE, 06Infrastructure-Foundations, 10netops: cr1-drmrs: No power on PEM 0 - https://phabricator.wikimedia.org/T432370#12129570 (10cmooney) [17:37:22] 06SRE, 06Infrastructure-Foundations, 10netops: cr1-drmrs: unexpected reboot Jul 16 2026 - https://phabricator.wikimedia.org/T432380#12129572 (10cmooney) 05Open→03Resolved [17:37:33] (03CR) 10BryanDavis: [C:04-2] "Adding a +2 seems not to have changed anything. Next I will try changing my +2 to a -2 to see if that removes this from the merge queue so" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311418 (https://phabricator.wikimedia.org/T432204) (owner: 10JavierMonton) [17:39:06] 10ops-drmrs, 06SRE, 06Infrastructure-Foundations, 10netops: cr1-drmrs: No power on PEM 0 - https://phabricator.wikimedia.org/T432370#12129577 (10cmooney) 05In progress→03Resolved Remote hands were able to restore power: ` Jul 16 17:12:41 cr1-drmrs alarmd[23564]: Alarm cleared: PS SFXPC id=53687098... [17:39:51] !log swfrench@cumin1003 START - Cookbook sre.hosts.reboot-single for host poolcounter2006.codfw.wmnet [17:41:09] (03CR) 10Dzahn: [V:03+1] "https://puppet-compiler.wmflabs.org/output/1311500/9024/gerrit1003.wikimedia.org/index.html" [puppet] - 10https://gerrit.wikimedia.org/r/1311500 (https://phabricator.wikimedia.org/T398401) (owner: 10Dzahn) [17:41:33] 06SRE, 10Cloud-VPS, 10Toolforge, 06tools-infrastructure-team, 06tools-platform-team: Wikimania 2026: Request to whitelist venue IP for all Wikimedia projects, Toolforge, and Cloud VPS etc... - https://phabricator.wikimedia.org/T432377#12129582 (10Michael_Barbereau_WMFr) [17:43:39] !log swfrench@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2006.codfw.wmnet [17:44:00] 06SRE, 10Cloud-VPS, 10Toolforge, 06tools-infrastructure-team, 06tools-platform-team: Wikimania 2026: Request to whitelist venue IP for all Wikimedia projects, Toolforge, and Cloud VPS etc... - https://phabricator.wikimedia.org/T432377#12129589 (10Michael_Barbereau_WMFr) [17:44:21] (03CR) 10BryanDavis: [C:03+2] "Trying another +2. If this is still an apparent no-op I will change tactics." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311418 (https://phabricator.wikimedia.org/T432204) (owner: 10JavierMonton) [17:44:39] !log kamila@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1267.eqiad.wmnet with OS trixie [17:44:46] 06SRE, 10Cloud-VPS, 10Toolforge, 06tools-infrastructure-team, 06tools-platform-team: Wikimania 2026: Request to whitelist venue IP for all Wikimedia projects, Toolforge, and Cloud VPS etc... - https://phabricator.wikimedia.org/T432377#12129592 (10Michael_Barbereau_WMFr) [17:45:01] !log kamila@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1268 [17:45:01] !log kamila@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1268 [17:45:47] (03CR) 10TrainBranchBot: [C:03+2] "Approved by swfrench@deploy2003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311468 (https://phabricator.wikimedia.org/T431705) (owner: 10Effie Mouzeli) [17:46:16] dancy: once this backport completes, we should be good to switch gears and do the mediawiki-deployments.yaml change. [17:47:04] (03PS1) 10Bking: wcqs: Add depool metadata to hieradata [puppet] - 10https://gerrit.wikimedia.org/r/1311502 (https://phabricator.wikimedia.org/T327300) [17:47:40] (03CR) 10Dzahn: "recheck" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311418 (https://phabricator.wikimedia.org/T432204) (owner: 10JavierMonton) [17:47:50] swfrench-wmf: Ok [17:48:01] hmmmm ... I guess that might be asking too much of the universe [17:48:31] (03PS2) 10JavierMonton: stream: pageview-trending-relative + webrequest-page-view [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311418 (https://phabricator.wikimedia.org/T432204) [17:48:36] (03CR) 10Jforrester: [C:03+2] "…" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311418 (https://phabricator.wikimedia.org/T432204) (owner: 10JavierMonton) [17:54:24] 06SRE, 06Infrastructure-Foundations, 10netops: JunOS: Investigate BGP PIC Edge / Protection - https://phabricator.wikimedia.org/T432381 (10cmooney) 03NEW p:05Triage→03Low [17:55:13] dancy: swfrench-wmf lmk if i should hold the train for your next change [17:55:49] (03PS3) 10Effie Mouzeli: ProductionServices: repool poolcounter2006 after reboot [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311468 (https://phabricator.wikimedia.org/T431705) [17:57:29] jeena: thanks! depending on how long this ongoing backport takes, let's sync shortly after your window starts. our change "should be a noop" so we could in theory sneak in after you if that's alright (but only when you're happy with how group2 looks) [17:57:50] (03CR) 10Scott French: [C:03+2] ProductionServices: repool poolcounter2006 after reboot [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311468 (https://phabricator.wikimedia.org/T431705) (owner: 10Effie Mouzeli) [17:57:54] (03Merged) 10jenkins-bot: admin_ng: add the postgresql-superset-metrics namespace on dse-k8s-eqiad [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311463 (https://phabricator.wikimedia.org/T432104) (owner: 10Btullis) [17:57:59] (03CR) 10TrainBranchBot: [C:03+2] "Approved by swfrench@deploy2003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311468 (https://phabricator.wikimedia.org/T431705) (owner: 10Effie Mouzeli) [17:58:25] kamila@cumin1003 renumber-node (PID 445401) is awaiting input [17:58:34] (03Merged) 10jenkins-bot: developer-portal: Bump container to 2026-07-16-140503-production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311498 (owner: 10BryanDavis) [17:58:46] (03Merged) 10jenkins-bot: stream: pageview-trending-relative + webrequest-page-view [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311418 (https://phabricator.wikimedia.org/T432204) (owner: 10JavierMonton) [17:59:00] (03Merged) 10jenkins-bot: ProductionServices: repool poolcounter2006 after reboot [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311468 (https://phabricator.wikimedia.org/T431705) (owner: 10Effie Mouzeli) [17:59:18] !log swfrench@deploy2003 Started scap sync-world: Backport for [[gerrit:1311468|ProductionServices: repool poolcounter2006 after reboot (T431705)]] [18:00:05] jeena and hashar: Deploy window MediaWiki train - Utc-7+Utc-0 Version (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260716T1800) [18:00:36] 06SRE, 10Cloud-VPS, 10Toolforge, 06tools-infrastructure-team, 06tools-platform-team: Wikimania 2026: Request to exempt venue IP for all Wikimedia projects, Toolforge, and Cloud VPS etc... - https://phabricator.wikimedia.org/T432377#12129642 (10taavi) [18:00:49] jeena: alas, still waiting on backport after CI woes ... ETA 5-6 minutes [18:01:16] when were you planning to roll to group2? [18:01:21] !log bd808@deploy2003 helmfile [staging] START helmfile.d/services/developer-portal: apply [18:01:24] !log swfrench@deploy2003 jiji, swfrench: Backport for [[gerrit:1311468|ProductionServices: repool poolcounter2006 after reboot (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [18:01:26] 06SRE, 10Cloud-VPS, 10Toolforge, 06tools-infrastructure-team, 06tools-platform-team: Wikimania 2026: Request to exempt venue IP for all Wikimedia projects, Toolforge, and Cloud VPS etc... - https://phabricator.wikimedia.org/T432377#12129647 (10taavi) Former Wikimania events have not needed exemptions fro... [18:01:40] !log bd808@deploy2003 helmfile [staging] DONE helmfile.d/services/developer-portal: apply [18:02:08] !log bd808@deploy2003 helmfile [codfw] START helmfile.d/services/developer-portal: apply [18:02:26] !log bd808@deploy2003 helmfile [codfw] DONE helmfile.d/services/developer-portal: apply [18:02:26] !log swfrench@deploy2003 jiji, swfrench: Continuing with deployment [18:02:49] !log bd808@deploy2003 helmfile [eqiad] START helmfile.d/services/developer-portal: apply [18:03:17] !log bd808@deploy2003 helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply [18:03:50] That's my window done (late). [18:06:08] !log kamila@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage [18:06:24] !log bking@deploy2003 Started deploy [wdqs/wdqs@e8fb00c] (wcqs): T430879 [18:06:27] T430879: Migrate WCQS to Bookworm or later - https://phabricator.wikimedia.org/T430879 [18:06:33] !log bking@deploy2003 Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): T430879 (duration: 00m 23s) [18:06:53] !log swfrench@deploy2003 Finished scap sync-world: Backport for [[gerrit:1311468|ProductionServices: repool poolcounter2006 after reboot (T431705)]] (duration: 07m 34s) [18:08:00] !log bking@cumin2003 START - Cookbook sre.wdqs.data-transfer (T430879, restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards [18:08:03] jeena: all yours to roll the train. if you'd be alright with us deploying our (should be noop) change once the dust settles after that, let me know :) [18:08:09] thanks for you patience! [18:09:04] Okay! [18:09:46] !log kamila@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage [18:09:47] and yes I will be rolling to group 2 [18:10:18] oh, when 😆 Well, now I suppose [18:11:11] * swfrench-wmf thumbs up [18:11:31] (03PS1) 10TrainBranchBot: group2 to 1.47.0-wmf.11 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311507 (https://phabricator.wikimedia.org/T430830) [18:11:34] (03CR) 10TrainBranchBot: [C:03+2] "Initiated by jhuneidi@deploy2003" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311507 (https://phabricator.wikimedia.org/T430830) (owner: 10TrainBranchBot) [18:12:15] (03PS3) 10Scott French: P:kubernetes::deployment_server::mediawiki: mediawiki-deployments v2 [puppet] - 10https://gerrit.wikimedia.org/r/1311081 (https://phabricator.wikimedia.org/T428971) [18:12:27] (03Merged) 10jenkins-bot: group2 to 1.47.0-wmf.11 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311507 (https://phabricator.wikimedia.org/T430830) (owner: 10TrainBranchBot) [18:16:02] !log kamila@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1267.eqiad.wmnet [18:16:03] !log kamila@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1267.eqiad.wmnet [18:16:05] !log kamila@cumin1003 END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1267.eqiad.wmnet [18:17:50] (03CR) 10Cathal Mooney: [C:03+1] wcqs: Add depool metadata to hieradata [puppet] - 10https://gerrit.wikimedia.org/r/1311502 (https://phabricator.wikimedia.org/T327300) (owner: 10Bking) [18:18:46] !log jhuneidi@deploy2003 rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.11 refs T430830 [18:18:50] T430830: 1.47.0-wmf.11 deployment blockers - https://phabricator.wikimedia.org/T430830 [18:20:11] (03CR) 10BryanDavis: [C:03+2] "The actions that eventually worked were rebasing the change and giving another +2. Just the rebase may have been enough, but the +2 wouldn" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311418 (https://phabricator.wikimedia.org/T432204) (owner: 10JavierMonton) [18:20:43] FIRING: [2x] JobUnavailable: Reduced availability for job jmx_wcqs_blazegraph in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [18:23:06] swfrench-wmf: ready for you now [18:23:22] jeena: oh, great! thank you very much :) [18:23:34] let's see how my friend jenkins is doing ... [18:28:29] :) [18:30:29] (03CR) 10Jeena Huneidi: [C:03+1] Update update_version.py to be compatible with ruamel >=0.15.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310090 (owner: 10Mvolz) [18:31:48] !log kamila@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1268.eqiad.wmnet with OS trixie [18:35:48] kamila@cumin1003 renumber-node (PID 451981) is awaiting input [18:37:26] (03CR) 10CI reject: [V:04-1] kubernetes: Add a debug deployment for mw-pretrain. [puppet] - 10https://gerrit.wikimedia.org/r/1311049 (https://phabricator.wikimedia.org/T427668) (owner: 10Blake) [18:37:39] (03CR) 10CI reject: [V:04-1] gerrit: add SSH config for replication between servers [puppet] - 10https://gerrit.wikimedia.org/r/1311500 (https://phabricator.wikimedia.org/T398401) (owner: 10Dzahn) [18:37:58] (03CR) 10CI reject: [V:04-1] Add zsinger to the analytics-privatedata-users group [puppet] - 10https://gerrit.wikimedia.org/r/1311476 (https://phabricator.wikimedia.org/T426458) (owner: 10Btullis) [18:40:42] CI for operations/puppet is back [18:40:43] dancy: I think we're ready to go, thanks to your jenkins wrangling [18:41:15] dancy: I'll get that merged and we can test this out? [18:41:21] Yep, ready [18:41:24] (03CR) 10Scott French: [C:03+2] P:kubernetes::deployment_server::mediawiki: mediawiki-deployments v2 [puppet] - 10https://gerrit.wikimedia.org/r/1311081 (https://phabricator.wikimedia.org/T428971) (owner: 10Scott French) [18:43:18] 10ops-eqiad, 06SRE, 06DC-Ops: hw troubleshooting: DIMM module in slot A7 for db1208.eqiad.wmnet - https://phabricator.wikimedia.org/T432116#12129798 (10VRiley-WMF) Hey @BTullis we should be able to swap the DIMM in this unit. I'm currently having trouble seeing what type of DIMM this is, but we should have s... [18:44:17] alright, puppet-agent is running [18:45:43] (03PS3) 10JHathaway: nftables: add support for the VRRP protocol [puppet] - 10https://gerrit.wikimedia.org/r/1308177 (https://phabricator.wikimedia.org/T427799) [18:46:10] (03CR) 10JHathaway: nftables: add support for the VRRP protocol (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1308177 (https://phabricator.wikimedia.org/T427799) (owner: 10JHathaway) [18:46:42] TIL puppet-agent runs are sorta non-hermetic :grimace: [18:47:51] (i.e., HEAD of production moved, and a puppet type changed names, which confused things) [18:49:25] (03PS4) 10Dzahn: gerrit: add SSH config for replication between servers [puppet] - 10https://gerrit.wikimedia.org/r/1311500 (https://phabricator.wikimedia.org/T398401) [18:49:58] (03CR) 10CI reject: [V:04-1] gerrit: add SSH config for replication between servers [puppet] - 10https://gerrit.wikimedia.org/r/1311500 (https://phabricator.wikimedia.org/T398401) (owner: 10Dzahn) [18:51:19] (03PS5) 10Dzahn: gerrit: add SSH config for replication between servers [puppet] - 10https://gerrit.wikimedia.org/r/1311500 (https://phabricator.wikimedia.org/T398401) [18:51:51] dancy: alright, we're ready to test. shall I do that, or would you prefer to? (e.g., to avoid a game of telephone in the event scap sneezes at the config) [18:52:13] 06SRE, 10SRE-Access-Requests: Requesting access to Superset for jniren-ctr - https://phabricator.wikimedia.org/T432273#12129814 (10KFrancis) Hi all, as Jonathan Niren is currently a contractor with the WMF, an NDA is not needed. It's covered under the contractor contract. Thank you! [18:52:30] haha, sure I'll run it. [18:52:50] (03CR) 10RLazarus: abstractwiki-rust: Ship the wasm32-wasip1 Rust std, built by the image's own rustc (031 comment) [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1307232 (https://phabricator.wikimedia.org/T430145) (owner: 10Jforrester) [18:52:51] * swfrench-wmf thumbs up [18:53:01] !log dancy@deploy2003 Started scap sync-world: testing T428971 [18:53:05] T428971: Allow configuration of canary and production checks based on deployment target - https://phabricator.wikimedia.org/T428971 [18:53:33] FIRING: KubernetesCalicoDown: wikikube-worker1264.eqiad.wmnet is not running calico-node Pod - https://wikitech.wikimedia.org/wiki/Calico#Operations - https://grafana.wikimedia.org/d/G8zPL7-Wz/?var-dc=eqiad%20prometheus%2Fk8s&var-instance=wikikube-worker1264.eqiad.wmnet - https://alerts.wikimedia.org/?q=alertname%3DKubernetesCalicoDown [18:55:30] (03CR) 10Dzahn: [C:03+2] gerrit: add SSH config for replication between servers [puppet] - 10https://gerrit.wikimedia.org/r/1311500 (https://phabricator.wikimedia.org/T398401) (owner: 10Dzahn) [18:55:42] !log dancy@deploy2003 Finished scap sync-world: testing T428971 (duration: 02m 41s) [18:56:07] Success! [18:56:13] \i/ [18:57:04] FIRING: HelmReleaseBadStatus: Helm release wdqs/scholarly-external on k8s-dse@codfw in state pending-rollback - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [18:57:17] dancy: awesome, thank you very much for all of your work on this :) [18:59:43] (03PS1) 10Ahmon Dancy: buildkitd: Bump buildkit image to wmf-v0.31.2 [puppet] - 10https://gerrit.wikimedia.org/r/1311523 (https://phabricator.wikimedia.org/T432360) [19:00:15] 06SRE, 06cloud-services-team, 10Infrastructure Security, 06Infrastructure-Foundations, and 5 others: July 2026 Bookworm/Trixie reboots: Data Platform SRE - https://phabricator.wikimedia.org/T431826#12129847 (10bking) [19:03:00] !log kamila@cumin1003 START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1269.eqiad.wmnet [19:03:03] !log kamila@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1269.eqiad.wmnet [19:03:36] !log kamila@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1269.eqiad.wmnet [19:05:39] FIRING: CertAlmostExpired: gNMI TLS certificate for lsw1-f6-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [19:05:50] !log kamila@cumin1003 START - Cookbook sre.hosts.reimage for host wikikube-worker1269.eqiad.wmnet with OS trixie [19:06:20] !log kamila@cumin1003 START - Cookbook sre.hosts.move-vlan for host wikikube-worker1269 [19:09:22] kamila@cumin1003 renumber-node (PID 468986) is awaiting input [19:10:24] (03PS1) 10Ahmon Dancy: scap.cfg.erb: Drop testservers_check_cmd_k8s [puppet] - 10https://gerrit.wikimedia.org/r/1311526 (https://phabricator.wikimedia.org/T428971) [19:12:09] (03PS2) 10Ahmon Dancy: scap.cfg.erb: Drop testservers_check_cmd_k8s [puppet] - 10https://gerrit.wikimedia.org/r/1311526 (https://phabricator.wikimedia.org/T428971) [19:12:23] !log bking@cumin2003 START - Cookbook sre.hosts.reboot-single for host wdqs1029.eqiad.wmnet [19:12:27] kamila@cumin1003 renumber-node (PID 451981) is awaiting input [19:15:11] (03CR) 10Scott French: [C:03+1] scap.cfg.erb: Drop testservers_check_cmd_k8s [puppet] - 10https://gerrit.wikimedia.org/r/1311526 (https://phabricator.wikimedia.org/T428971) (owner: 10Ahmon Dancy) [19:16:10] (03CR) 10Scott French: [C:03+2] scap.cfg.erb: Drop testservers_check_cmd_k8s [puppet] - 10https://gerrit.wikimedia.org/r/1311526 (https://phabricator.wikimedia.org/T428971) (owner: 10Ahmon Dancy) [19:17:19] !log kamila@cumin1003 START - Cookbook sre.dns.netbox [19:19:46] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1029.eqiad.wmnet [19:19:51] !log bking@cumin2003 START - Cookbook sre.hosts.reboot-single for host wdqs1030.eqiad.wmnet [19:21:40] (03PS1) 10Urbanecm: AddImage: Request only standard thumbnail sizes [extensions/GrowthExperiments] (wmf/1.47.0-wmf.11) - 10https://gerrit.wikimedia.org/r/1311529 (https://phabricator.wikimedia.org/T428797) [19:22:54] (03CR) 10Bking: [C:03+2] wcqs: Add depool metadata to hieradata [puppet] - 10https://gerrit.wikimedia.org/r/1311502 (https://phabricator.wikimedia.org/T327300) (owner: 10Bking) [19:23:26] !log bking@cumin2003 END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) (T430879, restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards [19:23:26] kamila@cumin1003 renumber-node (PID 468986) is awaiting input [19:23:29] T430879: Migrate WCQS to Bookworm or later - https://phabricator.wikimedia.org/T430879 [19:25:43] FIRING: [2x] JobUnavailable: Reduced availability for job jmx_wcqs_blazegraph in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [19:26:21] jouncebot: nowandnext [19:26:21] For the next 0 hour(s) and 33 minute(s): MediaWiki train - Utc-7+Utc-0 Version (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260716T1800) [19:26:21] In 0 hour(s) and 33 minute(s): UTC late backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260716T2000) [19:26:49] unless there are any concerns, I might sneak in some low-risk shellbox service updates shortly [19:27:33] kamila@cumin1003 renumber-node (PID 451981) is awaiting input [19:27:39] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1030.eqiad.wmnet [19:27:43] !log bking@cumin2003 START - Cookbook sre.hosts.reboot-single for host wdqs1031.eqiad.wmnet [19:28:04] (03CR) 10Scott French: [C:03+2] shellbox: Pick up new images [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311135 (https://phabricator.wikimedia.org/T431099) (owner: 10Scott French) [19:30:54] (03Merged) 10jenkins-bot: shellbox: Pick up new images [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311135 (https://phabricator.wikimedia.org/T431099) (owner: 10Scott French) [19:31:49] !log swfrench@deploy2003 helmfile [staging] START helmfile.d/services/shellbox: apply [19:32:20] !log swfrench@deploy2003 helmfile [staging] DONE helmfile.d/services/shellbox: apply [19:32:21] !log swfrench@deploy2003 helmfile [staging] START helmfile.d/services/shellbox-constraints: apply [19:32:27] (03CR) 10Ladsgroup: "recheck" [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1310618 (https://phabricator.wikimedia.org/T431767) (owner: 10Ladsgroup) [19:32:36] !log swfrench@deploy2003 helmfile [staging] DONE helmfile.d/services/shellbox-constraints: apply [19:32:37] !log swfrench@deploy2003 helmfile [staging] START helmfile.d/services/shellbox-media: apply [19:32:52] !log swfrench@deploy2003 helmfile [staging] DONE helmfile.d/services/shellbox-media: apply [19:32:54] !log swfrench@deploy2003 helmfile [staging] START helmfile.d/services/shellbox-syntaxhighlight: apply [19:33:09] !log swfrench@deploy2003 helmfile [staging] DONE helmfile.d/services/shellbox-syntaxhighlight: apply [19:33:10] !log swfrench@deploy2003 helmfile [staging] START helmfile.d/services/shellbox-timeline: apply [19:33:31] !log swfrench@deploy2003 helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply [19:33:32] !log swfrench@deploy2003 helmfile [staging] START helmfile.d/services/shellbox-video: apply [19:33:59] !log swfrench@deploy2003 helmfile [staging] DONE helmfile.d/services/shellbox-video: apply [19:35:33] !log swfrench@deploy2003 helmfile [codfw] START helmfile.d/services/shellbox: apply [19:35:49] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1031.eqiad.wmnet [19:35:54] !log bking@cumin2003 START - Cookbook sre.hosts.reboot-single for host wdqs1032.eqiad.wmnet [19:36:11] !log swfrench@deploy2003 helmfile [codfw] DONE helmfile.d/services/shellbox: apply [19:36:42] !log swfrench@deploy2003 helmfile [codfw] START helmfile.d/services/shellbox-constraints: apply [19:36:47] jouncebot: nowandnext [19:36:47] For the next 0 hour(s) and 23 minute(s): MediaWiki train - Utc-7+Utc-0 Version (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260716T1800) [19:36:47] In 0 hour(s) and 23 minute(s): UTC late backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260716T2000) [19:36:48] (03PS1) 10Bking: datahubsearch: move to 'insetup' role [puppet] - 10https://gerrit.wikimedia.org/r/1311530 (https://phabricator.wikimedia.org/T432148) [19:38:49] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1311530 (https://phabricator.wikimedia.org/T432148) (owner: 10Bking) [19:39:46] !log swfrench@deploy2003 helmfile [codfw] DONE helmfile.d/services/shellbox-constraints: apply [19:40:17] !log swfrench@deploy2003 helmfile [codfw] START helmfile.d/services/shellbox-media: apply [19:40:30] !log swfrench@deploy2003 helmfile [codfw] DONE helmfile.d/services/shellbox-media: apply [19:41:02] !log swfrench@deploy2003 helmfile [codfw] START helmfile.d/services/shellbox-syntaxhighlight: apply [19:41:17] !log swfrench@deploy2003 helmfile [codfw] DONE helmfile.d/services/shellbox-syntaxhighlight: apply [19:41:48] !log swfrench@deploy2003 helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply [19:42:10] !log swfrench@deploy2003 helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply [19:42:41] !log swfrench@deploy2003 helmfile [codfw] START helmfile.d/services/shellbox-video: apply [19:42:49] (03PS2) 10Bking: datahubsearch: move to 'insetup' role [puppet] - 10https://gerrit.wikimedia.org/r/1311530 (https://phabricator.wikimedia.org/T432148) [19:43:21] FIRING: CirrusSearchBackendMemoryIssue: CirrusSearch backend failed 67.77777777777779 times in the last 10 minutes due to memory usage (mw@codfw to dnsdisc) - https://wikitech.wikimedia.org/wiki/Search#Health/Activity_Monitoring - https://grafana.wikimedia.org/d/dc04b9f2-b8d5-4ab6-9482-5d9a75728951/elasticsearch-percentiles?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchBackendMemoryIssue [19:43:30] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1032.eqiad.wmnet [19:43:34] !log bking@cumin2003 START - Cookbook sre.hosts.reboot-single for host wdqs1033.eqiad.wmnet [19:43:37] !log swfrench@deploy2003 helmfile [codfw] DONE helmfile.d/services/shellbox-video: apply [19:48:54] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1033.eqiad.wmnet [19:48:58] !log bking@cumin2003 START - Cookbook sre.hosts.reboot-single for host wdqs1034.eqiad.wmnet [19:49:21] PROBLEM - Host db2207 #page is DOWN: PING CRITICAL - Packet loss = 100% [19:50:07] That's s2 master codfw [19:50:22] :( [19:50:32] PROBLEM - MariaDB Replica IO: s2 #page on db2225 is CRITICAL: CRITICAL slave_io_state Slave_IO_Running: No, Errno: 2013, Errmsg: error reconnecting to master repl2024@db2207.codfw.wmnet:3306 - retry-time: 60 maximum-retries: 100000 message: Lost connection to server at waiting for initial communication packet, system error: 110 Connection timed out https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [19:50:41] PROBLEM - MariaDB Replica IO: s2 on db2197 is CRITICAL: CRITICAL slave_io_state Slave_IO_Running: No, Errno: 2013, Errmsg: error reconnecting to master repl2024@db2207.codfw.wmnet:3306 - retry-time: 60 maximum-retries: 100000 message: Lost connection to server at waiting for initial communication packet, system error: 110 Connection timed out https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [19:50:54] also db2197 ? [19:50:56] PROBLEM - MariaDB Replica IO: s2 #page on db2204 is CRITICAL: CRITICAL slave_io_state Slave_IO_Running: No, Errno: 2013, Errmsg: error reconnecting to master repl2024@db2207.codfw.wmnet:3306 - retry-time: 60 maximum-retries: 100000 message: Lost connection to server at waiting for initial communication packet, system error: 110 Connection timed out https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [19:50:57] PROBLEM - MariaDB Replica IO: s2 #page on db2226 is CRITICAL: CRITICAL slave_io_state Slave_IO_Running: No, Errno: 2013, Errmsg: error reconnecting to master repl2024@db2207.codfw.wmnet:3306 - retry-time: 60 maximum-retries: 100000 message: Lost connection to server at waiting for initial communication packet, system error: 110 Connection timed out https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [19:50:58] (03PS1) 10Gerrit maintenance bot: mariadb: Promote db2204 to s2 master [puppet] - 10https://gerrit.wikimedia.org/r/1311532 (https://phabricator.wikimedia.org/T432396) [19:51:06] PROBLEM - MariaDB Replica IO: s2 #page on db2238 is CRITICAL: CRITICAL slave_io_state Slave_IO_Running: No, Errno: 2013, Errmsg: error reconnecting to master repl2024@db2207.codfw.wmnet:3306 - retry-time: 60 maximum-retries: 100000 message: Lost connection to server at waiting for initial communication packet, system error: 110 Connection timed out https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [19:51:07] PROBLEM - MariaDB Replica IO: s2 #page on db2175 is CRITICAL: CRITICAL slave_io_state Slave_IO_Running: No, Errno: 2013, Errmsg: error reconnecting to master repl2024@db2207.codfw.wmnet:3306 - retry-time: 60 maximum-retries: 100000 message: Lost connection to server at waiting for initial communication packet, system error: 110 Connection timed out https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [19:51:46] PROBLEM - MariaDB Replica IO: s2 #page on db2189 is CRITICAL: CRITICAL slave_io_state Slave_IO_Running: No, Errno: 2003, Errmsg: error reconnecting to master repl2024@db2207.codfw.wmnet:3306 - retry-time: 60 maximum-retries: 100000 message: Cant connect to server on db2207.codfw.wmnet (110 Connection timed out) https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [19:52:58] RECOVERY - Host db2207 #page is UP: PING OK - Packet loss = 0%, RTA = 31.79 ms [19:53:00] PROBLEM - MariaDB Replica IO: s2 #page on db2207 is CRITICAL: CRITICAL slave_io_state could not connect https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [19:53:01] PROBLEM - MariaDB Replica SQL: s2 #page on db2207 is CRITICAL: CRITICAL slave_sql_state could not connect https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [19:53:59] PROBLEM - pt-heartbeat-wikimedia process on db2207 is CRITICAL: PROCS CRITICAL: 0 processes with args pt-heartbeat-wikimedia https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23pt-heartbeat [19:54:06] PROBLEM - MariaDB read only s2 #page on db2207 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Master_comes_back_in_read_only [19:54:06] PROBLEM - MariaDB Event Scheduler s2 on db2207 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Event_Scheduler [19:54:06] PROBLEM - MariaDB Events s2 on db2207 is CRITICAL: CRITICAL - Failed to query events: ERROR 2002 (HY000): Cant connect to local server through socket /run/mysqld/mysqld.sock (2) https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Event_Scheduler [19:54:27] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1034.eqiad.wmnet [19:54:31] !log bking@cumin2003 START - Cookbook sre.hosts.reboot-single for host wdqs1035.eqiad.wmnet [19:54:46] RECOVERY - MariaDB Replica IO: s2 #page on db2189 is OK: OK slave_io_state Slave_IO_Running: Yes https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [19:55:05] RECOVERY - MariaDB Event Scheduler s2 on db2207 is OK: Version 10.11.16-MariaDB-log, Uptime 74s, read_only: True, event_scheduler: True, 24.21 QPS, connection latency: 0.026847s, query latency: 0.001308s https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Event_Scheduler [19:55:06] RECOVERY - MariaDB read only s2 #page on db2207 is OK: Version 10.11.16-MariaDB-log, Uptime 74s, read_only: True, event_scheduler: True, 24.23 QPS, connection latency: 0.027440s, query latency: 0.001081s https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Master_comes_back_in_read_only [19:55:07] RECOVERY - MariaDB Replica IO: s2 #page on db2238 is OK: OK slave_io_state Slave_IO_Running: Yes https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [19:55:07] RECOVERY - MariaDB Events s2 on db2207 is OK: OK - All 2 events in ops database are ENABLED https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Event_Scheduler [19:55:10] RECOVERY - MariaDB Replica IO: s2 #page on db2175 is OK: OK slave_io_state Slave_IO_Running: Yes https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [19:55:32] RECOVERY - MariaDB Replica IO: s2 #page on db2225 is OK: OK slave_io_state Slave_IO_Running: Yes https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [19:55:41] RECOVERY - MariaDB Replica IO: s2 on db2197 is OK: OK slave_io_state Slave_IO_Running: Yes https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [19:55:54] !log marostegui@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: Primary switchover s2 T432396 [19:55:56] RECOVERY - MariaDB Replica IO: s2 #page on db2204 is OK: OK slave_io_state Slave_IO_Running: Yes https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [19:55:57] RECOVERY - MariaDB Replica IO: s2 #page on db2226 is OK: OK slave_io_state Slave_IO_Running: Yes https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [19:55:57] T432396: Switchover s2 master (db2207 -> db2204) - https://phabricator.wikimedia.org/T432396 [19:55:59] RECOVERY - pt-heartbeat-wikimedia process on db2207 is OK: PROCS OK: 1 process with args pt-heartbeat-wikimedia https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23pt-heartbeat [19:56:29] !log marostegui@cumin1003 dbctl commit (dc=all): 'Set db2204 with weight 0 T432396', diff saved to https://phabricator.wikimedia.org/P94891 and previous config saved to /var/cache/conftool/dbconfig/20260716-195628-marostegui.json [19:56:32] 06SRE, 10SRE-Access-Requests: Requesting access to ml-lab-users for ebernhardson - https://phabricator.wikimedia.org/T432345#12130039 (10calbon) I approve [19:56:58] * swfrench-wmf is holding on further shellbox updates until this is resolved [19:57:00] RECOVERY - MariaDB Replica SQL: s2 #page on db2207 is OK: OK slave_sql_state Slave_SQL_Running: Yes https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [19:57:02] RECOVERY - MariaDB Replica IO: s2 #page on db2207 is OK: OK slave_io_state Slave_IO_Running: Yes https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [19:59:16] (03CR) 10Dzahn: [C:03+2] buildkitd: Bump buildkit image to wmf-v0.31.2 [puppet] - 10https://gerrit.wikimedia.org/r/1311523 (https://phabricator.wikimedia.org/T432360) (owner: 10Ahmon Dancy) [19:59:24] (03PS3) 10Btullis: Add zsinger to the analytics-privatedata-users group [puppet] - 10https://gerrit.wikimedia.org/r/1311476 (https://phabricator.wikimedia.org/T426458) [19:59:34] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1035.eqiad.wmnet [20:00:04] RoanKattouw, urbanecm, TheresNoTime, kindrobot, and cjming: Your horoscope predicts another UTC late backport window deploy. May Zuul be (nice) with you. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260716T2000). [20:00:04] aude and sbassett: A patch you scheduled for UTC late backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [20:00:24] (03CR) 10Marostegui: [C:03+2] mariadb: Promote db2204 to s2 master [puppet] - 10https://gerrit.wikimedia.org/r/1311532 (https://phabricator.wikimedia.org/T432396) (owner: 10Gerrit maintenance bot) [20:00:50] !log Starting emergency s2 codfw failover from db2207 to db2204 - T432396 [20:00:53] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [20:01:57] !log marostegui@cumin1003 dbctl commit (dc=all): 'Promote db2204 to s2 primary T432396', diff saved to https://phabricator.wikimedia.org/P94892 and previous config saved to /var/cache/conftool/dbconfig/20260716-200157-marostegui.json [20:02:00] T432396: Switchover s2 master (db2207 -> db2204) - https://phabricator.wikimedia.org/T432396 [20:02:17] (03CR) 10Btullis: [C:04-1] "I think that you need to take down the LVS service first:" [puppet] - 10https://gerrit.wikimedia.org/r/1311530 (https://phabricator.wikimedia.org/T432148) (owner: 10Bking) [20:02:58] !log marostegui@cumin1003 dbctl commit (dc=all): 'Depool db2207 T432396', diff saved to https://phabricator.wikimedia.org/P94893 and previous config saved to /var/cache/conftool/dbconfig/20260716-200257-marostegui.json [20:03:05] (03PS1) 10Bking: datahubsearch: remove references to Puppet plans [puppet] - 10https://gerrit.wikimedia.org/r/1311533 (https://phabricator.wikimedia.org/T432148) [20:03:17] !log btullis@cumin1003 START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. [20:04:00] (03CR) 10Btullis: [C:03+2] Add zsinger to the analytics-privatedata-users group [puppet] - 10https://gerrit.wikimedia.org/r/1311476 (https://phabricator.wikimedia.org/T426458) (owner: 10Btullis) [20:05:14] (03PS2) 10Bking: datahubsearch: remove references to Puppet plans [puppet] - 10https://gerrit.wikimedia.org/r/1311533 (https://phabricator.wikimedia.org/T432148) [20:05:26] !log btullis@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. [20:07:01] !log btullis@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. [20:09:41] !log btullis@cumin1003 END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons. [20:13:09] !log kamila@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" [20:13:14] !log kamila@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003" [20:13:14] !log kamila@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [20:13:14] !log kamila@cumin1003 START - Cookbook sre.dns.wipe-cache wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [20:13:18] !log kamila@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [20:13:19] !log kamila@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1269 [20:13:31] 06SRE, 10SRE-Access-Requests: Requesting access to ml-lab-users for dcausse - https://phabricator.wikimedia.org/T432347#12130119 (10calbon) I approve [20:13:51] RESOLVED: CirrusSearchBackendMemoryIssue: CirrusSearch backend failed 33.333333333333336 times in the last 10 minutes due to memory usage (mw@codfw to dnsdisc) - https://wikitech.wikimedia.org/wiki/Search#Health/Activity_Monitoring - https://grafana.wikimedia.org/d/dc04b9f2-b8d5-4ab6-9482-5d9a75728951/elasticsearch-percentiles?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchBackendMemoryIssue [20:14:26] PROBLEM - Host an-worker1188 is DOWN: PING CRITICAL - Packet loss = 100% [20:15:42] * swfrench-wmf will proceed to deploy remaining shellbox diffs shortly [20:16:41] !log swfrench@deploy2003 helmfile [eqiad] START helmfile.d/services/shellbox: apply [20:17:19] !log swfrench@deploy2003 helmfile [eqiad] DONE helmfile.d/services/shellbox: apply [20:17:50] !log swfrench@deploy2003 helmfile [eqiad] START helmfile.d/services/shellbox-constraints: apply [20:18:19] !log swfrench@deploy2003 helmfile [eqiad] DONE helmfile.d/services/shellbox-constraints: apply [20:18:50] !log swfrench@deploy2003 helmfile [eqiad] START helmfile.d/services/shellbox-media: apply [20:19:12] !log swfrench@deploy2003 helmfile [eqiad] DONE helmfile.d/services/shellbox-media: apply [20:19:32] (03PS1) 10Marostegui: db2207: Disable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1311535 (https://phabricator.wikimedia.org/T432398) [20:19:43] !log swfrench@deploy2003 helmfile [eqiad] START helmfile.d/services/shellbox-syntaxhighlight: apply [20:20:04] !log swfrench@deploy2003 helmfile [eqiad] DONE helmfile.d/services/shellbox-syntaxhighlight: apply [20:20:32] (03CR) 10Marostegui: [C:03+2] db2207: Disable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1311535 (https://phabricator.wikimedia.org/T432398) (owner: 10Marostegui) [20:20:35] !log swfrench@deploy2003 helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply [20:20:59] !log swfrench@deploy2003 helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply [20:21:30] !log swfrench@deploy2003 helmfile [eqiad] START helmfile.d/services/shellbox-video: apply [20:21:38] i am late for the backport window, but here for my backport (i can do it whenever) [20:22:19] kamila@cumin1003 renumber-node (PID 451981) is awaiting input [20:22:32] aude: do you need a deployer? [20:22:37] !log swfrench@deploy2003 helmfile [eqiad] DONE helmfile.d/services/shellbox-video: apply [20:22:51] i can deploy mine and don't see anything in progress [20:23:27] but saw sbassett had backports also (and idk if someone is handling those) [20:24:37] Nothing is being backported at the moment and I haven't seen them here on the channel so I think it's okay for you to go ahead [20:25:09] ok i am proceeding [20:25:58] (03CR) 10TrainBranchBot: [C:03+2] "Approved by aude@deploy2003 using scap backport" [skins/Vector] (wmf/1.47.0-wmf.11) - 10https://gerrit.wikimedia.org/r/1311472 (https://phabricator.wikimedia.org/T432316) (owner: 10Aude) [20:26:13] !log marostegui@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down [20:26:25] !incidents [20:26:25] 8210 (RESOLVED) db2207 (paged)/MariaDB Replica IO: s2 (paged) [20:26:26] 8211 (RESOLVED) db2207 (paged)/MariaDB Replica SQL: s2 (paged) [20:26:26] 8206 (RESOLVED) db2204 (paged)/MariaDB Replica IO: s2 (paged) [20:26:26] 8205 (RESOLVED) db2226 (paged)/MariaDB Replica IO: s2 (paged) [20:26:26] 8204 (RESOLVED) db2225 (paged)/MariaDB Replica IO: s2 (paged) [20:26:26] 8208 (RESOLVED) db2238 (paged)/MariaDB Replica IO: s2 (paged) [20:26:27] 8207 (RESOLVED) db2175 (paged)/MariaDB Replica IO: s2 (paged) [20:26:27] 8212 (RESOLVED) db2207 (paged)/MariaDB read only s2 (paged) [20:26:27] 8209 (RESOLVED) db2189 (paged)/MariaDB Replica IO: s2 (paged) [20:26:28] 8203 (RESOLVED) Host db2207 (paged) [20:26:28] 8202 (RESOLVED) NELHigh sre (thanos-rule@main tcp.timed_out) [20:26:29] 8201 (RESOLVED) Host cr1-drmrs [20:28:54] (03PS1) 10Krinkle: Set $wgMathInternalRestbaseURL explicitly (take 2) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311536 (https://phabricator.wikimedia.org/T349582) [20:32:51] 06SRE, 06Data-Platform-SRE: Make the shell group analytics-privatedata-users less confusing - https://phabricator.wikimedia.org/T405517#12130162 (10BTullis) It may be helpful to know that the #data-engineering and #data-platform-sre teams are currently working on a project to migrate our Data Lake away from Ha... [20:33:24] (03Merged) 10jenkins-bot: Preserve menus after toolbox (e.g. print/export) in page tools [skins/Vector] (wmf/1.47.0-wmf.11) - 10https://gerrit.wikimedia.org/r/1311472 (https://phabricator.wikimedia.org/T432316) (owner: 10Aude) [20:33:36] !log aude@deploy2003 Started scap sync-world: Backport for [[gerrit:1311472|Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] [20:33:40] T432316: Print/export section missing in 1.47.0-wmf.11 - https://phabricator.wikimedia.org/T432316 [20:35:16] !log aude@deploy2003 aude: Backport for [[gerrit:1311472|Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [20:35:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [20:36:50] !log aude@deploy2003 aude: Continuing with deployment [20:38:10] (03PS1) 10Cathal Mooney: Arelion transport: enable OSPF from cr2-drmrs to cr2-eqiad [homer/public] - 10https://gerrit.wikimedia.org/r/1311537 (https://phabricator.wikimedia.org/T424839) [20:39:51] !log kamila@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1269 [20:39:51] !log kamila@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1269 [20:39:59] (03CR) 10Ayounsi: [C:03+1] Arelion transport: enable OSPF from cr2-drmrs to cr2-eqiad [homer/public] - 10https://gerrit.wikimedia.org/r/1311537 (https://phabricator.wikimedia.org/T424839) (owner: 10Cathal Mooney) [20:41:08] (03CR) 10Cathal Mooney: [C:03+2] Arelion transport: enable OSPF from cr2-drmrs to cr2-eqiad [homer/public] - 10https://gerrit.wikimedia.org/r/1311537 (https://phabricator.wikimedia.org/T424839) (owner: 10Cathal Mooney) [20:41:10] !log aude@deploy2003 Finished scap sync-world: Backport for [[gerrit:1311472|Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] (duration: 07m 34s) [20:41:14] T432316: Print/export section missing in 1.47.0-wmf.11 - https://phabricator.wikimedia.org/T432316 [20:42:24] (03Merged) 10jenkins-bot: Arelion transport: enable OSPF from cr2-drmrs to cr2-eqiad [homer/public] - 10https://gerrit.wikimedia.org/r/1311537 (https://phabricator.wikimedia.org/T424839) (owner: 10Cathal Mooney) [20:43:06] (03PS1) 10Bking: datahubsearch: remove DNS record [dns] - 10https://gerrit.wikimedia.org/r/1311538 (https://phabricator.wikimedia.org/T432148) [20:45:18] (03CR) 10JHathaway: [C:03+2] "I did another manual review, looks good, I'm comfortable merging as is." [puppet] - 10https://gerrit.wikimedia.org/r/1305986 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [20:46:48] btullis@cumin1003 reboot-workers (PID 394778) is awaiting input [20:46:56] !log swfrench@cumin1003 START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1271.eqiad.wmnet [20:46:59] !log swfrench@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1271.eqiad.wmnet [20:47:35] !log swfrench@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1271.eqiad.wmnet [20:48:04] !log swfrench@cumin1003 START - Cookbook sre.hosts.reimage for host wikikube-worker1271.eqiad.wmnet with OS trixie [20:48:33] !log swfrench@cumin1003 START - Cookbook sre.hosts.move-vlan for host wikikube-worker1271 [20:49:01] !log kamila@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1268.eqiad.wmnet [20:49:02] !log kamila@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1268.eqiad.wmnet [20:49:04] !log kamila@cumin1003 END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1268.eqiad.wmnet [20:49:33] !log swfrench@cumin1003 START - Cookbook sre.dns.netbox [20:50:55] !log arlolra@deploy2003 helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply [20:51:25] !log arlolra@deploy2003 helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply [20:51:26] !log arlolra@deploy2003 helmfile [codfw] START helmfile.d/services/mw-parsoid: apply [20:51:53] !log arlolra@deploy2003 helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply [20:51:57] (03PS3) 10Bking: datahubsearch: move to 'insetup' role [puppet] - 10https://gerrit.wikimedia.org/r/1311530 (https://phabricator.wikimedia.org/T432148) [20:54:22] !log swfrench@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" [20:54:27] !log swfrench@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003" [20:54:27] !log swfrench@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [20:54:27] !log swfrench@cumin1003 START - Cookbook sre.dns.wipe-cache wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [20:54:30] !log swfrench@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [20:54:31] !log swfrench@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1271 [20:55:35] (03PS1) 10BCornwall: varnish: Add SPDX license header to VCL files [puppet] - 10https://gerrit.wikimedia.org/r/1311540 [20:55:35] (03PS1) 10BCornwall: varnish: Remove nonfunctional vim modelines [puppet] - 10https://gerrit.wikimedia.org/r/1311541 [20:55:59] !log swfrench@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1271 [20:55:59] !log swfrench@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1271 [20:59:21] (03PS2) 10BCornwall: varnish: Remove nonfunctional vim modelines [puppet] - 10https://gerrit.wikimedia.org/r/1311541 [21:00:05] Deploy window Readers deployment window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260716T2100) [21:00:13] (03PS1) 10Bking: opensearch: remove references to soon-to-be-decom'd service [alerts] - 10https://gerrit.wikimedia.org/r/1311542 (https://phabricator.wikimedia.org/T432148) [21:00:19] (03CR) 10BCornwall: [C:03+1] datahubsearch: remove DNS record [dns] - 10https://gerrit.wikimedia.org/r/1311538 (https://phabricator.wikimedia.org/T432148) (owner: 10Bking) [21:00:35] !log kamila@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage [21:03:45] (03CR) 10CDobbins: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1311540 (owner: 10BCornwall) [21:04:24] !log kamila@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage [21:06:44] jeena aude - sorry, was massively distracted. I can deploy my two patches now as long as nobody is using the readers web window... [21:07:03] I am done with mine [21:07:48] ok [21:10:26] (03CR) 10TrainBranchBot: [C:03+2] "Approved by sbassett@deploy2003 using scap backport" [core] (wmf/1.47.0-wmf.11) - 10https://gerrit.wikimedia.org/r/1311481 (https://phabricator.wikimedia.org/T432042) (owner: 10SBassett) [21:10:26] (03CR) 10TrainBranchBot: [C:03+2] "Approved by sbassett@deploy2003 using scap backport" [core] (wmf/1.47.0-wmf.11) - 10https://gerrit.wikimedia.org/r/1311482 (https://phabricator.wikimedia.org/T432042) (owner: 10SBassett) [21:11:03] (03CR) 10BCornwall: [V:03+1] "PCC SUCCESS (NOOP 2): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9025/console" [puppet] - 10https://gerrit.wikimedia.org/r/1311540 (owner: 10BCornwall) [21:13:31] (03CR) 10BCornwall: [V:03+1] "PCC SUCCESS (CORE_DIFF 2): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9026/co" [puppet] - 10https://gerrit.wikimedia.org/r/1311541 (owner: 10BCornwall) [21:13:34] (03CR) 10JHathaway: "yeah I agree, it will be easier to manage that way." [puppet] - 10https://gerrit.wikimedia.org/r/1305989 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:14:27] (03Abandoned) 10JHathaway: Puppet 8: Replace legacy facts [puppet] - 10https://gerrit.wikimedia.org/r/1305989 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:16:28] (03PS1) 10Eevans: cassandra: introduce Cassandra 5.0.x configuration [puppet] - 10https://gerrit.wikimedia.org/r/1311545 (https://phabricator.wikimedia.org/T418419) [21:17:11] (03CR) 10CI reject: [V:04-1] cassandra: introduce Cassandra 5.0.x configuration [puppet] - 10https://gerrit.wikimedia.org/r/1311545 (https://phabricator.wikimedia.org/T418419) (owner: 10Eevans) [21:17:15] !log swfrench@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage [21:18:34] (03Merged) 10jenkins-bot: Cleanup: remove reauth indicator from log message [core] (wmf/1.47.0-wmf.11) - 10https://gerrit.wikimedia.org/r/1311481 (https://phabricator.wikimedia.org/T432042) (owner: 10SBassett) [21:18:40] (03CR) 10CI reject: [V:04-1] Use non-sampled authentication log channel instead of authevents [core] (wmf/1.47.0-wmf.11) - 10https://gerrit.wikimedia.org/r/1311482 (https://phabricator.wikimedia.org/T432042) (owner: 10SBassett) [21:19:59] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host wdqs2025.codfw.wmnet with OS bookworm [21:20:27] !log bking@cumin2003 START - Cookbook sre.hosts.move-vlan for host wdqs2025 [21:21:46] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in module alertmanager [puppet] - 10https://gerrit.wikimedia.org/r/1311546 (https://phabricator.wikimedia.org/T372666) [21:21:59] (03PS1) 10Lerickson: Add main/scholarly and internal/external specific configs for EventGate streams. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311547 (https://phabricator.wikimedia.org/T429407) [21:22:08] !log sbassett@deploy2003 Started scap sync-world: Backport for [[gerrit:1311481|Cleanup: remove reauth indicator from log message (T432042)]] [21:22:12] T432042: Add logging for various re-authentication methods (AuthPopup vs DataStash flow) - https://phabricator.wikimedia.org/T432042 [21:22:13] !log bking@cumin2003 START - Cookbook sre.dns.netbox [21:22:29] (03PS2) 10Eevans: cassandra: introduce Cassandra 5.0.x configuration [puppet] - 10https://gerrit.wikimedia.org/r/1311545 (https://phabricator.wikimedia.org/T418419) [21:22:29] (03PS1) 10Eevans: cassandra-dev2001: upgrade to Cassandra 5.0.8 [puppet] - 10https://gerrit.wikimedia.org/r/1311548 (https://phabricator.wikimedia.org/T418419) [21:22:41] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in module graphite [puppet] - 10https://gerrit.wikimedia.org/r/1311549 (https://phabricator.wikimedia.org/T372666) [21:23:11] (03CR) 10CI reject: [V:04-1] cassandra: introduce Cassandra 5.0.x configuration [puppet] - 10https://gerrit.wikimedia.org/r/1311545 (https://phabricator.wikimedia.org/T418419) (owner: 10Eevans) [21:23:15] (03PS2) 10Eevans: cassandra-dev2001: upgrade to Cassandra 5.0.8 [puppet] - 10https://gerrit.wikimedia.org/r/1311548 (https://phabricator.wikimedia.org/T418419) [21:23:16] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in module icinga [puppet] - 10https://gerrit.wikimedia.org/r/1311550 (https://phabricator.wikimedia.org/T372666) [21:23:31] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in module logstash [puppet] - 10https://gerrit.wikimedia.org/r/1311551 (https://phabricator.wikimedia.org/T372666) [21:23:37] (03CR) 10Eevans: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1311548 (https://phabricator.wikimedia.org/T418419) (owner: 10Eevans) [21:23:51] !log sbassett@deploy2003 sbassett: Backport for [[gerrit:1311481|Cleanup: remove reauth indicator from log message (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [21:24:01] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in module monitoring [puppet] - 10https://gerrit.wikimedia.org/r/1311552 (https://phabricator.wikimedia.org/T372666) [21:24:28] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in module prometheus [puppet] - 10https://gerrit.wikimedia.org/r/1311553 (https://phabricator.wikimedia.org/T372666) [21:24:38] !log swfrench@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage [21:24:41] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in module statsite [puppet] - 10https://gerrit.wikimedia.org/r/1311554 (https://phabricator.wikimedia.org/T372666) [21:24:43] !log kamila@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1269.eqiad.wmnet with OS trixie [21:24:55] (03PS2) 10Lerickson: Add main/scholarly and internal/external specific configs for EventGate streams. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311547 (https://phabricator.wikimedia.org/T429407) [21:25:08] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in profile monitoring [puppet] - 10https://gerrit.wikimedia.org/r/1311555 (https://phabricator.wikimedia.org/T372666) [21:25:38] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in profile alertmanager [puppet] - 10https://gerrit.wikimedia.org/r/1311556 (https://phabricator.wikimedia.org/T372666) [21:26:11] !log sbassett@deploy2003 sbassett: Continuing with deployment [21:26:17] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in profile prometheus [puppet] - 10https://gerrit.wikimedia.org/r/1311558 (https://phabricator.wikimedia.org/T372666) [21:26:27] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in profile corto [puppet] - 10https://gerrit.wikimedia.org/r/1311559 (https://phabricator.wikimedia.org/T372666) [21:26:36] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in profile grafana [puppet] - 10https://gerrit.wikimedia.org/r/1311560 (https://phabricator.wikimedia.org/T372666) [21:26:43] (03CR) 10CI reject: [V:04-1] Puppet 8: Replace legacy facts in module monitoring [puppet] - 10https://gerrit.wikimedia.org/r/1311552 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:26:46] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in profile graphite [puppet] - 10https://gerrit.wikimedia.org/r/1311561 (https://phabricator.wikimedia.org/T372666) [21:26:59] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in profile icinga [puppet] - 10https://gerrit.wikimedia.org/r/1311562 (https://phabricator.wikimedia.org/T372666) [21:27:31] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in profile logstash [puppet] - 10https://gerrit.wikimedia.org/r/1311563 (https://phabricator.wikimedia.org/T372666) [21:27:46] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in profile statsd [puppet] - 10https://gerrit.wikimedia.org/r/1311564 (https://phabricator.wikimedia.org/T372666) [21:28:00] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in profile syslog [puppet] - 10https://gerrit.wikimedia.org/r/1311565 (https://phabricator.wikimedia.org/T372666) [21:28:10] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in profile pontoon [puppet] - 10https://gerrit.wikimedia.org/r/1311566 (https://phabricator.wikimedia.org/T372666) [21:28:21] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in profile netconsole [puppet] - 10https://gerrit.wikimedia.org/r/1311567 (https://phabricator.wikimedia.org/T372666) [21:28:28] bking@cumin2003 reimage (PID 2098710) is awaiting input [21:28:30] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in role alerting_host [puppet] - 10https://gerrit.wikimedia.org/r/1311568 (https://phabricator.wikimedia.org/T372666) [21:29:23] (03PS3) 10SBassett: Use non-sampled authentication log channel instead of authevents [core] (wmf/1.47.0-wmf.11) - 10https://gerrit.wikimedia.org/r/1311482 (https://phabricator.wikimedia.org/T432042) [21:29:29] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1311549 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:29:31] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1311551 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:29:34] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1311550 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:29:36] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1311546 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:29:38] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1311554 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:29:40] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1311553 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:29:42] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1311558 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:29:44] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1311559 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:29:47] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1311560 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:29:51] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1311552 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:29:56] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1311561 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:30:00] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1311562 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:30:04] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1311563 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:30:08] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1311555 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:30:12] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1311564 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:30:16] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1311565 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:30:20] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1311566 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:30:24] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1311567 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:30:28] !log sbassett@deploy2003 Finished scap sync-world: Backport for [[gerrit:1311481|Cleanup: remove reauth indicator from log message (T432042)]] (duration: 08m 19s) [21:30:28] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1311556 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:30:31] T432042: Add logging for various re-authentication methods (AuthPopup vs DataStash flow) - https://phabricator.wikimedia.org/T432042 [21:30:32] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1311568 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:31:11] (03PS3) 10Eevans: cassandra-dev2001: upgrade to Cassandra 5.0.8 [puppet] - 10https://gerrit.wikimedia.org/r/1311548 (https://phabricator.wikimedia.org/T418419) [21:31:32] (03CR) 10TrainBranchBot: [C:03+2] "Approved by sbassett@deploy2003 using scap backport" [core] (wmf/1.47.0-wmf.11) - 10https://gerrit.wikimedia.org/r/1311482 (https://phabricator.wikimedia.org/T432042) (owner: 10SBassett) [21:31:42] (03CR) 10JHathaway: "recheck" [puppet] - 10https://gerrit.wikimedia.org/r/1305984 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:37:05] !log bking@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" [21:37:10] !log bking@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003" [21:37:10] !log bking@cumin2003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [21:37:11] !log bking@cumin2003 START - Cookbook sre.dns.wipe-cache wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [21:37:14] !log bking@cumin2003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [21:37:15] !log bking@cumin2003 START - Cookbook sre.network.configure-switch-interfaces for host wdqs2025 [21:38:18] (03Merged) 10jenkins-bot: Use non-sampled authentication log channel instead of authevents [core] (wmf/1.47.0-wmf.11) - 10https://gerrit.wikimedia.org/r/1311482 (https://phabricator.wikimedia.org/T432042) (owner: 10SBassett) [21:38:29] !log sbassett@deploy2003 Started scap sync-world: Backport for [[gerrit:1311482|Use non-sampled authentication log channel instead of authevents (T432042)]] [21:38:33] T432042: Add logging for various re-authentication methods (AuthPopup vs DataStash flow) - https://phabricator.wikimedia.org/T432042 [21:40:13] !log sbassett@deploy2003 sbassett: Backport for [[gerrit:1311482|Use non-sampled authentication log channel instead of authevents (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [21:40:19] bking@cumin2003 reimage (PID 2098710) is awaiting input [21:40:32] !log bking@cumin2003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2025 [21:40:32] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2025 [21:40:49] !log sbassett@deploy2003 sbassett: Continuing with deployment [21:42:36] (03PS3) 10Eevans: cassandra: introduce Cassandra 5.0.x configuration [puppet] - 10https://gerrit.wikimedia.org/r/1311545 (https://phabricator.wikimedia.org/T418419) [21:42:36] (03PS4) 10Eevans: cassandra-dev2001: upgrade to Cassandra 5.0.8 [puppet] - 10https://gerrit.wikimedia.org/r/1311548 (https://phabricator.wikimedia.org/T418419) [21:43:17] (03CR) 10CI reject: [V:04-1] cassandra: introduce Cassandra 5.0.x configuration [puppet] - 10https://gerrit.wikimedia.org/r/1311545 (https://phabricator.wikimedia.org/T418419) (owner: 10Eevans) [21:43:38] kamila@cumin1003 renumber-node (PID 468986) is awaiting input [21:45:00] !log sbassett@deploy2003 Finished scap sync-world: Backport for [[gerrit:1311482|Use non-sampled authentication log channel instead of authevents (T432042)]] (duration: 06m 31s) [21:45:04] T432042: Add logging for various re-authentication methods (AuthPopup vs DataStash flow) - https://phabricator.wikimedia.org/T432042 [21:46:19] !log swfrench@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1271.eqiad.wmnet with OS trixie [21:47:52] Ok, I should be done, thanks. [21:49:22] (03CR) 10Eevans: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1311548 (https://phabricator.wikimedia.org/T418419) (owner: 10Eevans) [21:55:42] !log swfrench@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1271.eqiad.wmnet [21:55:43] !log swfrench@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1271.eqiad.wmnet [21:55:44] !log swfrench@cumin1003 END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1271.eqiad.wmnet [21:57:14] (03PS2) 10JHathaway: Puppet 8: Replace legacy facts in module graphite [puppet] - 10https://gerrit.wikimedia.org/r/1311549 (https://phabricator.wikimedia.org/T372666) [21:59:28] !log swfrench@cumin1003 START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1272.eqiad.wmnet [21:59:32] !log swfrench@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1272.eqiad.wmnet [22:00:04] !log swfrench@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1272.eqiad.wmnet [22:00:26] !log swfrench@cumin1003 START - Cookbook sre.hosts.reimage for host wikikube-worker1272.eqiad.wmnet with OS trixie [22:00:38] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1311549 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [22:00:54] !log swfrench@cumin1003 START - Cookbook sre.hosts.move-vlan for host wikikube-worker1272 [22:01:15] !log swfrench@cumin1003 START - Cookbook sre.dns.netbox [22:03:33] !log bking@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage [22:03:54] jouncebot: nowandnext [22:03:54] No deployments scheduled for the next 7 hour(s) and 56 minute(s) [22:03:54] In 7 hour(s) and 56 minute(s): MediaWiki infrastructure (UTC early) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260717T0600) [22:05:09] (03CR) 10Ladsgroup: [C:03+2] Upload: Do not throw for failure to save a chunk file [core] (wmf/1.47.0-wmf.11) - 10https://gerrit.wikimedia.org/r/1311420 (https://phabricator.wikimedia.org/T430986) (owner: 10Ladsgroup) [22:05:42] !log swfrench@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" [22:05:46] !log swfrench@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003" [22:05:46] !log swfrench@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [22:05:46] !log swfrench@cumin1003 START - Cookbook sre.dns.wipe-cache wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [22:05:50] !log swfrench@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [22:05:51] !log swfrench@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1272 [22:06:43] !log swfrench@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1272 [22:06:43] !log swfrench@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1272 [22:06:57] jouncebot: nowandnext [22:06:57] No deployments scheduled for the next 7 hour(s) and 53 minute(s) [22:06:57] In 7 hour(s) and 53 minute(s): MediaWiki infrastructure (UTC early) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260717T0600) [22:07:18] (03CR) 10Urbanecm: [C:03+2] AddImage: Request only standard thumbnail sizes [extensions/GrowthExperiments] (wmf/1.47.0-wmf.11) - 10https://gerrit.wikimedia.org/r/1311529 (https://phabricator.wikimedia.org/T428797) (owner: 10Urbanecm) [22:08:17] Amir1: failed to notice your +2 earlier... [22:08:32] hehehe [22:08:37] ...should i cancel? i can also test together with yours if you're comfortable [22:08:38] don't worry. Mine is pretty minor [22:08:45] ack ack [22:08:50] sounds good to me [22:09:14] mine is quite hard to test [22:09:42] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage [22:13:49] (03Merged) 10jenkins-bot: Upload: Do not throw for failure to save a chunk file [core] (wmf/1.47.0-wmf.11) - 10https://gerrit.wikimedia.org/r/1311420 (https://phabricator.wikimedia.org/T430986) (owner: 10Ladsgroup) [22:13:55] (03Merged) 10jenkins-bot: AddImage: Request only standard thumbnail sizes [extensions/GrowthExperiments] (wmf/1.47.0-wmf.11) - 10https://gerrit.wikimedia.org/r/1311529 (https://phabricator.wikimedia.org/T428797) (owner: 10Urbanecm) [22:15:30] Amir1: are you calling the scap, or should i? [22:15:38] can do [22:15:43] ok, waiting [22:17:32] !log ladsgroup@deploy2003 Started scap sync-world: Backport for [[gerrit:1311420|Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529|AddImage: Request only standard thumbnail sizes (T428797)]] [22:17:37] T430986: MediaWiki\Upload\Exception\UploadChunkFileException: Error storing file in '{chunkPath}': backend-fail-internal; local-swift-codfw - https://phabricator.wikimedia.org/T430986 [22:17:37] T428797: [wmf.5-regression] Add image dialog doesn't display image - https://phabricator.wikimedia.org/T428797 [22:18:10] hehe I broke that [22:18:17] sowwy [22:18:40] now you're fixing it, so...credit earned ;) [22:19:19] !log ladsgroup@deploy2003 ladsgroup, urbanecm: Backport for [[gerrit:1311420|Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529|AddImage: Request only standard thumbnail sizes (T428797)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [22:19:27] testing mine [22:20:03] starting the mwdebug extension might help with that... [22:20:30] haha [22:21:08] and after a couple of hard refreshes, it work [22:21:11] *works [22:21:12] lgtm :) [22:22:07] !log ladsgroup@deploy2003 ladsgroup, urbanecm: Continuing with deployment [22:22:13] okay dokie, pushing forward [22:26:23] !log ladsgroup@deploy2003 Finished scap sync-world: Backport for [[gerrit:1311420|Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529|AddImage: Request only standard thumbnail sizes (T428797)]] (duration: 08m 51s) [22:26:28] T430986: MediaWiki\Upload\Exception\UploadChunkFileException: Error storing file in '{chunkPath}': backend-fail-internal; local-swift-codfw - https://phabricator.wikimedia.org/T430986 [22:26:29] T428797: [wmf.5-regression] Add image dialog doesn't display image - https://phabricator.wikimedia.org/T428797 [22:27:29] !log swfrench@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage [22:28:01] !log kamila@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1269.eqiad.wmnet [22:28:02] !log kamila@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1269.eqiad.wmnet [22:28:04] !log kamila@cumin1003 END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1269.eqiad.wmnet [22:32:02] !log bking@deploy2003 Started deploy [wdqs/wdqs@e8fb00c]: T430880 [22:32:07] T430880: Migrate WDQS hosts to Bookworm or later - https://phabricator.wikimedia.org/T430880 [22:32:08] !log bking@deploy2003 Finished deploy [wdqs/wdqs@e8fb00c]: T430880 (duration: 00m 27s) [22:35:08] !log swfrench@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage [22:39:53] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2025.codfw.wmnet with OS bookworm [22:53:48] FIRING: KubernetesCalicoDown: wikikube-worker1264.eqiad.wmnet is not running calico-node Pod - https://wikitech.wikimedia.org/wiki/Calico#Operations - https://grafana.wikimedia.org/d/G8zPL7-Wz/?var-dc=eqiad%20prometheus%2Fk8s&var-instance=wikikube-worker1264.eqiad.wmnet - https://alerts.wikimedia.org/?q=alertname%3DKubernetesCalicoDown [22:56:52] !log deleting echo notifications from 2015 in group0 [22:56:53] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [22:57:04] FIRING: HelmReleaseBadStatus: Helm release wdqs/scholarly-external on k8s-dse@codfw in state pending-rollback - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [22:57:28] !log swfrench@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1272.eqiad.wmnet with OS trixie [22:59:36] (03PS1) 10Dzahn: Revert "gerrit: add SSH config for replication between servers" [puppet] - 10https://gerrit.wikimedia.org/r/1311577 [23:01:08] !log ryankemper@cumin2003 START - Cookbook sre.wdqs.data-transfer (T430880, xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards [23:01:12] T430880: Migrate WDQS hosts to Bookworm or later - https://phabricator.wikimedia.org/T430880 [23:01:33] (03PS2) 10Dzahn: Revert "gerrit: add SSH config for replication between servers" [puppet] - 10https://gerrit.wikimedia.org/r/1311577 (https://phabricator.wikimedia.org/T432413) [23:02:00] (03CR) 10Dzahn: [C:03+2] Revert "gerrit: add SSH config for replication between servers" [puppet] - 10https://gerrit.wikimedia.org/r/1311577 (https://phabricator.wikimedia.org/T432413) (owner: 10Dzahn) [23:02:25] (03CR) 10Dzahn: [V:03+2 C:03+2] Revert "gerrit: add SSH config for replication between servers" [puppet] - 10https://gerrit.wikimedia.org/r/1311577 (https://phabricator.wikimedia.org/T432413) (owner: 10Dzahn) [23:05:39] FIRING: CertAlmostExpired: gNMI TLS certificate for lsw1-f6-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [23:08:40] !log swfrench@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1272.eqiad.wmnet [23:08:41] !log swfrench@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1272.eqiad.wmnet [23:08:44] !log swfrench@cumin1003 END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1272.eqiad.wmnet [23:09:13] !log ryankemper@cumin2003 conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad [23:12:17] !log T430880 depooled dnsdisc of wdqs-internal-scholarly-eqiad bc we only have 1 host there [23:12:20] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [23:12:21] T430880: Migrate WDQS hosts to Bookworm or later - https://phabricator.wikimedia.org/T430880 [23:12:51] !log swfrench@cumin1003 START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1274.eqiad.wmnet [23:12:54] !log swfrench@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1274.eqiad.wmnet [23:13:26] !log swfrench@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1274.eqiad.wmnet [23:13:57] !log swfrench@cumin1003 START - Cookbook sre.hosts.reimage for host wikikube-worker1274.eqiad.wmnet with OS trixie [23:14:26] !log swfrench@cumin1003 START - Cookbook sre.hosts.move-vlan for host wikikube-worker1274 [23:14:36] !log swfrench@cumin1003 START - Cookbook sre.dns.netbox [23:19:00] !log ryankemper@cumin2003 START - Cookbook sre.hosts.reimage for host wdqs1027.eqiad.wmnet with OS bookworm [23:19:01] !log swfrench@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" [23:19:05] !log swfrench@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003" [23:19:05] !log swfrench@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [23:19:05] !log swfrench@cumin1003 START - Cookbook sre.dns.wipe-cache wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [23:19:09] !log swfrench@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [23:19:10] !log swfrench@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1274 [23:19:37] !log ryankemper@cumin2003 START - Cookbook sre.hosts.reimage for host wdqs1025.eqiad.wmnet with OS bookworm [23:19:45] !log swfrench@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1274 [23:19:45] !log swfrench@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1274 [23:22:23] !log ryankemper@cumin2003 START - Cookbook sre.hosts.move-vlan for host wdqs1027 [23:22:23] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1027 [23:23:01] !log ryankemper@cumin2003 START - Cookbook sre.hosts.move-vlan for host wdqs1025 [23:23:01] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1025 [23:23:10] PROBLEM - PyBal IPVS diff check on lvs1019 is CRITICAL: (CRITICAL: Mismatch between IPVS and PyBal https://wikitech.wikimedia.org/wiki/PyBal [23:25:58] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [23:26:09] PROBLEM - PyBal IPVS diff check on lvs1020 is CRITICAL: (CRITICAL: Mismatch between IPVS and PyBal https://wikitech.wikimedia.org/wiki/PyBal [23:26:41] FIRING: [2x] ConfdResourceFailed: confd resource _srv_config-master_pybal_eqiad_wdqs-internal-scholarly.toml has errors - https://wikitech.wikimedia.org/wiki/Confd#Monitoring - https://grafana.wikimedia.org/d/OUJF1VI4k/confd - https://alerts.wikimedia.org/?q=alertname%3DConfdResourceFailed [23:38:37] !log swfrench@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage [23:39:41] !log ryankemper@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage [23:41:47] !log ryankemper@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage [23:42:33] (03PS1) 10TrainBranchBot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1311579 [23:42:33] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1311579 (owner: 10TrainBranchBot) [23:43:55] !log swfrench@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage [23:47:32] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage [23:50:08] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage [23:50:15] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) (T430880, xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards [23:50:18] T430880: Migrate WDQS hosts to Bookworm or later - https://phabricator.wikimedia.org/T430880 [23:50:48] (03Merged) 10jenkins-bot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1311579 (owner: 10TrainBranchBot)