[00:18:11] (03PS1) 10Raymond Ndibe: webservice-runner: default to 8000 only if PORT and TOOL_WEB_PORT has no value [docker-images/toollabs-images] - 10https://gerrit.wikimedia.org/r/1310212 (https://phabricator.wikimedia.org/T432078) [00:28:36] FIRING: [5x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-3/1/2 (Transport: Hurricane Electric (dc4841.dal5) {#122601}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [00:35:42] FIRING: [2x] JobUnavailable: Reduced availability for job atlas_exporter in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [00:40:43] FIRING: [2x] JobUnavailable: Reduced availability for job atlas_exporter in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [01:10:58] (03PS1) 10TrainBranchBot: Branch commit for wmf/1.47.0-wmf.11 [core] (wmf/1.47.0-wmf.11) - 10https://gerrit.wikimedia.org/r/1310219 (https://phabricator.wikimedia.org/T430830) [01:11:01] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/1.47.0-wmf.11 [core] (wmf/1.47.0-wmf.11) - 10https://gerrit.wikimedia.org/r/1310219 (https://phabricator.wikimedia.org/T430830) (owner: 10TrainBranchBot) [01:12:07] (03PS1) 10TrainBranchBot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1310220 [01:12:07] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1310220 (owner: 10TrainBranchBot) [01:19:19] (03Merged) 10jenkins-bot: Branch commit for wmf/1.47.0-wmf.11 [core] (wmf/1.47.0-wmf.11) - 10https://gerrit.wikimedia.org/r/1310219 (https://phabricator.wikimedia.org/T430830) (owner: 10TrainBranchBot) [01:20:07] (03Merged) 10jenkins-bot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1310220 (owner: 10TrainBranchBot) [02:00:05] Deploy window Automatic branching of MediaWiki, extensions, skins, and vendor – see Heterogeneous deployment/Train deploys (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260714T0200) [02:00:29] !log mwpresync@deploy2003 Started scap build-images: Publishing wmf/next image [02:06:58] !log mwpresync@deploy2003 Finished scap build-images: Publishing wmf/next image (duration: 06m 29s) [02:10:42] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:15:42] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [03:00:05] Deploy window Automatic deployment of MediaWiki, extensions, skins, and vendor to testwikis only – see Heterogeneous deployment/Train deploys (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260714T0300) [03:01:47] (03PS1) 10TrainBranchBot: testwikis to 1.47.0-wmf.11 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310237 (https://phabricator.wikimedia.org/T430830) [03:01:50] (03CR) 10TrainBranchBot: [C:03+2] "Initiated by mwpresync@deploy2003" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310237 (https://phabricator.wikimedia.org/T430830) (owner: 10TrainBranchBot) [03:02:46] FIRING: Not accepting/receiving prefixes from anycast BGP peer: Alert for device asw1-b4-magru.mgmt.magru.wmnet - Not accepting/receiving prefixes from anycast BGP peer - https://alerts.wikimedia.org/?q=alertname%3DNot+accepting%2Freceiving+prefixes+from+anycast+BGP+peer [03:02:51] (03Merged) 10jenkins-bot: testwikis to 1.47.0-wmf.11 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310237 (https://phabricator.wikimedia.org/T430830) (owner: 10TrainBranchBot) [03:03:08] !log mwpresync@deploy2003 Started scap sync-world: testwikis to 1.47.0-wmf.11 refs T430830 [03:03:11] T430830: 1.47.0-wmf.11 deployment blockers - https://phabricator.wikimedia.org/T430830 [03:10:38] FIRING: [2x] CertAlmostExpired: gNMI TLS certificate for lsw1-d2-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [03:31:39] FIRING: CoreBGPDown: Core BGP session down between cr1-drmrs and cr2-eqiad (185.15.58.138) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=drmrs&var-device=cr1-drmrs:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [03:36:31] FIRING: Traffic on tunnel link: Alert for device cr1-drmrs.wikimedia.org - Traffic on tunnel link - https://alerts.wikimedia.org/?q=alertname%3DTraffic+on+tunnel+link [03:36:39] RESOLVED: CoreBGPDown: Core BGP session down between cr1-drmrs and cr2-eqiad (185.15.58.138) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=drmrs&var-device=cr1-drmrs:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [03:39:09] !log mwpresync@deploy2003 Finished scap sync-world: testwikis to 1.47.0-wmf.11 refs T430830 (duration: 36m 01s) [03:39:13] T430830: 1.47.0-wmf.11 deployment blockers - https://phabricator.wikimedia.org/T430830 [03:41:31] RESOLVED: Traffic on tunnel link: Device cr1-drmrs.wikimedia.org recovered from Traffic on tunnel link - https://alerts.wikimedia.org/?q=alertname%3DTraffic+on+tunnel+link [04:00:05] Deploy window Automatic removal of all obsolete MediaWiki versions from the deployment and bare metal servers (except the most-recent obsolete version) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260714T0400) [04:00:43] FIRING: [2x] JobUnavailable: Reduced availability for job atlas_exporter in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [04:01:09] !log mwpresync@deploy2003 Pruned MediaWiki: 1.47.0-wmf.8 (duration: 01m 07s) [04:05:42] FIRING: [2x] JobUnavailable: Reduced availability for job atlas_exporter in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [04:28:37] FIRING: [5x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-3/1/2 (Transport: Hurricane Electric (dc4841.dal5) {#122601}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [05:07:29] !log jelto@cumin1003 START - Cookbook sre.hosts.reboot-single for host gitlab1004.wikimedia.org [05:08:54] PROBLEM - Host gitlab.wikimedia.org is DOWN: PING CRITICAL - Packet loss = 100% [05:10:50] ^ thats me [05:10:54] RECOVERY - Host gitlab.wikimedia.org is UP: PING OK - Packet loss = 0%, RTA = 0.34 ms [05:12:52] PROBLEM - Gitlab HTTPS SSL Expiry on gitlab.wikimedia.org is CRITICAL: connect to address gitlab.wikimedia.org and port 443: Connection refused https://wikitech.wikimedia.org/wiki/GitLab%23Monitoring [05:13:12] PROBLEM - Gitlab HTTPS healthcheck on gitlab.wikimedia.org is CRITICAL: HTTP CRITICAL: HTTP/1.1 502 Bad Gateway - 2353 bytes in 0.019 second response time https://wikitech.wikimedia.org/wiki/GitLab%23Monitoring [05:13:46] !log jelto@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1004.wikimedia.org [05:13:52] RECOVERY - Gitlab HTTPS SSL Expiry on gitlab.wikimedia.org is OK: OK - Certificate gitlab.wikimedia.org will expire on Tue 01 Sep 2026 09:02:37 AM GMT +0000. https://wikitech.wikimedia.org/wiki/GitLab%23Monitoring [05:14:14] RECOVERY - Gitlab HTTPS healthcheck on gitlab.wikimedia.org is OK: HTTP OK: HTTP/1.1 200 OK - 28690 bytes in 1.529 second response time https://wikitech.wikimedia.org/wiki/GitLab%23Monitoring [05:15:43] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [05:15:43] FIRING: [4x] ProbeDown: Service gitlab1004:22 has failed probes (tcp_gitlab_wikimedia_org_ssh_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [05:16:25] FIRING: SystemdUnitFailed: gitlab-package-puller.service on apt-staging2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [05:20:43] RESOLVED: [4x] ProbeDown: Service gitlab1004:22 has failed probes (tcp_gitlab_wikimedia_org_ssh_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [05:21:25] RESOLVED: SystemdUnitFailed: gitlab-package-puller.service on apt-staging2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [05:23:03] (03PS1) 10Marostegui: wmnet: Failover m3-master [dns] - 10https://gerrit.wikimedia.org/r/1310256 (https://phabricator.wikimedia.org/T431660) [05:24:29] (03CR) 10Marostegui: [C:03+2] wmnet: Failover m3-master [dns] - 10https://gerrit.wikimedia.org/r/1310256 (https://phabricator.wikimedia.org/T431660) (owner: 10Marostegui) [05:24:33] !log marostegui@dns1004 START - running authdns-update [05:24:40] !log marostegui@dns1004 START - running authdns-update [05:26:32] !log marostegui@dns1004 END - running authdns-update [05:31:00] (03PS1) 10Marostegui: installserver: Do not format clouddb1029 [puppet] - 10https://gerrit.wikimedia.org/r/1310258 [05:32:39] (03CR) 10JavierMonton: [C:03+2] stream: pageview-trending-relative [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309908 (https://phabricator.wikimedia.org/T430675) (owner: 10JavierMonton) [05:32:44] (03CR) 10JavierMonton: [C:03+2] topic: webrequest-page-view-next [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310025 (https://phabricator.wikimedia.org/T430675) (owner: 10JavierMonton) [05:34:43] (03Merged) 10jenkins-bot: stream: pageview-trending-relative [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309908 (https://phabricator.wikimedia.org/T430675) (owner: 10JavierMonton) [05:34:55] (03Merged) 10jenkins-bot: topic: webrequest-page-view-next [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310025 (https://phabricator.wikimedia.org/T430675) (owner: 10JavierMonton) [05:38:30] (03CR) 10Marostegui: [C:03+2] installserver: Do not format clouddb1029 [puppet] - 10https://gerrit.wikimedia.org/r/1310258 (owner: 10Marostegui) [05:39:46] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply [05:40:21] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply [05:58:18] (03PS1) 10Kosta Harlan: SiteStats: Add temporary account columns to site_stats [core] (wmf/1.47.0-wmf.11) - 10https://gerrit.wikimedia.org/r/1310470 (https://phabricator.wikimedia.org/T339291) [05:59:20] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply [05:59:55] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply [06:00:04] Deploy window MediaWiki infrastructure (UTC early) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260714T0600) [06:00:05] marostegui, Amir1, and federico3: #bothumor When your hammer is PHP, everything starts looking like a thumb. Rise for Primary database switchover. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260714T0600). [06:00:54] !log javiermonton@deploy2003 helmfile [eqiad] START helmfile.d/services/eventgate-main: sync [06:01:20] !log javiermonton@deploy2003 helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync [06:01:37] !log javiermonton@deploy2003 helmfile [codfw] START helmfile.d/services/eventgate-main: sync [06:02:00] !log javiermonton@deploy2003 helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync [06:03:17] !log javiermonton@deploy2003 helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync [06:03:50] !log javiermonton@deploy2003 helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync [06:04:17] !log javiermonton@deploy2003 helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync [06:04:40] !log javiermonton@deploy2003 helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync [06:05:43] FIRING: [2x] JobUnavailable: Reduced availability for job atlas_exporter in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [06:10:43] FIRING: [2x] JobUnavailable: Reduced availability for job atlas_exporter in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [06:17:16] (03PS1) 10Marostegui: Revert "wmnet: Failover m3-master" [dns] - 10https://gerrit.wikimedia.org/r/1310474 [06:17:17] !log marostegui@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1026.eqiad.wmnet with reason: reboot [06:20:01] (03CR) 10Marostegui: [C:03+2] Revert "wmnet: Failover m3-master" [dns] - 10https://gerrit.wikimedia.org/r/1310474 (owner: 10Marostegui) [06:20:08] !log marostegui@dns1004 START - running authdns-update [06:21:30] 10SRE-Access-Requests, 06Data-Engineering, 13Patch-For-Review: Add catherinekelsey to analytics-wmde-users group - https://phabricator.wikimedia.org/T432002#12117946 (10catherine.kelsey.wmde) Hey @Lena_WMDE - when you get a moment, please could you approve? Thanks! [06:21:54] (03PS1) 10Marostegui: wmnet: Failover m5-master [dns] - 10https://gerrit.wikimedia.org/r/1310475 (https://phabricator.wikimedia.org/T431660) [06:22:01] !log marostegui@dns1004 END - running authdns-update [06:22:42] 10SRE-Access-Requests, 06Data-Engineering, 13Patch-For-Review: Add catherinekelsey to analytics-wmde-users group - https://phabricator.wikimedia.org/T432002#12117952 (10Lena_WMDE) Hi all, I approve @catherine.kelsey.wmde 's request. Thanks! [06:23:05] (03CR) 10Marostegui: [C:03+2] wmnet: Failover m5-master [dns] - 10https://gerrit.wikimedia.org/r/1310475 (https://phabricator.wikimedia.org/T431660) (owner: 10Marostegui) [06:23:10] !log marostegui@dns1004 START - running authdns-update [06:25:04] !log marostegui@dns1004 END - running authdns-update [06:37:53] 10ops-codfw, 06SRE, 06DBA, 06DC-Ops: db2209 firmware upgrade after a crash - https://phabricator.wikimedia.org/T431952#12117967 (10Marostegui) p:05Triage→03Medium @Jhancock.wm this is what i got from idrac logs during the time of the crash: ` ------------------------------------------------------------... [06:44:09] !log jelto@cumin1003 START - Cookbook sre.hosts.reboot-single for host lists1004.wikimedia.org [06:50:42] !log jelto@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists1004.wikimedia.org [07:00:04] Amir1, urbanecm, and awight: Time to snap out of that daydream and deploy UTC morning backport window. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260714T0700). [07:00:04] No Gerrit patches in the queue for this window AFAICS. [07:02:46] FIRING: Not accepting/receiving prefixes from anycast BGP peer: Alert for device asw1-b4-magru.mgmt.magru.wmnet - Not accepting/receiving prefixes from anycast BGP peer - https://alerts.wikimedia.org/?q=alertname%3DNot+accepting%2Freceiving+prefixes+from+anycast+BGP+peer [07:06:32] 06SRE, 06Infrastructure-Foundations, 10Kafka-Infrastructure, 06ServiceOps new, 10ServiceOps-Datastores: Upgrade Kafka to version 3.x - https://phabricator.wikimedia.org/T416669#12118002 (10elukey) 05Open→03Resolved a:03elukey There is still one subtask opened but it is a cookbook follow up. All... [07:10:38] FIRING: [2x] CertAlmostExpired: gNMI TLS certificate for lsw1-d2-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [07:11:43] (03PS3) 10Elukey: profile::puppetserver::volatile: add metrics for airflow webrequest [puppet] - 10https://gerrit.wikimedia.org/r/1310118 (https://phabricator.wikimedia.org/T402512) [07:12:18] (03CR) 10Elukey: profile::puppetserver::volatile: add metrics for airflow webrequest (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1310118 (https://phabricator.wikimedia.org/T402512) (owner: 10Elukey) [07:12:48] (03CR) 10Elukey: [C:03+2] sre.hosts.bmc-user-mgmt: add more logging [cookbooks] - 10https://gerrit.wikimedia.org/r/1309226 (https://phabricator.wikimedia.org/T426180) (owner: 10Elukey) [07:13:51] (03CR) 10Elukey: "recheck" [software/spicerack] - 10https://gerrit.wikimedia.org/r/1310131 (https://phabricator.wikimedia.org/T426180) (owner: 10Elukey) [07:14:15] (03PS2) 10Elukey: redfish: skip setting AccountTypes if the admin user doesn't expose it [software/spicerack] - 10https://gerrit.wikimedia.org/r/1309120 (https://phabricator.wikimedia.org/T426180) [07:18:53] (03PS2) 10Gerrit maintenance bot: mariadb: Promote db2220 to s7 master [puppet] - 10https://gerrit.wikimedia.org/r/1307081 (https://phabricator.wikimedia.org/T430920) [07:23:40] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply [07:26:07] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply [07:26:18] !log elukey@cumin1003 START - Cookbook sre.k8s.reboot-nodes rolling reboot on P{aux-k8s-ctrl200*} and (A:aux-master-codfw or A:aux-worker-codfw) [07:26:21] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2002.codfw.wmnet [07:26:23] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2002.codfw.wmnet [07:26:29] 10SRE-Access-Requests, 06Data-Engineering, 13Patch-For-Review: Add catherinekelsey to analytics-wmde-users group - https://phabricator.wikimedia.org/T432002#12118034 (10jcrespo) [07:28:15] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host ms-fe2024.codfw.wmnet with OS trixie [07:28:30] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12118035 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host ms-fe2024.codfw.wmnet with OS trixie [07:29:40] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-fe1024.eqiad.wmnet with OS trixie [07:29:55] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12118037 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-fe1024.eqiad.wmnet with OS trixie [07:31:19] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2002.codfw.wmnet [07:31:21] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2002.codfw.wmnet [07:31:27] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2003.codfw.wmnet [07:31:28] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2003.codfw.wmnet [07:34:52] !log elukey@cumin1003 START - Cookbook sre.misc-clusters.restart-reboot-config-master rolling reboot on P{config-master*} and (A:config-master or A:config-master-eqiad or A:config-master-codfw) [07:35:12] !log elukey@cumin1003 START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors [07:35:16] !log elukey@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors [07:36:33] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2003.codfw.wmnet [07:36:35] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2003.codfw.wmnet [07:36:35] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P{aux-k8s-ctrl200*} and (A:aux-master-codfw or A:aux-worker-codfw) [07:39:20] !log elukey@cumin1003 START - Cookbook sre.k8s.reboot-nodes rolling reboot on P{aux-k8s-worker2*} and (A:aux-master-codfw or A:aux-worker-codfw) [07:39:23] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2002.codfw.wmnet [07:39:41] !log elukey@cumin1003 START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors [07:39:45] !log elukey@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors [07:39:56] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2002.codfw.wmnet [07:44:01] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2002.codfw.wmnet [07:44:03] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2002.codfw.wmnet [07:44:09] !log elukey@cumin1003 END (PASS) - Cookbook sre.misc-clusters.restart-reboot-config-master (exit_code=0) rolling reboot on P{config-master*} and (A:config-master or A:config-master-eqiad or A:config-master-codfw) [07:44:10] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2003.codfw.wmnet [07:45:01] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply [07:45:01] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage [07:45:11] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply [07:46:03] !log mvernon@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage [07:47:46] 06SRE, 10Wikimedia Australia, 10Wikimedia-Mailing-lists: Create new private announce-only mailing list for the ICIP project (Wikimedia Australia) - https://phabricator.wikimedia.org/T432082#12118082 (10AlphaLemur) [07:48:46] !log elukey@cumin1003 START - Cookbook sre.pki.restart-reboot rolling reboot on P{pki*} and (A:pki) [07:48:47] (03PS1) 10Filippo Giunchedi: dumps: create /mnt/nfs/dumps directory [puppet] - 10https://gerrit.wikimedia.org/r/1310482 (https://phabricator.wikimedia.org/T411248) [07:48:54] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage [07:49:10] !log elukey@cumin1003 START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors [07:49:12] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2003.codfw.wmnet [07:49:14] !log elukey@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors [07:51:55] (03CR) 10Filippo Giunchedi: [C:03+2] "Trivial and currently testing, self-merging" [puppet] - 10https://gerrit.wikimedia.org/r/1310482 (https://phabricator.wikimedia.org/T411248) (owner: 10Filippo Giunchedi) [07:53:12] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage [07:53:13] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2003.codfw.wmnet [07:53:15] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2003.codfw.wmnet [07:53:20] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2004.codfw.wmnet [07:53:52] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2004.codfw.wmnet [07:54:55] !log elukey@cumin1003 START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors [07:54:59] !log elukey@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors [07:56:25] FIRING: [23x] SystemdUnitFailed: cfssl-ocsprefresh-Wikimedia_Internal_Root_CA.service on pki2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [07:57:57] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2004.codfw.wmnet [07:57:59] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2004.codfw.wmnet [07:58:04] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2005.codfw.wmnet [07:58:36] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2005.codfw.wmnet [08:01:25] FIRING: [23x] SystemdUnitFailed: cfssl-ocsprefresh-Wikimedia_Internal_Root_CA.service on pki2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [08:02:43] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2005.codfw.wmnet [08:02:45] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2005.codfw.wmnet [08:02:51] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2006.codfw.wmnet [08:03:27] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2006.codfw.wmnet [08:05:43] FIRING: [2x] JobUnavailable: Reduced availability for job cfssl in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [08:06:19] (03PS7) 10Mvolz: Update update_version.py to be compatible with ruamel >=0.15.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310090 [08:06:25] RESOLVED: [23x] SystemdUnitFailed: cfssl-ocsprefresh-Wikimedia_Internal_Root_CA.service on pki2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [08:08:03] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2024.codfw.wmnet with OS trixie [08:08:22] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12118106 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-fe2024.codfw.wmnet with OS trixie completed... [08:08:35] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2006.codfw.wmnet [08:08:36] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2006.codfw.wmnet [08:08:42] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet [08:09:14] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet [08:10:43] FIRING: [2x] JobUnavailable: Reduced availability for job cfssl in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [08:10:58] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply [08:11:07] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply [08:11:54] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply [08:12:18] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply [08:12:55] !log elukey@cumin1003 END (PASS) - Cookbook sre.pki.restart-reboot (exit_code=0) rolling reboot on P{pki*} and (A:pki) [08:13:04] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1024.eqiad.wmnet with OS trixie [08:13:16] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12118131 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-fe1024.eqiad.wmnet with OS trixie completed... [08:14:23] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet [08:14:24] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet [08:14:30] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2008.codfw.wmnet [08:15:04] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2008.codfw.wmnet [08:20:24] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2008.codfw.wmnet [08:20:25] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2008.codfw.wmnet [08:20:31] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2009.codfw.wmnet [08:21:00] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply [08:21:01] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply [08:21:08] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2009.codfw.wmnet [08:21:11] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply [08:21:19] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply [08:23:43] !log marostegui@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot [08:24:11] !log marostegui@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1029.eqiad.wmnet with reason: reboot [08:24:26] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host ms-fe2023.codfw.wmnet with OS trixie [08:24:32] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply [08:24:42] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12118149 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host ms-fe2023.codfw.wmnet with OS trixie [08:24:44] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply [08:25:20] (03PS1) 10Marostegui: Revert "wmnet: Failover m5-master" [dns] - 10https://gerrit.wikimedia.org/r/1310517 [08:25:22] !log elukey@cumin1003 START - Cookbook sre.hosts.reboot-single for host netbox-dev2003.codfw.wmnet [08:25:36] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-fe1023.eqiad.wmnet with OS trixie [08:25:54] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12118151 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-fe1023.eqiad.wmnet with OS trixie [08:26:56] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2009.codfw.wmnet [08:26:58] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2009.codfw.wmnet [08:26:58] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P{aux-k8s-worker2*} and (A:aux-master-codfw or A:aux-worker-codfw) [08:27:40] (03CR) 10Marostegui: [C:03+2] Revert "wmnet: Failover m5-master" [dns] - 10https://gerrit.wikimedia.org/r/1310517 (owner: 10Marostegui) [08:27:44] !log marostegui@dns1004 START - running authdns-update [08:28:37] FIRING: [5x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-3/1/2 (Transport: Hurricane Electric (dc4841.dal5) {#122601}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [08:29:19] 06SRE, 10Wikimedia Australia, 10Wikimedia-Mailing-lists: Create new private announce-only mailing list for the ICIP project (Wikimedia Australia) - https://phabricator.wikimedia.org/T432082#12118155 (10jcrespo) 05Open→03Resolved a:03jcrespo An announcements-style list has been created at https://li... [08:29:22] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox-dev2003.codfw.wmnet [08:29:38] !log marostegui@dns1004 END - running authdns-update [08:31:17] (03PS1) 10Hnowlan: karma: consistently order some labels first [puppet] - 10https://gerrit.wikimedia.org/r/1310519 (https://phabricator.wikimedia.org/T431980) [08:34:37] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply [08:34:44] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply [08:38:17] (03CR) 10Hnowlan: [C:03+1] "One last fix (Sorry!), but feel free to ship once that's addressed!" [puppet] - 10https://gerrit.wikimedia.org/r/1305774 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [08:40:43] 06SRE, 10Wikimedia-Mailing-lists: Request for a Mailing List for Wikimedia Community User Group Uganda - https://phabricator.wikimedia.org/T432087#12118229 (10jcrespo) 05Open→03Resolved a:03jcrespo The list has been created with basic default settings at: https://lists.wikimedia.org/postorius/lists/... [08:41:30] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage [08:42:07] !log mvernon@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage [08:43:54] (03Abandoned) 10Mvolz: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1301732 (owner: 10PipelineBot) [08:44:08] (03PS3) 10Gerrit maintenance bot: mariadb: Promote db2220 to s7 master [puppet] - 10https://gerrit.wikimedia.org/r/1307081 (https://phabricator.wikimedia.org/T430920) [08:44:37] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage [08:45:29] !log cwilliams@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 T430920 [08:45:32] T430920: Switchover s7 master (db2159 -> db2220) - https://phabricator.wikimedia.org/T430920 [08:45:53] !log cwilliams@cumin1003 dbctl commit (dc=all): 'Set db2220 with weight 0 T430920', diff saved to https://phabricator.wikimedia.org/P94809 and previous config saved to /var/cache/conftool/dbconfig/20260714-084553-cwilliams.json [08:54:06] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage [08:54:06] 10SRE-Access-Requests, 10LDAP-Access-Requests, 13Patch-For-Review: Superset data access request for abibendall - https://phabricator.wikimedia.org/T430938#12118257 (10jcrespo) I ask for patience, sadly the people usually in charge of quickly approving WMF LDAP access are temporarily out of office, and the re... [08:54:07] (03CR) 10CWilliams: [C:03+2] mariadb: Promote db2220 to s7 master [puppet] - 10https://gerrit.wikimedia.org/r/1307081 (https://phabricator.wikimedia.org/T430920) (owner: 10Gerrit maintenance bot) [08:54:07] !log Starting s7 codfw failover from db2159 to db2220 - T430920 [08:54:07] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [08:54:07] T430920: Switchover s7 master (db2159 -> db2220) - https://phabricator.wikimedia.org/T430920 [08:54:07] (03CR) 10Atsuko: [V:03+1] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/8990/co" [puppet] - 10https://gerrit.wikimedia.org/r/1310129 (https://phabricator.wikimedia.org/T431538) (owner: 10Atsuko) [08:54:07] !log cwilliams@cumin1003 dbctl commit (dc=all): 'Promote db2220 to s7 primary T430920', diff saved to https://phabricator.wikimedia.org/P94810 and previous config saved to /var/cache/conftool/dbconfig/20260714-085239-cwilliams.json [08:54:08] !log blake@cumin1003 START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1055.eqiad.wmnet [08:54:08] !log blake@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1055.eqiad.wmnet [08:54:08] !log blake@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1055.eqiad.wmnet [08:54:45] !log blake@cumin1003 START - Cookbook sre.hosts.reimage for host wikikube-worker1055.eqiad.wmnet with OS trixie [08:55:13] !log blake@cumin1003 START - Cookbook sre.hosts.move-vlan for host wikikube-worker1055 [08:55:52] !log blake@cumin1003 START - Cookbook sre.dns.netbox [08:56:25] !log cwilliams@cumin1003 dbctl commit (dc=all): 'Depool db2159 T430920', diff saved to https://phabricator.wikimedia.org/P94811 and previous config saved to /var/cache/conftool/dbconfig/20260714-085624-cwilliams.json [08:59:07] !log a-pizzata@deploy2003 Started deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] [09:00:04] !log blake@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" [09:00:09] !log blake@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003" [09:00:09] !log blake@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [09:00:09] !log blake@cumin1003 START - Cookbook sre.dns.wipe-cache wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [09:00:12] !log blake@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [09:00:14] !log blake@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1055 [09:01:09] !log a-pizzata@deploy2003 Finished deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] (duration: 02m 01s) [09:01:52] !log a-pizzata@deploy2003 Started deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] [09:02:58] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2023.codfw.wmnet with OS trixie [09:03:11] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12118303 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-fe2023.codfw.wmnet with OS trixie completed... [09:04:44] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host ms-fe2022.codfw.wmnet with OS trixie [09:05:01] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12118306 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host ms-fe2022.codfw.wmnet with OS trixie [09:06:41] !log a-pizzata@deploy2003 Finished deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] (duration: 04m 49s) [09:07:10] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1023.eqiad.wmnet with OS trixie [09:07:20] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12118320 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-fe1023.eqiad.wmnet with OS trixie completed... [09:10:37] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-fe1022.eqiad.wmnet with OS trixie [09:10:52] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12118342 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-fe1022.eqiad.wmnet with OS trixie [09:11:11] !log a-pizzata@deploy2003 Started deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] [09:13:19] !log a-pizzata@deploy2003 Finished deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] (duration: 02m 07s) [09:14:46] (03CR) 10Mvolz: Update update_version.py to be compatible with ruamel >=0.15.0 (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310090 (owner: 10Mvolz) [09:15:43] FIRING: [2x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [09:16:42] !log cwilliams@cumin1003 START - Cookbook sre.mysql.multiinstance_reboot for db2159.codfw.wmnet [09:17:55] !log jgiannelos@deploy2003 helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply [09:18:29] !log jgiannelos@deploy2003 helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply [09:18:30] !log jgiannelos@deploy2003 helmfile [codfw] START helmfile.d/services/mw-parsoid: apply [09:18:55] !log jgiannelos@deploy2003 helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply [09:19:51] cwilliams@cumin1003 multiinstance_reboot (PID 4009282) is awaiting input [09:20:02] !log cwilliams@cumin1003 START - Cookbook sre.mysql.depool depool db2159: Rebooting db2159.codfw.wmnet [09:20:14] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: Rebooting db2159.codfw.wmnet [09:21:24] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage [09:21:50] 10SRE-Access-Requests, 06Infrastructure-Foundations, 10LDAP-Access-Requests, 13Patch-For-Review: Superset data access request for abibendall - https://phabricator.wikimedia.org/T430938#12118356 (10jcrespo) This is blocked on LDAP new workflow approval before the rest of the access is handled by clinic duty. [09:23:49] RECOVERY - jenkins_service_running on contint1002 is OK: PROCS OK: 1 process with regex args .*/bin/java .*-jar /usr/share/java/jenkins.war https://wikitech.wikimedia.org/wiki/Jenkins [09:24:36] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage [09:26:49] PROBLEM - jenkins_service_running on contint1002 is CRITICAL: PROCS CRITICAL: 0 processes with regex args .*/bin/java .*-jar /usr/share/java/jenkins.war https://wikitech.wikimedia.org/wiki/Jenkins [09:27:04] !log mvernon@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage [09:27:08] !log mvernon@cumin1003 END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage [09:28:33] !log blake@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1055 [09:28:33] !log blake@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1055 [09:29:13] (03CR) 10Marostegui: "Pending doc and then it would be good to go" [cookbooks] - 10https://gerrit.wikimedia.org/r/1277076 (https://phabricator.wikimedia.org/T419874) (owner: 10Federico Ceratto) [09:31:33] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2159.codfw.wmnet [09:32:12] 06SRE, 10Wikimedia-Mailing-lists: Request for a Mailing List for Wikimedia Community User Group Uganda - https://phabricator.wikimedia.org/T432087#12118391 (10MichealKaluba) Thank you team, much appreciated! [09:40:43] FIRING: [2x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [09:42:11] PROBLEM - Host ms-fe1022 is DOWN: PING CRITICAL - Packet loss = 100% [09:42:31] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2022.codfw.wmnet with OS trixie [09:42:45] RECOVERY - Host ms-fe1022 is UP: PING OK - Packet loss = 0%, RTA = 0.34 ms [09:42:48] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12118404 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-fe2022.codfw.wmnet with OS trixie completed... [09:44:06] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host ms-fe2021.codfw.wmnet with OS trixie [09:44:14] !log cwilliams@cumin1003 START - Cookbook sre.mysql.pool pool db2159: Repooling after switchover [09:44:17] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1022.eqiad.wmnet with OS trixie [09:44:24] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12118405 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host ms-fe2021.codfw.wmnet with OS trixie [09:44:27] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12118406 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-fe1022.eqiad.wmnet with OS trixie completed... [09:45:44] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-fe1021.eqiad.wmnet with OS trixie [09:45:55] !log blake@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage [09:46:00] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12118409 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-fe1021.eqiad.wmnet with OS trixie [09:48:14] (03PS1) 10Genoveva Galarza: abstractwiki: Make cacheAbstractContentFragment throttling a global setting [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310530 (https://phabricator.wikimedia.org/T430898) [09:49:36] (03CR) 10Cathal Mooney: [C:03+1] "LGTM!" [software/spicerack] - 10https://gerrit.wikimedia.org/r/1308680 (https://phabricator.wikimedia.org/T426180) (owner: 10Elukey) [09:50:14] !log blake@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage [09:51:43] (03CR) 10Atsuko: [V:03+1 C:03+2] service: change pybal check for opensearch clusters [puppet] - 10https://gerrit.wikimedia.org/r/1310129 (https://phabricator.wikimedia.org/T431538) (owner: 10Atsuko) [09:58:41] !log restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 [09:58:42] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [10:00:03] 07sre-alert-triage, 06Data-Platform-SRE (2026-07-03 - 2026-07-31), 06Machine-Learning-Team (Q1 FY2026-27): Alert in need of triage: SmartNotHealthy (instance ml-serve1001:9100) - https://phabricator.wikimedia.org/T414969#12118461 (10Gehel) [10:00:03] (03PS1) 10Atsuko: Revert "service: change pybal check for opensearch clusters" [puppet] - 10https://gerrit.wikimedia.org/r/1310531 [10:00:05] Deploy window MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260714T1000) [10:00:14] (03CR) 10Atsuko: [C:03+2] Revert "service: change pybal check for opensearch clusters" [puppet] - 10https://gerrit.wikimedia.org/r/1310531 (owner: 10Atsuko) [10:00:18] (03CR) 10Atsuko: [V:03+2 C:03+2] Revert "service: change pybal check for opensearch clusters" [puppet] - 10https://gerrit.wikimedia.org/r/1310531 (owner: 10Atsuko) [10:01:08] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage [10:01:42] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - search_9200: Servers cirrussearch1103.eqiad.wmnet, cirrussearch1124.eqiad.wmnet, cirrussearch1087.eqiad.wmnet, cirrussearch1079.eqiad.wmnet, cirrussearch1121.eqiad.wmnet, cirrussearch1123.eqiad.wmnet, cirrussearch1113.eqiad.wmnet, cirrussearch1073.eqiad.wmnet, cirrussearch1101.eqiad.wmnet, cirrussearch1108.eqiad.wmnet, cirrussearch1102.eqiad.wmnet, c [10:01:42] rch1088.eqiad.wmnet, cirrussearch1070.eqiad.wmnet, cirrussearch1086.eqiad.wmnet, cirrussearch1107.eqiad.wmnet, cirrussearch1097.eqiad.wmnet, cirrussearch1091.eqiad.wmnet, cirrussearch1080.eqiad.wmnet, cirrussearch1089.eqiad.wmnet, cirrussearch1119.eqiad.wmnet, cirrussearch1074.eqiad.wmnet, cirrussearch1083.eqiad.wmnet, cirrussearch1081.eqiad.wmnet, cirrussearch1076.eqiad.wmnet, cirrussearch1112.eqiad.wmnet, cirrussearch1096.eqiad.wmnet, c [10:01:42] rch1114.eqiad.wmnet, cirrussearch1116.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [10:01:56] ack [10:02:04] healthcheck has failed, I'm reverting [10:02:13] !log mvernon@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage [10:03:45] !log restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 revert [10:03:46] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [10:04:42] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [10:05:47] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage [10:09:33] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage [10:11:35] !log blake@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1055.eqiad.wmnet with OS trixie [10:14:12] (03PS1) 10Gmodena: wdqs: codfw: deploy test qlever image [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310533 (https://phabricator.wikimedia.org/T431271) [10:15:23] FIRING: [2x] CertAlmostExpired: gNMI TLS certificate for lsw1-d2-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [10:16:05] (03CR) 10Federico Ceratto: "Documented at https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting#Setting_all_sections_read-only_in_emergency and https://wikitech" [cookbooks] - 10https://gerrit.wikimedia.org/r/1277076 (https://phabricator.wikimedia.org/T419874) (owner: 10Federico Ceratto) [10:17:35] blake@cumin1003 renumber-node (PID 4005988) is awaiting input [10:18:31] (03CR) 10Trueg: wdqs: codfw: deploy test qlever image (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310533 (https://phabricator.wikimedia.org/T431271) (owner: 10Gmodena) [10:19:46] (03CR) 10Marostegui: [C:03+1] "Thank you!" [cookbooks] - 10https://gerrit.wikimedia.org/r/1277076 (https://phabricator.wikimedia.org/T419874) (owner: 10Federico Ceratto) [10:20:42] jouncebot: nowandnext [10:20:42] For the next 0 hour(s) and 39 minute(s): MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260714T1000) [10:20:42] In 1 hour(s) and 39 minute(s): Mobileapps/RESTBase/Wikifeeds (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260714T1200) [10:22:30] (03CR) 10TrainBranchBot: [C:03+2] "Approved by kharlan@deploy2003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1307716 (https://phabricator.wikimedia.org/T431023) (owner: 10Kosta Harlan) [10:23:24] (03Merged) 10jenkins-bot: extension-list: Add WikimediaAntiAbuse [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1307716 (https://phabricator.wikimedia.org/T431023) (owner: 10Kosta Harlan) [10:24:02] !log kharlan@deploy2003 Started scap sync-world: Backport for [[gerrit:1307716|extension-list: Add WikimediaAntiAbuse (T431023)]] [10:24:02] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2021.codfw.wmnet with OS trixie [10:24:06] T431023: Create and deploy Extension:WikimediaAntiAbuse - https://phabricator.wikimedia.org/T431023 [10:24:13] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12118527 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-fe2021.codfw.wmnet with OS trixie completed... [10:24:53] (03PS1) 10Atsuko: service: change pybal check for opensearch clusters [puppet] - 10https://gerrit.wikimedia.org/r/1310535 (https://phabricator.wikimedia.org/T431538) [10:26:07] !log blake@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1055.eqiad.wmnet [10:26:08] !log blake@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1055.eqiad.wmnet [10:26:10] !log blake@cumin1003 END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1055.eqiad.wmnet [10:26:10] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host ms-fe2020.codfw.wmnet with OS trixie [10:26:22] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12118550 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host ms-fe2020.codfw.wmnet with OS trixie [10:26:49] !log blake@cumin1003 START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1067.eqiad.wmnet [10:26:53] !log blake@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1067.eqiad.wmnet [10:27:23] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply [10:27:25] !log blake@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1067.eqiad.wmnet [10:27:33] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply [10:27:41] !log blake@cumin1003 START - Cookbook sre.hosts.reimage for host wikikube-worker1067.eqiad.wmnet with OS trixie [10:27:50] (03CR) 10Atsuko: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1310535 (https://phabricator.wikimedia.org/T431538) (owner: 10Atsuko) [10:28:13] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1021.eqiad.wmnet with OS trixie [10:28:32] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12118559 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-fe1021.eqiad.wmnet with OS trixie completed... [10:29:39] !log blake@cumin1003 START - Cookbook sre.hosts.move-vlan for host wikikube-worker1067 [10:29:43] (03CR) 10Btullis: [C:03+1] "Looks good." [puppet] - 10https://gerrit.wikimedia.org/r/1310535 (https://phabricator.wikimedia.org/T431538) (owner: 10Atsuko) [10:29:44] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: Repooling after switchover [10:32:15] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-fe1020.eqiad.wmnet with OS trixie [10:32:35] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12118567 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-fe1020.eqiad.wmnet with OS trixie [10:32:42] blake@cumin1003 renumber-node (PID 4022368) is awaiting input [10:42:22] !log kharlan@deploy2003 kharlan: Backport for [[gerrit:1307716|extension-list: Add WikimediaAntiAbuse (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [10:42:25] T431023: Create and deploy Extension:WikimediaAntiAbuse - https://phabricator.wikimedia.org/T431023 [10:42:50] (03PS1) 10Marostegui: aliases.yaml: Add clouddb1029 to db-clouddb-sanitization [puppet] - 10https://gerrit.wikimedia.org/r/1310542 (https://phabricator.wikimedia.org/T409557) [10:43:02] !log kharlan@deploy2003 kharlan: Continuing with deployment [10:45:58] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage [10:48:18] !log mvernon@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage [10:48:53] (03PS8) 10Effie Mouzeli: trafficserver: Remove XWD routing for /w/rest.php mw-debug backend [puppet] - 10https://gerrit.wikimedia.org/r/1310103 (https://phabricator.wikimedia.org/T428909) [10:49:43] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply [10:49:49] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage [10:49:52] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply [10:50:12] (03PS2) 10Marostegui: aliases.yaml: Add clouddb1029 to db-clouddb-sanitization [puppet] - 10https://gerrit.wikimedia.org/r/1310542 (https://phabricator.wikimedia.org/T409557) [10:51:01] (03CR) 10CI reject: [V:04-1] trafficserver: Remove XWD routing for /w/rest.php mw-debug backend [puppet] - 10https://gerrit.wikimedia.org/r/1310103 (https://phabricator.wikimedia.org/T428909) (owner: 10Effie Mouzeli) [10:51:48] (03CR) 10FNegri: [C:03+1] aliases.yaml: Add clouddb1029 to db-clouddb-sanitization [puppet] - 10https://gerrit.wikimedia.org/r/1310542 (https://phabricator.wikimedia.org/T409557) (owner: 10Marostegui) [10:52:39] (03CR) 10Marostegui: [C:03+2] aliases.yaml: Add clouddb1029 to db-clouddb-sanitization [puppet] - 10https://gerrit.wikimedia.org/r/1310542 (https://phabricator.wikimedia.org/T409557) (owner: 10Marostegui) [10:52:44] !log marostegui@dns1004 START - running authdns-update [10:52:50] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage [10:55:42] !log kharlan@deploy2003 Finished scap sync-world: Backport for [[gerrit:1307716|extension-list: Add WikimediaAntiAbuse (T431023)]] (duration: 31m 40s) [10:55:46] T431023: Create and deploy Extension:WikimediaAntiAbuse - https://phabricator.wikimedia.org/T431023 [10:56:04] (03PS1) 10Marostegui: eqiad.yaml: Add clouddb1029 [puppet] - 10https://gerrit.wikimedia.org/r/1310547 (https://phabricator.wikimedia.org/T409557) [10:56:04] (03CR) 10TrainBranchBot: [C:03+2] "Approved by kharlan@deploy2003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1307717 (https://phabricator.wikimedia.org/T431023) (owner: 10Kosta Harlan) [10:56:24] (03CR) 10Kosta Harlan: WikimediaAntiAbuse: Enable everywhere [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1307751 (https://phabricator.wikimedia.org/T431023) (owner: 10Kosta Harlan) [10:56:41] (03CR) 10Marostegui: "@fnegri@wikimedia.org I want to get clouddb1029 added to the LB, I will not pool it yet." [puppet] - 10https://gerrit.wikimedia.org/r/1310547 (https://phabricator.wikimedia.org/T409557) (owner: 10Marostegui) [10:57:15] (03Merged) 10jenkins-bot: WikimediaAntiAbuse: Register wmgUse config and load extension [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1307717 (https://phabricator.wikimedia.org/T431023) (owner: 10Kosta Harlan) [10:57:19] 07sre-alert-triage, 06Data-Platform-SRE (2026-07-03 - 2026-07-31), 06Machine-Learning-Team (Q1 FY2026-27): Alert in need of triage: SmartNotHealthy (instance ml-serve1001:9100) - https://phabricator.wikimedia.org/T414969#12118642 (10klausman) a:03klausman [10:57:36] !log kharlan@deploy2003 Started scap sync-world: Backport for [[gerrit:1307717|WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] [10:57:48] (03PS9) 10Effie Mouzeli: trafficserver: Remove XWD routing for /w/rest.php mw-debug backend [puppet] - 10https://gerrit.wikimedia.org/r/1310103 (https://phabricator.wikimedia.org/T428909) [11:00:24] (03CR) 10Effie Mouzeli: trafficserver: Remove XWD routing for /w/rest.php mw-debug backend (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1310103 (https://phabricator.wikimedia.org/T428909) (owner: 10Effie Mouzeli) [11:01:20] !log kharlan@deploy2003 kharlan: Backport for [[gerrit:1307717|WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [11:01:23] T431023: Create and deploy Extension:WikimediaAntiAbuse - https://phabricator.wikimedia.org/T431023 [11:01:54] !log blake@cumin1003 START - Cookbook sre.dns.netbox [11:03:17] !log kharlan@deploy2003 kharlan: Continuing with deployment [11:05:00] (03PS10) 10Effie Mouzeli: trafficserver: Remove XWD routing for /w/rest.php mw-debug backend [puppet] - 10https://gerrit.wikimedia.org/r/1310103 (https://phabricator.wikimedia.org/T428909) [11:07:56] blake@cumin1003 renumber-node (PID 4022368) is awaiting input [11:08:58] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2020.codfw.wmnet with OS trixie [11:09:09] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12118707 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-fe2020.codfw.wmnet with OS trixie completed... [11:09:48] !log kharlan@deploy2003 Finished scap sync-world: Backport for [[gerrit:1307717|WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] (duration: 12m 12s) [11:09:52] T431023: Create and deploy Extension:WikimediaAntiAbuse - https://phabricator.wikimedia.org/T431023 [11:10:01] !log blake@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" [11:10:13] (03CR) 10TrainBranchBot: [C:03+2] "Approved by kharlan@deploy2003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1307750 (https://phabricator.wikimedia.org/T431023) (owner: 10Kosta Harlan) [11:10:39] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1020.eqiad.wmnet with OS trixie [11:10:48] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host ms-fe2019.codfw.wmnet with OS trixie [11:10:53] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12118717 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-fe1020.eqiad.wmnet with OS trixie completed... [11:11:00] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12118720 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host ms-fe2019.codfw.wmnet with OS trixie [11:12:11] i'm seeing an unexpected diff while running homer to re-ip a wikikube-worker - is anyone currently working on ml-serve1001? [11:12:37] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-fe1019.eqiad.wmnet with OS trixie [11:12:50] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12118730 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-fe1019.eqiad.wmnet with OS trixie [11:13:05] blake@cumin1003 renumber-node (PID 4022368) is awaiting input [11:13:07] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply [11:13:19] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply [11:15:01] !log btullis@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet [11:15:29] oh, this isn't homer, this is... ml-serve1001.yaml being updated [11:17:57] (03Merged) 10jenkins-bot: Enable WikimediaAntiAbuse on Beta Cluster [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1307750 (https://phabricator.wikimedia.org/T431023) (owner: 10Kosta Harlan) [11:18:12] 07sre-alert-triage, 06Data-Platform-SRE (2026-07-03 - 2026-07-31), 06Machine-Learning-Team (Q1 FY2026-27): Alert in need of triage: SmartNotHealthy (instance ml-serve1001:9100) - https://phabricator.wikimedia.org/T414969#12118754 (10klausman) Disk is indeed bad, have filed T432105 with DCOps to get a replace... [11:18:43] 07sre-alert-triage, 06Data-Platform-SRE (2026-07-03 - 2026-07-31), 06Machine-Learning-Team (Q1 FY2026-27): Alert in need of triage: SmartNotHealthy (instance ml-serve1001:9100) - https://phabricator.wikimedia.org/T414969#12118756 (10klausman) p:05Triage→03Medium [11:19:55] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply [11:20:03] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply [11:20:10] !log btullis@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet [11:21:37] (03PS2) 10Kosta Harlan: WikimediaAntiAbuse: Enable everywhere [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1307751 (https://phabricator.wikimedia.org/T431023) [11:23:16] jouncebot: nowandnext [11:23:16] No deployments scheduled for the next 0 hour(s) and 36 minute(s) [11:23:16] In 0 hour(s) and 36 minute(s): Mobileapps/RESTBase/Wikifeeds (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260714T1200) [11:23:25] I've got a quick config fix to deploy; any objections? [11:23:34] James_F: I just finished with my deploys [11:23:46] Excellent, thanks, will get it out quickly. [11:23:49] (03CR) 10TrainBranchBot: [C:03+2] "Approved by jforrester@deploy2003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310530 (https://phabricator.wikimedia.org/T430898) (owner: 10Genoveva Galarza) [11:25:01] (03CR) 10AikoChou: [C:03+2] changeprop: add liftwing revertrisk-wikidata to production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1307434 (https://phabricator.wikimedia.org/T420883) (owner: 10AikoChou) [11:26:27] (03Merged) 10jenkins-bot: abstractwiki: Make cacheAbstractContentFragment throttling a global setting [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310530 (https://phabricator.wikimedia.org/T430898) (owner: 10Genoveva Galarza) [11:26:45] !log jforrester@deploy2003 Started scap sync-world: Backport for [[gerrit:1310530|abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] [11:26:48] T430898: AW now calls for all fragments at once and overloads WF - https://phabricator.wikimedia.org/T430898 [11:27:52] (03Merged) 10jenkins-bot: changeprop: add liftwing revertrisk-wikidata to production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1307434 (https://phabricator.wikimedia.org/T420883) (owner: 10AikoChou) [11:28:37] !log jforrester@deploy2003 jforrester, gengh: Backport for [[gerrit:1310530|abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [11:28:54] !log mvernon@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage [11:29:15] (03CR) 10Gmodena: wdqs: codfw: deploy test qlever image (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310533 (https://phabricator.wikimedia.org/T431271) (owner: 10Gmodena) [11:29:58] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage [11:30:38] (03CR) 10Elukey: [C:03+2] redfish: don't assume that the Allow header is always present [software/spicerack] - 10https://gerrit.wikimedia.org/r/1308680 (https://phabricator.wikimedia.org/T426180) (owner: 10Elukey) [11:32:11] !log jforrester@deploy2003 jforrester, gengh: Continuing with deployment [11:32:43] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage [11:35:47] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply [11:35:49] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply [11:36:11] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply [11:36:25] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply [11:36:26] !log jforrester@deploy2003 Finished scap sync-world: Backport for [[gerrit:1310530|abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] (duration: 09m 41s) [11:36:30] T430898: AW now calls for all fragments at once and overloads WF - https://phabricator.wikimedia.org/T430898 [11:36:30] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage [11:38:33] o/ I'll be deploying this changeprop patch: https://gerrit.wikimedia.org/r/c/operations/deployment-charts/+/1307434 [11:38:40] (03CR) 10Elukey: [C:03+2] redfish: skip setting AccountTypes if the admin user doesn't expose it (031 comment) [software/spicerack] - 10https://gerrit.wikimedia.org/r/1309120 (https://phabricator.wikimedia.org/T426180) (owner: 10Elukey) [11:38:50] (03CR) 10Elukey: [C:03+2] redfish: add delete_account method [software/spicerack] - 10https://gerrit.wikimedia.org/r/1310131 (https://phabricator.wikimedia.org/T426180) (owner: 10Elukey) [11:38:58] (03PS3) 10Elukey: redfish: add delete_account method [software/spicerack] - 10https://gerrit.wikimedia.org/r/1310131 (https://phabricator.wikimedia.org/T426180) [11:41:48] !log aikochou@deploy2003 helmfile [eqiad] START helmfile.d/services/changeprop: sync [11:42:29] !log aikochou@deploy2003 helmfile [eqiad] DONE helmfile.d/services/changeprop: sync [11:42:42] !log elukey@cumin1003 START - Cookbook sre.k8s.reboot-nodes rolling reboot on P{aux-k8s-ctrl100*} and (A:aux-master-eqiad or A:aux-worker-eqiad) [11:42:46] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1002.eqiad.wmnet [11:42:46] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1002.eqiad.wmnet [11:43:41] (03CR) 10Atsuko: [C:03+2] service: change pybal check for opensearch clusters [puppet] - 10https://gerrit.wikimedia.org/r/1310535 (https://phabricator.wikimedia.org/T431538) (owner: 10Atsuko) [11:44:36] !log marostegui@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5:00:00 on es1046.eqiad.wmnet with reason: Reimage to Trixie [11:44:39] !log marostegui@cumin1003 START - Cookbook sre.mysql.depool depool es1046: Reimage to Trixie [11:46:09] !log marostegui@cumin1003 END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1046: Reimage to Trixie [11:48:18] (03PS1) 10Cathal Mooney: pki: add profile::server_depool to document depool actions [puppet] - 10https://gerrit.wikimedia.org/r/1310554 (https://phabricator.wikimedia.org/T327300) [11:48:20] !log restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310535 [11:48:21] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [11:48:47] !log aikochou@deploy2003 helmfile [codfw] START helmfile.d/services/changeprop: sync [11:48:59] !log marostegui@cumin1003 START - Cookbook sre.hosts.reimage for host es1046.eqiad.wmnet with OS trixie [11:49:19] !log aikochou@deploy2003 helmfile [codfw] DONE helmfile.d/services/changeprop: sync [11:49:39] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1002.eqiad.wmnet [11:49:40] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1002.eqiad.wmnet [11:49:45] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1003.eqiad.wmnet [11:49:46] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1003.eqiad.wmnet [11:52:21] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1019.eqiad.wmnet with OS trixie [11:52:33] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12118876 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-fe1019.eqiad.wmnet with OS trixie completed... [11:53:12] (03PS1) 10JavierMonton: streams: webrequest-page-view - pageview-relative-trending [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310555 (https://phabricator.wikimedia.org/T430134) [11:54:46] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1003.eqiad.wmnet [11:54:47] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1003.eqiad.wmnet [11:54:47] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P{aux-k8s-ctrl100*} and (A:aux-master-eqiad or A:aux-worker-eqiad) [11:56:32] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-fe1018.eqiad.wmnet with OS trixie [11:56:45] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12118917 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-fe1018.eqiad.wmnet with OS trixie [11:56:47] !log elukey@cumin1003 START - Cookbook sre.k8s.reboot-nodes rolling reboot on P{aux-k8s-worker100*} and (A:aux-master-eqiad or A:aux-worker-eqiad) [11:56:50] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1002.eqiad.wmnet [11:57:03] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2019.codfw.wmnet with OS trixie [11:57:15] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12118919 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-fe2019.codfw.wmnet with OS trixie completed... [11:57:25] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1002.eqiad.wmnet [11:58:22] !log restarting pybal on lvs1018 `high-traffic2` for https://gerrit.wikimedia.org/r/1310535 [11:58:23] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [11:59:19] !log blake@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003" [11:59:19] !log blake@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [11:59:19] !log blake@cumin1003 START - Cookbook sre.dns.wipe-cache wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [11:59:22] !log blake@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [11:59:23] !log blake@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1067 [11:59:42] !log blake@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1067 [11:59:42] !log blake@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1067 [11:59:58] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host ms-fe2018.codfw.wmnet with OS trixie [12:00:04] Deploy window Mobileapps/RESTBase/Wikifeeds (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260714T1200) [12:00:11] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12118929 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host ms-fe2018.codfw.wmnet with OS trixie [12:00:51] (03PS1) 10Elukey: kafka: remove leftovers for Kafka 1.1 [puppet] - 10https://gerrit.wikimedia.org/r/1310558 (https://phabricator.wikimedia.org/T432089) [12:01:13] !log cmooney@cumin1003 START - Cookbook sre.dns.netbox [12:01:15] (03CR) 10Elukey: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1310558 (https://phabricator.wikimedia.org/T432089) (owner: 10Elukey) [12:01:28] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1002.eqiad.wmnet [12:01:29] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1002.eqiad.wmnet [12:01:35] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1003.eqiad.wmnet [12:02:10] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1003.eqiad.wmnet [12:02:16] (03PS2) 10Elukey: kafka: remove leftovers for Kafka 1.1 [puppet] - 10https://gerrit.wikimedia.org/r/1310558 (https://phabricator.wikimedia.org/T432089) [12:04:13] !log cmooney@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [12:04:38] !log cmooney@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" [12:05:16] 10ops-eqiad, 06Data-Platform-SRE, 06DC-Ops, 06Machine-Learning-Team: hw troubleshooting: failing disk in ml-serve1001.eqiad.wmnet - https://phabricator.wikimedia.org/T432105#12118961 (10cmooney) @VRiley-WMF I set the Netbox status for this one back to 'active'. I'm not 100% if that was the correct thing t... [12:05:16] !log cmooney@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003" [12:05:40] !log marostegui@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage [12:06:15] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1003.eqiad.wmnet [12:06:17] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1003.eqiad.wmnet [12:06:22] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1004.eqiad.wmnet [12:06:56] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1004.eqiad.wmnet [12:07:25] FIRING: SystemdUnitFailed: dump_proxy_ranges.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [12:08:29] !log restarting pybal on lvs1019 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 [12:08:30] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [12:09:52] !log marostegui@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage [12:10:54] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1004.eqiad.wmnet [12:10:55] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1004.eqiad.wmnet [12:11:00] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1005.eqiad.wmnet [12:11:34] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1005.eqiad.wmnet [12:12:44] !log mvernon@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage [12:15:12] !log restarting pybal on lvs2014 for https://gerrit.wikimedia.org/r/1310535 [12:15:13] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [12:15:29] (03CR) 10Jelto: [C:03+1] "lgtm, one comment regarding `helm3.11` in-line" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310146 (https://phabricator.wikimedia.org/T388390) (owner: 10Kamila Součková) [12:15:37] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1005.eqiad.wmnet [12:15:38] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1005.eqiad.wmnet [12:15:44] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1006.eqiad.wmnet [12:16:10] (03CR) 10Elukey: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1310558 (https://phabricator.wikimedia.org/T432089) (owner: 10Elukey) [12:16:16] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1006.eqiad.wmnet [12:18:42] PROBLEM - Check unit status of statograph_post on alert1002 is CRITICAL: CRITICAL: Status of the systemd unit statograph_post https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [12:18:49] !log blake@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage [12:19:15] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage [12:20:40] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage [12:21:18] (03CR) 10Elukey: [C:03+1] pki: add profile::server_depool to document depool actions [puppet] - 10https://gerrit.wikimedia.org/r/1310554 (https://phabricator.wikimedia.org/T327300) (owner: 10Cathal Mooney) [12:21:55] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1006.eqiad.wmnet [12:21:56] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1006.eqiad.wmnet [12:22:01] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1007.eqiad.wmnet [12:22:25] !log restarting pybal on lvs2013 `low-traffic` for https://gerrit.wikimedia.org/r/1310535 [12:22:26] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [12:22:32] !log elukey@cumin1003 START - Cookbook sre.hosts.reboot-single for host idp-test1005.wikimedia.org [12:22:33] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1007.eqiad.wmnet [12:23:08] !log mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model goodfaith --old (T431159) [12:23:11] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [12:23:11] T431159: ORES flagging not working on simplewiki - https://phabricator.wikimedia.org/T431159 [12:23:23] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test1005.wikimedia.org [12:24:08] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 - https://phabricator.wikimedia.org/T431682#12119042 (10fgiunchedi) @VRiley-WMF I thought we could start with one host, cloudvirt1048 and test-run the whole procedure there. Once we have nailed the process then we can d... [12:24:09] !log elukey@cumin1003 START - Cookbook sre.hosts.reboot-single for host idp-test2005.wikimedia.org [12:24:14] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage [12:25:24] !log ladsgroup@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1068.eqiad.wmnet with reason: vacuum overlarge container dbs [12:25:31] 06SRE, 10SRE-swift-storage: Disk near-full warnings on ms swift backends for container filesystems due to some bloated sqlite files - https://phabricator.wikimedia.org/T377827#12119051 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=b29b1198-5356-4ee3-bbfb-d6c22e5f7df4) set by ladsgroup@cum... [12:27:51] !log blake@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage [12:27:57] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1007.eqiad.wmnet [12:27:58] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1007.eqiad.wmnet [12:28:04] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1008.eqiad.wmnet [12:28:08] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test2005.wikimedia.org [12:28:42] RECOVERY - Check unit status of statograph_post on alert1002 is OK: OK: Status of the systemd unit statograph_post https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [12:29:43] !log elukey@cumin1003 START - Cookbook sre.hosts.reboot-single for host urldownloader1006.wikimedia.org [12:29:47] !log marostegui@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es1046.eqiad.wmnet with OS trixie [12:30:50] !log marostegui@cumin1003 START - Cookbook sre.mysql.pool pool es1046: es1046 after reimage [12:31:02] (03PS1) 10Effie Mouzeli: hiera: add profile::server_depool keys for redis and memcached. [puppet] - 10https://gerrit.wikimedia.org/r/1310566 (https://phabricator.wikimedia.org/T430930) [12:31:15] (03PS2) 10Effie Mouzeli: hiera: add profile::server_depool keys for redis and memcached. [puppet] - 10https://gerrit.wikimedia.org/r/1310566 (https://phabricator.wikimedia.org/T430930) [12:33:07] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1008.eqiad.wmnet [12:34:10] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1006.wikimedia.org [12:35:06] !log elukey@cumin1003 START - Cookbook sre.hosts.reboot-single for host urldownloader2006.wikimedia.org [12:38:09] (03PS1) 10Elukey: Failover url-dowloaders in eqiad and codfw [dns] - 10https://gerrit.wikimedia.org/r/1310567 [12:38:23] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1008.eqiad.wmnet [12:38:24] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1008.eqiad.wmnet [12:38:30] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1009.eqiad.wmnet [12:38:33] !log elukey@cumin1003 START - Cookbook sre.hosts.reboot-single for host pki-root1002.eqiad.wmnet [12:39:02] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1009.eqiad.wmnet [12:39:22] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1018.eqiad.wmnet with OS trixie [12:39:34] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12119125 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-fe1018.eqiad.wmnet with OS trixie completed... [12:39:37] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2006.wikimedia.org [12:42:06] (03PS1) 10AikoChou: EventStreams - Expose mediawiki.page_revert_risk_wikidata_prediction_change.v1 stream [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310570 (https://phabricator.wikimedia.org/T420883) [12:42:36] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-fe1017.eqiad.wmnet with OS trixie [12:42:46] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2018.codfw.wmnet with OS trixie [12:42:55] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12119136 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-fe1017.eqiad.wmnet with OS trixie [12:42:57] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12119137 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-fe2018.codfw.wmnet with OS trixie completed... [12:43:32] !log sukhe@cumin1003 START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling reboot on A:wikidough [12:43:49] !log sukhe@cumin1003 START - Cookbook sre.cdn.roll-restart-reboot-ncredir rolling reboot on A:ncredir and A:ncredir [12:43:54] RECOVERY - jenkins_service_running on contint1002 is OK: PROCS OK: 1 process with regex args .*/bin/java .*-jar /usr/share/java/jenkins.war https://wikitech.wikimedia.org/wiki/Jenkins [12:44:06] !log sukhe@cumin1003 START - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy [12:44:13] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1009.eqiad.wmnet [12:44:14] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1009.eqiad.wmnet [12:44:14] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P{aux-k8s-worker100*} and (A:aux-master-eqiad or A:aux-worker-eqiad) [12:44:20] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host pki-root1002.eqiad.wmnet [12:44:59] !log sukhe@cumin1003 START - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy rolling reboot on A:tcpproxy and A:tcpproxy [12:45:14] !log sukhe@cumin1003 START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum [12:45:48] !log mvernon@cumin2003 conftool action : set/pooled=inactive; selector: name=ms-fe2018.codfw.wmnet [12:46:54] PROBLEM - jenkins_service_running on contint1002 is CRITICAL: PROCS CRITICAL: 0 processes with regex args .*/bin/java .*-jar /usr/share/java/jenkins.war https://wikitech.wikimedia.org/wiki/Jenkins [12:47:05] !log mvernon@cumin2003 conftool action : set/pooled=yes; selector: name=ms-fe2018.codfw.wmnet [12:47:28] 10ops-eqiad, 06SRE, 06DC-Ops: Degraded RAID on backup1019 - https://phabricator.wikimedia.org/T431367#12119146 (10VRiley-WMF) Hey @jcrespo I just got the HDD back. I do apologize, that drive just got on the 10th, and I was on vacation at the time and just got back in today. I can still replace this drive whe... [12:47:37] !log sukhe@cumin1003 START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) [12:47:37] !log sukhe@cumin1003 cookbooks.sre.dns.roll-reboot begin reboot of dns4003.wikimedia.org [12:47:40] (03CR) 10FNegri: "x3 is already in 1022 so I don't think we need it in 1029." [puppet] - 10https://gerrit.wikimedia.org/r/1310547 (https://phabricator.wikimedia.org/T409557) (owner: 10Marostegui) [12:47:55] !log depool codfw pki in dns discovery ahead of lsw1-b5-codfw maintenance T430918 [12:47:58] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [12:47:58] T430918: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918 [12:48:09] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host ms-fe2017.codfw.wmnet with OS trixie [12:48:22] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12119151 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host ms-fe2017.codfw.wmnet with OS trixie [12:48:38] !log elukey@cumin1003 START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet [12:49:03] (03CR) 10Marostegui: "Indeed! It is on both 1022 and 1023" [puppet] - 10https://gerrit.wikimedia.org/r/1310547 (https://phabricator.wikimedia.org/T409557) (owner: 10Marostegui) [12:49:10] FIRING: [2x] BFDdown: BFD session down between cr1-codfw and 2620:0:860:103:10:192:32:58 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [12:49:24] !log blake@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1067.eqiad.wmnet with OS trixie [12:49:39] (03PS2) 10Marostegui: eqiad.yaml: Add clouddb1029 [puppet] - 10https://gerrit.wikimedia.org/r/1310547 (https://phabricator.wikimedia.org/T409557) [12:49:58] !log cmooney@cumin1003 conftool action : set/pooled=false; selector: dnsdisc=pki,name=codfw [12:50:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.56% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [12:50:32] 10ops-eqiad, 06SRE, 06Data-Platform-SRE, 06DC-Ops, 06Machine-Learning-Team: hw troubleshooting: failing disk in ml-serve1001.eqiad.wmnet - https://phabricator.wikimedia.org/T432105#12119155 (10VRiley-WMF) Thanks @cmooney, currently looking into this right now. It's not under warrenty, so I'll need to sou... [12:50:45] (03CR) 10Marostegui: "Sent a new patch removing x3." [puppet] - 10https://gerrit.wikimedia.org/r/1310547 (https://phabricator.wikimedia.org/T409557) (owner: 10Marostegui) [12:51:30] 10ops-eqiad, 06SRE, 06DC-Ops: Degraded RAID on backup1019 - https://phabricator.wikimedia.org/T431367#12119158 (10jcrespo) Please go ahead, if you can do it in a hot way. Otherwise let me know and I can stop the server (it will take me 15 minutes to do so). [12:54:10] FIRING: [4x] BFDdown: BFD session down between cr1-codfw and 2620:0:860:104:10:192:48:14 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [12:54:51] (03PS1) 10Marostegui: mariadb: Remove x3 from clouddb1029 [puppet] - 10https://gerrit.wikimedia.org/r/1310571 (https://phabricator.wikimedia.org/T409557) [12:54:57] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet [12:55:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.27% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [12:55:17] !log cmooney@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b5-codfw,lsw1-b5-codfw IPv6,lsw1-b5-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: switch upgade lsw1-b5-codfw [12:55:25] (03CR) 10Marostegui: "I will remove the instance from zarcillo db" [puppet] - 10https://gerrit.wikimedia.org/r/1310571 (https://phabricator.wikimedia.org/T409557) (owner: 10Marostegui) [12:56:34] (03CR) 10Marostegui: "REmoved from zarcillo DB the instance with x3" [puppet] - 10https://gerrit.wikimedia.org/r/1310571 (https://phabricator.wikimedia.org/T409557) (owner: 10Marostegui) [12:56:53] 10SRE-Access-Requests, 06Data-Engineering, 13Patch-For-Review: Add catherinekelsey to analytics-wmde-users group - https://phabricator.wikimedia.org/T432002#12119186 (10jcrespo) Still blocked on @Gehel approval. [12:57:07] !log cmooney@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 30 hosts with reason: lsw1-b5-codfw JunOS upgrade [12:57:21] blake@cumin1003 renumber-node (PID 4022368) is awaiting input [12:57:38] !log cmooney@cumin1003 START - Cookbook sre.mysql.depool depool db2159: codfw rack B5 depool for maintenance [12:57:54] !log elukey@cumin1003 conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=codfw [12:58:15] !log elukey@cumin1003 conftool action : set/pooled=false; selector: dnsdisc=tegola,name=codfw [12:58:17] !log mvernon@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage [12:58:20] !log cmooney@cumin1003 END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: codfw rack B5 depool for maintenance [12:58:29] (03CR) 10Cathal Mooney: [C:03+1] "LGTM!" [dns] - 10https://gerrit.wikimedia.org/r/1310567 (owner: 10Elukey) [12:58:42] !log elukey@cumin1003 conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=codfw [12:59:07] !log ladsgroup@cumin1003 START - Cookbook sre.hosts.remove-downtime for ms-be1068.eqiad.wmnet [12:59:08] !log ladsgroup@cumin1003 END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1068.eqiad.wmnet [12:59:10] FIRING: [6x] BFDdown: BFD session down between cr1-codfw and 2620:0:860:104:10:192:48:14 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [12:59:14] (03CR) 10FNegri: [C:03+1] mariadb: Remove x3 from clouddb1029 [puppet] - 10https://gerrit.wikimedia.org/r/1310571 (https://phabricator.wikimedia.org/T409557) (owner: 10Marostegui) [12:59:22] (03CR) 10Marostegui: [C:03+2] mariadb: Remove x3 from clouddb1029 [puppet] - 10https://gerrit.wikimedia.org/r/1310571 (https://phabricator.wikimedia.org/T409557) (owner: 10Marostegui) [12:59:56] (03CR) 10FNegri: [C:03+1] eqiad.yaml: Add clouddb1029 [puppet] - 10https://gerrit.wikimedia.org/r/1310547 (https://phabricator.wikimedia.org/T409557) (owner: 10Marostegui) [13:00:05] urbanecm and TheresNoTime: It is that lovely time of the day again! You are hereby commanded to deploy UTC afternoon backport window. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260714T1300). [13:00:05] No Gerrit patches in the queue for this window AFAICS. [13:00:07] 10SRE-Access-Requests, 06Data-Engineering, 13Patch-For-Review: Add catherinekelsey to analytics-wmde-users group - https://phabricator.wikimedia.org/T432002#12119207 (10Gehel) My bad. Approved ! [13:00:58] !log cmooney@cumin1003 START - Cookbook sre.mysql.depool depool db2177: codfw rack B5 depool for maintenance [13:01:11] 10ops-eqiad, 06SRE, 06Data-Platform-SRE, 06DC-Ops, 06Machine-Learning-Team: hw troubleshooting: failing disk in ml-serve1001.eqiad.wmnet - https://phabricator.wikimedia.org/T432105#12119208 (10VRiley-WMF) @cmooney I just checked our invintory of hdds, and it seems like we don't have this one covered. We... [13:01:18] !log cmooney@cumin1003 END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2177: codfw rack B5 depool for maintenance [13:01:22] !log cmooney@cumin1003 START - Cookbook sre.mysql.depool depool db2178: codfw rack B5 depool for maintenance [13:01:30] 10ops-eqiad, 06SRE, 06DC-Ops: Degraded RAID on an-worker1191 - https://phabricator.wikimedia.org/T431828#12119209 (10VRiley-WMF) a:03VRiley-WMF [13:01:39] !log elukey@cumin1003 START - Cookbook sre.hosts.reboot-single for host idm-test1001.wikimedia.org [13:01:42] !log cmooney@cumin1003 END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2178: codfw rack B5 depool for maintenance [13:01:45] !log sukhe@cumin1003 START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_magru [13:01:46] !log cmooney@cumin1003 START - Cookbook sre.mysql.depool depool db2188: codfw rack B5 depool for maintenance [13:02:01] !log sukhe@cumin1003 START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_magru [13:02:07] !log cmooney@cumin1003 END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2188: codfw rack B5 depool for maintenance [13:02:53] 10ops-eqiad, 06SRE, 06DC-Ops: Degraded RAID on backup1019 - https://phabricator.wikimedia.org/T431367#12119226 (10VRiley-WMF) Thanks, I will proceed with that. I should be able to hot swap it. I'll let you know when it's completed. [13:03:16] !log sukhe@cumin1003 START - Cookbook sre.loadbalancer.admin rebooting P{lvs7003*} and A:liberica [13:03:20] (03CR) 10Ottomata: [C:03+1] streams: webrequest-page-view - pageview-relative-trending [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310555 (https://phabricator.wikimedia.org/T430134) (owner: 10JavierMonton) [13:03:54] sukhe@cumin1003 roll-reboot (PID 4048116) is awaiting input [13:04:10] FIRING: [8x] BFDdown: BFD session down between cr1-codfw and 2620:0:860:104:10:192:48:14 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [13:04:10] 10SRE-Access-Requests, 06Data-Engineering, 13Patch-For-Review: Add catherinekelsey to analytics-wmde-users group - https://phabricator.wikimedia.org/T432002#12119228 (10jcrespo) [13:04:42] !log elukey@cumin1003 START - Cookbook sre.hosts.reboot-single for host maps2014.codfw.wmnet [13:04:50] (03CR) 10Jcrespo: [C:03+1] "Approvals ok." [puppet] - 10https://gerrit.wikimedia.org/r/1310121 (https://phabricator.wikimedia.org/T432002) (owner: 10Jcrespo) [13:04:59] (03PS2) 10Jcrespo: admin: Add catherinekelsey to the analytics-wmde-users group [puppet] - 10https://gerrit.wikimedia.org/r/1310121 (https://phabricator.wikimedia.org/T432002) [13:05:04] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage [13:05:32] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage [13:05:36] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm-test1001.wikimedia.org [13:06:28] !log blake@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1067.eqiad.wmnet [13:06:29] !log blake@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1067.eqiad.wmnet [13:06:30] !log blake@cumin1003 END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1067.eqiad.wmnet [13:06:36] !log elukey@cumin1003 START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2003.codfw.wmnet [13:06:55] 06SRE, 06cloud-services-team, 10Cloud-VPS, 06tools-infrastructure-team: Modernise memcached systemd unit / sync, and make it presentable - https://phabricator.wikimedia.org/T273950#12119254 (10jijiki) a:05jijiki→03None [13:07:03] !log sukhe@cumin1003 cookbooks.sre.dns.roll-reboot finished rebooting dns4003.wikimedia.org [13:07:06] 06SRE, 10SRE-swift-storage: Disk near-full warnings on ms swift backends for container filesystems due to some bloated sqlite files - https://phabricator.wikimedia.org/T377827#12119257 (10Ladsgroup) 05Open→03Resolved I'm closing this ticket this time for real. There is no backend with more than 40% use... [13:07:21] !log sukhe@cumin1003 END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting P{lvs7003*} and A:liberica [13:08:12] 06SRE, 06cloud-services-team, 10Cloud-VPS, 06tools-infrastructure-team: Modernise memcached systemd unit / sync, and make it presentable - https://phabricator.wikimedia.org/T273950#12119267 (10taavi) [13:09:10] FIRING: [10x] BFDdown: BFD session down between cr1-codfw and 2620:0:860:104:10:192:48:14 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [13:09:41] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage [13:10:18] (03CR) 10Btullis: [C:03+1] "Looks good to me." [puppet] - 10https://gerrit.wikimedia.org/r/1310121 (https://phabricator.wikimedia.org/T432002) (owner: 10Jcrespo) [13:10:36] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2003.codfw.wmnet [13:11:21] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2014.codfw.wmnet [13:11:45] !log cmooney@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet [13:12:45] FIRING: WidespreadPuppetFailure: Puppet has failed in magru - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [13:12:45] (03CR) 10Marostegui: [C:03+2] eqiad.yaml: Add clouddb1029 [puppet] - 10https://gerrit.wikimedia.org/r/1310547 (https://phabricator.wikimedia.org/T409557) (owner: 10Marostegui) [13:12:58] !log elukey@cumin1003 START - Cookbook sre.hosts.reboot-single for host maps2013.codfw.wmnet [13:13:28] !log elukey@cumin1003 START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2004.codfw.wmnet [13:13:43] !log marostegui@cumin1003 conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s8 [13:13:50] !log marostegui@cumin1003 conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s5 [13:14:10] FIRING: [12x] BFDdown: BFD session down between cr1-codfw and 2620:0:860:104:10:192:48:14 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [13:14:15] !log marostegui@cumin1003 conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s5 [13:14:20] !log marostegui@cumin1003 conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s8 [13:14:29] (03CR) 10Jcrespo: [C:03+2] admin: Add catherinekelsey to the analytics-wmde-users group [puppet] - 10https://gerrit.wikimedia.org/r/1310121 (https://phabricator.wikimedia.org/T432002) (owner: 10Jcrespo) [13:15:50] (03CR) 10Elukey: [C:03+2] Failover url-dowloaders in eqiad and codfw [dns] - 10https://gerrit.wikimedia.org/r/1310567 (owner: 10Elukey) [13:16:18] !log marostegui@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1046: es1046 after reimage [13:16:50] !log elukey@dns1004 START - running authdns-update [13:17:03] (03CR) 10JavierMonton: [C:03+2] streams: webrequest-page-view - pageview-relative-trending [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310555 (https://phabricator.wikimedia.org/T430134) (owner: 10JavierMonton) [13:17:27] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2004.codfw.wmnet [13:17:30] 10SRE-Access-Requests, 06Data-Engineering, 13Patch-For-Review: Add catherinekelsey to analytics-wmde-users group - https://phabricator.wikimedia.org/T432002#12119319 (10jcrespo) 05Open→03Resolved a:03jcrespo @catherine.kelsey.wmde Your access has been deployed. It may take up to 30 minutes to be ap... [13:17:37] !log elukey@cumin1003 START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2005.codfw.wmnet [13:17:58] !log cmooney@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet [13:18:30] 10ops-eqiad, 06SRE, 06DC-Ops: Degraded RAID on backup1019 - https://phabricator.wikimedia.org/T431367#12119325 (10VRiley-WMF) @jcrespo This is completed and it looks like it's currently rebuilding. [13:18:46] !log elukey@dns1004 END - running authdns-update [13:18:47] (03CR) 10Kamila Součková: ml-services/*: don't hardcode helmBinary (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310146 (https://phabricator.wikimedia.org/T388390) (owner: 10Kamila Součková) [13:19:17] (03Merged) 10jenkins-bot: streams: webrequest-page-view - pageview-relative-trending [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310555 (https://phabricator.wikimedia.org/T430134) (owner: 10JavierMonton) [13:19:50] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2013.codfw.wmnet [13:20:34] 10ops-eqiad, 06SRE, 06DC-Ops: Degraded RAID on backup1019 - https://phabricator.wikimedia.org/T431367#12119336 (10jcrespo) Nice! I will keep an eye on the progress. [13:20:49] 10ops-eqiad, 06SRE, 06Data-Platform-SRE, 06DC-Ops, 06Machine-Learning-Team: hw troubleshooting: failing disk in ml-serve1001.eqiad.wmnet - https://phabricator.wikimedia.org/T432105#12119338 (10klausman) >>! In T432105#12119208, @VRiley-WMF wrote: > @cmooney I just checked our invintory of hdds, and it se... [13:21:22] 10ops-eqiad, 06SRE, 06DC-Ops: Degraded RAID on backup1019 - https://phabricator.wikimedia.org/T431367#12119342 (10VRiley-WMF) Awesome, let me know when it's okay to close this! [13:21:37] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2005.codfw.wmnet [13:21:57] (03CR) 10Kamila Součková: [C:03+2] ml-services/*: don't hardcode helmBinary [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310146 (https://phabricator.wikimedia.org/T388390) (owner: 10Kamila Součková) [13:22:03] !log sukhe@cumin1003 cookbooks.sre.dns.roll-reboot begin reboot of dns4004.wikimedia.org [13:22:05] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply [13:22:08] (03PS2) 10Santiago Faci: Test Kitchen UI: Deploying v1.4.8 release to production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310139 (https://phabricator.wikimedia.org/T428984) [13:22:18] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply [13:22:37] !log cmooney@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet [13:23:08] PROBLEM - haproxy process on cp7009 is CRITICAL: PROCS CRITICAL: 0 processes with command name haproxy https://wikitech.wikimedia.org/wiki/HAProxy [13:23:08] PROBLEM - haproxy process on cp7001 is CRITICAL: PROCS CRITICAL: 0 processes with command name haproxy https://wikitech.wikimedia.org/wiki/HAProxy [13:23:10] PROBLEM - HAProxy HTTPS measure-eqiad.wikimedia.org ECDSA on cp7009 is CRITICAL: SSL CRITICAL - failed to connect or SSL handshake:Connection refused https://wikitech.wikimedia.org/wiki/HTTPS [13:23:10] PROBLEM - HAProxy HTTPS upload.wikimedia.org ECDSA on cp7009 is CRITICAL: SSL CRITICAL - failed to connect or SSL handshake:Connection refused https://wikitech.wikimedia.org/wiki/HTTPS [13:23:21] hmm the reboot [13:24:02] PROBLEM - HAProxy HTTPS wikipedia.org ECDSA on cp7001 is CRITICAL: SSL CRITICAL - failed to connect or SSL handshake:Connection refused https://wikitech.wikimedia.org/wiki/HTTPS [13:24:02] PROBLEM - HAProxy HTTPS wikipedia25.org ECDSA on cp7001 is CRITICAL: SSL CRITICAL - failed to connect or SSL handshake:Connection refused https://wikitech.wikimedia.org/wiki/HTTPS [13:24:02] PROBLEM - HAProxy HTTPS wikiworkshop.org ECDSA on cp7001 is CRITICAL: SSL CRITICAL - failed to connect or SSL handshake:Connection refused https://wikitech.wikimedia.org/wiki/HTTPS [13:24:29] !log elukey@cumin1003 START - Cookbook sre.hosts.reboot-single for host maps2012.codfw.wmnet [13:25:45] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1017.eqiad.wmnet with OS trixie [13:25:59] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12119356 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-fe1017.eqiad.wmnet with OS trixie completed... [13:26:42] 10SRE-Access-Requests, 06Data-Engineering, 13Patch-For-Review: Add catherinekelsey to analytics-wmde-users group - https://phabricator.wikimedia.org/T432002#12119368 (10catherine.kelsey.wmde) Thank you very much everyone for the support for this! I also consider resolved, as can see my access updated: ht... [13:28:19] (03Merged) 10jenkins-bot: ml-services/*: don't hardcode helmBinary [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310146 (https://phabricator.wikimedia.org/T388390) (owner: 10Kamila Součková) [13:28:58] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2017.codfw.wmnet with OS trixie [13:29:19] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12119380 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-fe2017.codfw.wmnet with OS trixie completed... [13:29:22] 10ops-eqiad, 06SRE, 06DC-Ops: Degraded RAID on backup1019 - https://phabricator.wikimedia.org/T431367#12119381 (10jcrespo) We can recheck tomorrow (although sometimes there is a speedup to sync zeroes :-D): ` $ perccli64 /c0/e64/s8 show rebuild CLI Version = 007.1910.0000.0000 Oct 08, 2021 Operating system =... [13:29:52] 10ops-eqiad, 06SRE, 06Data-Platform-SRE, 06DC-Ops, 06Machine-Learning-Team: hw troubleshooting: failing disk in ml-serve1001.eqiad.wmnet - https://phabricator.wikimedia.org/T432105#12119383 (10VRiley-WMF) Thanks for the update @klausman (I was going by who last replied, sorry about that!) I'll check wi... [13:30:00] RECOVERY - Check if dnsdist.service has been restarted after /etc/dnsdist/dnsdist.conf was changed on doh5004 is OK: OK: dnsdist.service was restarted after /etc/dnsdist/dnsdist.conf was changed. https://wikitech.wikimedia.org/wiki/Wikidough/Monitoring%23Service_Restart_Check [13:31:08] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2012.codfw.wmnet [13:31:51] !log elukey@cumin1003 START - Cookbook sre.hosts.reboot-single for host maps2011.codfw.wmnet [13:31:53] (03PS8) 10CDobbins: varnish: put testwiki back into enforce mode [puppet] - 10https://gerrit.wikimedia.org/r/1309263 [13:32:19] (03CR) 10Tiziano Fogli: [C:03+1] karma: automatically linkify Phab task IDs in messages [puppet] - 10https://gerrit.wikimedia.org/r/1310141 (https://phabricator.wikimedia.org/T431980) (owner: 10Hnowlan) [13:32:44] !log cmooney@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet [13:34:17] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host ms-fe2016.codfw.wmnet with OS trixie [13:34:18] 10ops-eqiad, 06Data-Persistence, 06DBA, 06DC-Ops, and 2 others: es7 primary (es1039.eqiad.wmnet) crashed - https://phabricator.wikimedia.org/T430764#12119394 (10VRiley-WMF) @Marostegui Hopefully everything is looking good. I will close this ticket for now. Feel free to open it if there is another issue. [13:34:29] 10ops-eqiad, 06Data-Persistence, 06DBA, 06DC-Ops, and 2 others: es7 primary (es1039.eqiad.wmnet) crashed - https://phabricator.wikimedia.org/T430764#12119396 (10VRiley-WMF) 05Open→03Resolved [13:34:31] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12119395 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host ms-fe2016.codfw.wmnet with OS trixie [13:34:40] !log reboot lsw1-b5-codfw to upgrade JunOS T430918 [13:34:40] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-fe1016.eqiad.wmnet with OS trixie [13:34:43] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [13:34:43] T430918: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918 [13:34:53] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12119397 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-fe1016.eqiad.wmnet with OS trixie [13:35:26] 10ops-eqiad, 06DC-Ops: hw troubleshooting: DIMM module in slot A7 for db1208.eqiad.wmnet - https://phabricator.wikimedia.org/T432116 (10BTullis) 03NEW [13:36:13] 10ops-eqiad, 06Data-Persistence, 06DBA, 06DC-Ops, and 2 others: es7 primary (es1039.eqiad.wmnet) crashed - https://phabricator.wikimedia.org/T430764#12119412 (10Marostegui) Thank you for your help! [13:36:14] !log sukhe@cumin1003 cookbooks.sre.dns.roll-reboot finished rebooting dns4004.wikimedia.org [13:36:14] !log sukhe@cumin1003 END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox) [13:36:43] 10ops-eqiad, 06DC-Ops: hw troubleshooting: DIMM module in slot A7 for db1208.eqiad.wmnet - https://phabricator.wikimedia.org/T432116#12119413 (10BTullis) [13:37:02] 10ops-eqiad, 06DC-Ops: hw troubleshooting: DIMM module in slot A7 for db1208.eqiad.wmnet - https://phabricator.wikimedia.org/T432116#12119415 (10BTullis) [13:37:36] 10ops-eqiad, 06SRE, 06DC-Ops: Degraded RAID on backup1019 - https://phabricator.wikimedia.org/T431367#12119420 (10VRiley-WMF) 05Open→03Resolved @jcrespo awesome! I'll go ahead and close this ticket for now. If there are any other issues feel free to open it back up. [13:38:10] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2011.codfw.wmnet [13:38:51] FIRING: [3x] ProbeDown: Service restbase2036-a:9042 has failed probes (tcp_cassandra_a_cql_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [13:39:39] FIRING: CoreBGPDown: Core BGP session down between ssw1-a8-codfw and lsw1-b5-codfw (10.192.252.14) - group EVPN_IBGP - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=ssw1-a8-codfw:9804&var-bgp_group=EVPN_IBGP&var-bgp_neighbor=lsw1-b5-codfw - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [13:39:51] FIRING: [2x] SwitchCoreInterfaceDown: Switch core interface down - ssw1-a1-codfw:et-0/0/11 (Core: lsw1-b5-codfw:et-0/0/55 {#230403800004}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Switch_interface_down - https://alerts.wikimedia.org/?q=alertname%3DSwitchCoreInterfaceDown [13:40:58] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [13:42:28] !log sukhe@cumin1003 END (ERROR) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=97) rolling reboot on A:durum and A:durum [13:42:43] 06SRE, 06cloud-services-team, 06collaboration-services, 10Infrastructure Security, and 7 others: Reboot cirrussearch in CODFW - https://phabricator.wikimedia.org/T432119 (10atsuko) 03NEW [13:43:13] 06SRE, 06cloud-services-team, 06collaboration-services, 10Infrastructure Security, and 7 others: Reboot cirrussearch in CODFW - https://phabricator.wikimedia.org/T432119#12119478 (10atsuko) [13:43:28] !log sukhe@cumin1003 START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum [13:43:51] FIRING: [8x] ProbeDown: Service restbase2036-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [13:44:10] FIRING: [14x] BFDdown: BFD session down between cr1-codfw and 10.192.16.35 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [13:44:39] FIRING: [2x] CoreBGPDown: Core BGP session down between ssw1-a1-codfw and lsw1-b5-codfw (10.192.252.14) - group EVPN_IBGP - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [13:45:15] 10ops-eqiad, 06DC-Ops: hw troubleshooting: DIMM module in slot A7 for db1208.eqiad.wmnet - https://phabricator.wikimedia.org/T432116#12119489 (10VRiley-WMF) a:03VRiley-WMF [13:45:43] FIRING: [3x] JobUnavailable: Reduced availability for job cfssl in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [13:46:45] !log sukhe@cumin1003 END (PASS) - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy (exit_code=0) rolling reboot on A:tcpproxy and A:tcpproxy [13:48:43] FIRING: RedisInstanceDown: Redis instance down arclamp2001:9121 redis_arclamp - https://wikitech.wikimedia.org/wiki/Redis#Cluster_redis_arclamp - https://grafana.wikimedia.org/d/000000174/redis?orgId=1&var-site=codfw&var-job=redis_arclamp&var-instance=arclamp2001:9121 - https://alerts.wikimedia.org/?q=alertname%3DRedisInstanceDown [13:48:56] FIRING: [15x] RedisInstanceDown: Redis instance down rdb2011:16378 redis_misc - https://wikitech.wikimedia.org/wiki/Redis#Cluster_redis_misc - https://alerts.wikimedia.org/?q=alertname%3DRedisInstanceDown [13:49:05] FIRING: [9x] ProbeDown: Service wikikube-ctrl2001:6443 has failed probes (http_codfw_kube_apiserver_ip4) #page - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [13:49:07] !log sukhe@cumin1003 END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling reboot on A:wikidough [13:49:08] o/ [13:49:13] FIRING: [7x] SwaggerProbeHasFailures: Not all openapi/swagger endpoints returned healthy - https://grafana.wikimedia.org/d/_77ik484k/openapi-swagger-endpoint-state?var-site=codfw - https://alerts.wikimedia.org/?q=alertname%3DSwaggerProbeHasFailures [13:49:14] !incidents [13:49:15] 8190 (UNACKED) [9x] ProbeDown sre (probes/custom codfw) [13:49:15] 8191 (UNACKED) EtcdReplicationDown etcd sre (conf2005:8000 etcdmirror codfw) [13:49:22] FIRING: [17x] SwaggerProbeHasFailures: Not all openapi/swagger endpoints returned healthy - https://grafana.wikimedia.org/d/_77ik484k/openapi-swagger-endpoint-state?var-site=codfw - https://alerts.wikimedia.org/?q=alertname%3DSwaggerProbeHasFailures [13:49:31] FIRING: [2x] ProbeDown: Service doc1004.eqiad.wmnet:443 has failed probes (http_doc1004_eqiad_wmnet_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [13:49:37] !log sukhe@cumin1003 END (PASS) - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy (exit_code=0) rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy [13:49:41] FIRING: [2x] ProbeDown: Service idp2005:443 has failed probes (http_idp_wikimedia_org_ip4) - https://wikitech.wikimedia.org/wiki/CAS-SSO#Alerting - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [13:49:51] FIRING: RedisInstanceDown: Redis instance down gitlab2002:9121 redis_gitlab - https://wikitech.wikimedia.org/wiki/Redis#Cluster_redis_gitlab - https://grafana.wikimedia.org/d/000000174/redis?orgId=1&var-site=codfw&var-job=redis_gitlab&var-instance=gitlab2002:9121 - https://alerts.wikimedia.org/?q=alertname%3DRedisInstanceDown [13:49:56] FIRING: [2x] ProbeDown: Service gerrit2003:443 has failed probes (http_gerrit_tls_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#gerrit2003:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [13:50:12] FIRING: EtcdReplicationDown: etcd replication down on conf2005:8000 #page - https://wikitech.wikimedia.org/wiki/Etcd/Main_cluster#Replication - TODO - https://alerts.wikimedia.org/?q=alertname%3DEtcdReplicationDown [13:50:30] !ack 8190 [13:50:31] 8190 (ACKED) [9x] ProbeDown sre (probes/custom codfw) [13:50:32] !ack 8191 [13:50:33] 8191 (ACKED) EtcdReplicationDown etcd sre (conf2005:8000 etcdmirror codfw) [13:50:35] FIRING: [253x] ProbeDown: Service aqs2001-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [13:50:45] !log mvernon@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage [13:50:56] FIRING: [14x] BFDdown: BFD session down between cr1-codfw and 10.192.16.35 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [13:51:01] FIRING: [4x] ProbeDown: Service idm2001:443 has failed probes (http_idm_wikimedia_org_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [13:51:06] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage [13:51:07] (03CR) 10CDobbins: "upload: 0 tests failed, 0 tests skipped, 21 tests passed" [puppet] - 10https://gerrit.wikimedia.org/r/1309263 (owner: 10CDobbins) [13:51:14] FIRING: [2x] CoreBGPDown: Core BGP session down between ssw1-a1-codfw and lsw1-b5-codfw (10.192.252.14) - group EVPN_IBGP - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [13:51:24] FIRING: [2x] SwitchCoreInterfaceDown: Switch core interface down - ssw1-a1-codfw:et-0/0/11 (Core: lsw1-b5-codfw:et-0/0/55 {#230403800004}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Switch_interface_down - https://alerts.wikimedia.org/?q=alertname%3DSwitchCoreInterfaceDown [13:51:35] FIRING: [50x] CertAlmostExpired: gNMI TLS certificate for cloudsw1-b1-codfw.mgmt.codfw.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [13:51:39] FIRING: [103x] JobUnavailable: Reduced availability for job alertmanager in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [13:52:10] jhathaway: codfw issues is the B5 maintenance? [13:52:18] !log cmooney@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet [13:52:25] !log cmooney@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet [13:52:48] dcaro: ah thanks, that makes sense [13:52:58] dcaro: o/ the hosts in B5 are not the ones that alarmed though https://netbox.wikimedia.org/dcim/racks/55/ [13:53:02] (03CR) 10Sbisson: [C:03+2] Update cxserver to 2026-07-10-150240-production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310042 (https://phabricator.wikimedia.org/T290730) (owner: 10Nik Gkountas) [13:53:11] yep, we got some alert for things not in B5 either [13:53:32] topranks: ^ is it related? [13:53:40] he said no on other channel [13:53:43] RESOLVED: RedisInstanceDown: Redis instance down arclamp2001:9121 redis_arclamp - https://wikitech.wikimedia.org/wiki/Redis#Cluster_redis_arclamp - https://grafana.wikimedia.org/d/000000174/redis?orgId=1&var-site=codfw&var-job=redis_arclamp&var-instance=arclamp2001:9121 - https://alerts.wikimedia.org/?q=alertname%3DRedisInstanceDown [13:53:46] ack [13:53:56] I'll let you handle then [13:54:00] RESOLVED: [20x] RedisInstanceDown: Redis instance down rdb2011:16378 redis_misc - https://wikitech.wikimedia.org/wiki/Redis#Cluster_redis_misc - https://alerts.wikimedia.org/?q=alertname%3DRedisInstanceDown [13:54:04] RESOLVED: [64x] ProbeDown: Service aux-k8s-ctrl2002:6443 has failed probes (http_aux_k8s_codfw_kube_apiserver_ip4) #page - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [13:54:22] RESOLVED: [17x] SwaggerProbeHasFailures: Not all openapi/swagger endpoints returned healthy - https://grafana.wikimedia.org/d/_77ik484k/openapi-swagger-endpoint-state?var-site=codfw - https://alerts.wikimedia.org/?q=alertname%3DSwaggerProbeHasFailures [13:54:29] (03CR) 10Mforns: [C:03+1] "LGTM!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310555 (https://phabricator.wikimedia.org/T430134) (owner: 10JavierMonton) [13:54:31] RESOLVED: [25x] ProbeDown: Service doc1004.eqiad.wmnet:443 has failed probes (http_doc1004_eqiad_wmnet_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [13:54:46] RESOLVED: [4x] ProbeDown: Service idm2001:443 has failed probes (http_idm_wikimedia_org_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [13:54:51] RESOLVED: RedisInstanceDown: Redis instance down gitlab2002:9121 redis_gitlab - https://wikitech.wikimedia.org/wiki/Redis#Cluster_redis_gitlab - https://grafana.wikimedia.org/d/000000174/redis?orgId=1&var-site=codfw&var-job=redis_gitlab&var-instance=gitlab2002:9121 - https://alerts.wikimedia.org/?q=alertname%3DRedisInstanceDown [13:54:55] RESOLVED: [4x] ProbeDown: Service gerrit2003:29418 has failed probes (tcp_gerrit_ssh_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [13:54:57] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage [13:55:08] RESOLVED: EtcdReplicationDown: etcd replication down on conf2005:8000 #page - https://wikitech.wikimedia.org/wiki/Etcd/Main_cluster#Replication - TODO - https://alerts.wikimedia.org/?q=alertname%3DEtcdReplicationDown [13:55:16] (03Merged) 10jenkins-bot: Update cxserver to 2026-07-10-150240-production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310042 (https://phabricator.wikimedia.org/T290730) (owner: 10Nik Gkountas) [13:55:25] RESOLVED: [297x] ProbeDown: Service aqs2001-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [13:55:33] FIRING: [14x] BFDdown: BFD session down between cr1-codfw and 10.192.16.35 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [13:55:38] RESOLVED: [2x] CoreBGPDown: Core BGP session down between ssw1-a1-codfw and lsw1-b5-codfw (10.192.252.14) - group EVPN_IBGP - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [13:55:43] RESOLVED: [2x] SwitchCoreInterfaceDown: Switch core interface down - ssw1-a1-codfw:et-0/0/11 (Core: lsw1-b5-codfw:et-0/0/55 {#230403800004}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Switch_interface_down - https://alerts.wikimedia.org/?q=alertname%3DSwitchCoreInterfaceDown [13:56:04] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 - https://phabricator.wikimedia.org/T431682#12119596 (10fgiunchedi) I'm going through https://wikitech.wikimedia.org/wiki/Server_Lifecycle#Move_existing_server_between_rows/racks,_changing_IPs and figuring out ownership... [13:56:17] (03CR) 10CDobbins: [V:03+1] "PCC SUCCESS (CORE_DIFF 2): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/8995/co" [puppet] - 10https://gerrit.wikimedia.org/r/1310160 (owner: 10CDobbins) [13:56:41] !log sbisson@deploy2003 helmfile [staging] START helmfile.d/services/cxserver: sync [13:57:02] !log sbisson@deploy2003 helmfile [staging] DONE helmfile.d/services/cxserver: sync [13:57:04] !log sukhe@cumin1003 END (PASS) - Cookbook sre.cdn.roll-restart-reboot-ncredir (exit_code=0) rolling reboot on A:ncredir and A:ncredir [13:58:09] (03CR) 10CDobbins: "https://puppet-compiler.wmflabs.org/output/1310160/8995/dns7001.wikimedia.org/fulldiff.html" [puppet] - 10https://gerrit.wikimedia.org/r/1310160 (owner: 10CDobbins) [13:58:20] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage [13:58:23] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 - https://phabricator.wikimedia.org/T431682#12119607 (10VRiley-WMF) Thanks @fgiunchedi! I will be running through this today. I just got back from vacation and catching up on a few different things here. I will let you... [14:00:04] Deploy window Test Kitchen UI Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260714T1400) [14:00:08] !log elukey@cumin1003 START - Cookbook sre.hosts.reboot-single for host urldownloader2005.wikimedia.org [14:00:57] !log elukey@cumin1003 conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=codfw [14:01:05] !log elukey@cumin1003 conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=codfw [14:02:39] !log btullis@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet [14:02:41] !log btullis@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet [14:03:46] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 - https://phabricator.wikimedia.org/T431682#12119651 (10fgiunchedi) >>! In T431682#12119607, @VRiley-WMF wrote: > Thanks @fgiunchedi! I will be running through this today. I just got back from vacation and catching up o... [14:04:40] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2005.wikimedia.org [14:05:25] !log elukey@cumin1003 START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org [14:07:20] !log sbisson@deploy2003 helmfile [eqiad] START helmfile.d/services/cxserver: sync [14:07:30] (03PS1) 10Elukey: Revert "Failover url-dowloaders in eqiad and codfw" [dns] - 10https://gerrit.wikimedia.org/r/1310578 [14:07:53] !log sbisson@deploy2003 helmfile [eqiad] DONE helmfile.d/services/cxserver: sync [14:09:52] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org [14:11:21] 10ops-eqiad, 06SRE, 06DC-Ops: Degraded RAID on an-worker1191 - https://phabricator.wikimedia.org/T431828#12119674 (10VRiley-WMF) Opened and submitted a ticket for this server. Dell SR SR229034753 [14:11:39] !log sbisson@deploy2003 helmfile [codfw] START helmfile.d/services/cxserver: sync [14:11:45] Hi! I want to run a maintenance script to add wikidata support for a new language wiki. Let me know if this is a bad time, otherwise I will proceed. [14:11:51] !log cmooney@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet [14:11:52] !log cmooney@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet [14:12:07] !log sbisson@deploy2003 helmfile [codfw] DONE helmfile.d/services/cxserver: sync [14:12:27] !log cmooney@cumin1003 conftool action : set/pooled=true; selector: dnsdisc=pki,name=codfw [14:12:40] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1016.eqiad.wmnet with OS trixie [14:13:05] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12119711 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-fe1016.eqiad.wmnet with OS trixie completed... [14:14:34] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-fe1015.eqiad.wmnet with OS trixie [14:14:48] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12119736 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-fe1015.eqiad.wmnet with OS trixie [14:14:53] !log cmooney@cumin1003 START - Cookbook sre.mysql.pool pool db2159: repooling after rack b5 maintenance [14:15:34] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2016.codfw.wmnet with OS trixie [14:15:49] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12119738 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-fe2016.codfw.wmnet with OS trixie completed... [14:16:10] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host ms-fe2015.codfw.wmnet with OS trixie [14:16:22] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12119749 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host ms-fe2015.codfw.wmnet with OS trixie [14:16:41] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12119754 (10MatthewVernon) [14:19:31] (03CR) 10SBassett: [C:03+1] varnish: put testwiki back into enforce mode [puppet] - 10https://gerrit.wikimedia.org/r/1309263 (owner: 10CDobbins) [14:21:38] (03PS15) 10Federico Ceratto: cookbooks/sre/mysql/decommission: add cookbook [cookbooks] - 10https://gerrit.wikimedia.org/r/1291952 (https://phabricator.wikimedia.org/T426613) [14:24:57] !log sukhe@cumin1003 END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum [14:30:05] Deploy window Test Kitchen Experiment Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260714T1430) [14:30:47] !log seanleong-wmde@deploy2003 mwscript-k8s job started: foreachwikiindblist wikidataclient extensions/Wikibase/lib/maintenance/populateSitesTable.php --force-protocol https # T429939 [14:30:51] T429939: ✨Add Wikidata support for isvwiki (🏆) - https://phabricator.wikimedia.org/T429939 [14:30:53] !log mvernon@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage [14:32:25] (03CR) 10Cathal Mooney: [C:03+2] pki: add profile::server_depool to document depool actions [puppet] - 10https://gerrit.wikimedia.org/r/1310554 (https://phabricator.wikimedia.org/T327300) (owner: 10Cathal Mooney) [14:33:16] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage [14:33:27] PROBLEM - Host cloudweb1003 is DOWN: PING CRITICAL - Packet loss = 100% [14:33:51] (03CR) 10CWilliams: [C:03+1] "Seems OK for now" [cookbooks] - 10https://gerrit.wikimedia.org/r/1295480 (https://phabricator.wikimedia.org/T422361) (owner: 10Federico Ceratto) [14:33:55] RECOVERY - Host cloudweb1003 is UP: PING OK - Packet loss = 0%, RTA = 0.34 ms [14:34:35] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage [14:35:53] (03CR) 10Ottomata: [C:03+2] EventStreams - Expose mediawiki.page_revert_risk_wikidata_prediction_change.v1 stream [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310570 (https://phabricator.wikimedia.org/T420883) (owner: 10AikoChou) [14:36:01] PROBLEM - SSH on cloudweb1004 is CRITICAL: connect to address 208.80.155.117 and port 22: Connection refused https://wikitech.wikimedia.org/wiki/SSH/monitoring [14:36:55] RECOVERY - SSH on cloudweb1004 is OK: SSH OK - OpenSSH_10.0p2 Debian-7+deb13u4 (protocol 2.0) https://wikitech.wikimedia.org/wiki/SSH/monitoring [14:37:08] (03PS1) 10JavierMonton: stream: pageview.v1 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310587 (https://phabricator.wikimedia.org/T425624) [14:38:11] (03CR) 10Ottomata: [C:03+1] stream: pageview.v1 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310587 (https://phabricator.wikimedia.org/T425624) (owner: 10JavierMonton) [14:38:43] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage [14:38:45] (03Merged) 10jenkins-bot: EventStreams - Expose mediawiki.page_revert_risk_wikidata_prediction_change.v1 stream [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310570 (https://phabricator.wikimedia.org/T420883) (owner: 10AikoChou) [14:39:02] (03PS1) 10Cathal Mooney: Pometheus nodes: adjust message on server_depool key [puppet] - 10https://gerrit.wikimedia.org/r/1310588 (https://phabricator.wikimedia.org/T327300) [14:39:31] (03CR) 10Ssingh: "Looks good, I will +1 when we are ready to merge." [puppet] - 10https://gerrit.wikimedia.org/r/1310160 (owner: 10CDobbins) [14:39:39] !log otto@deploy2003 helmfile [staging] START helmfile.d/services/eventstreams: apply [14:40:13] !log otto@deploy2003 helmfile [staging] DONE helmfile.d/services/eventstreams: apply [14:40:15] (03CR) 10Ottomata: [C:03+2] "Deploying! thank you Aiko!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310570 (https://phabricator.wikimedia.org/T420883) (owner: 10AikoChou) [14:40:20] !log otto@deploy2003 helmfile [codfw] START helmfile.d/services/eventstreams: apply [14:40:24] (03PS1) 10Kamila Součková: modules/rsync: add pidfile to stunnel.conf [puppet] - 10https://gerrit.wikimedia.org/r/1310589 (https://phabricator.wikimedia.org/T432108) [14:41:02] (03CR) 10A-pizzata: [C:03+1] stream: pageview.v1 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310587 (https://phabricator.wikimedia.org/T425624) (owner: 10JavierMonton) [14:41:03] !log otto@deploy2003 helmfile [codfw] DONE helmfile.d/services/eventstreams: apply [14:41:14] !log otto@deploy2003 helmfile [eqiad] START helmfile.d/services/eventstreams: apply [14:42:05] !log otto@deploy2003 helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply [14:43:18] (03CR) 10Hnowlan: [C:03+1] Pometheus nodes: adjust message on server_depool key [puppet] - 10https://gerrit.wikimedia.org/r/1310588 (https://phabricator.wikimedia.org/T327300) (owner: 10Cathal Mooney) [14:44:26] (03CR) 10TrainBranchBot: [C:03+2] "Approved by javiermonton@deploy2003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310587 (https://phabricator.wikimedia.org/T425624) (owner: 10JavierMonton) [14:45:49] (03Merged) 10jenkins-bot: stream: pageview.v1 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310587 (https://phabricator.wikimedia.org/T425624) (owner: 10JavierMonton) [14:46:08] !log javiermonton@deploy2003 Started scap sync-world: Backport for [[gerrit:1310587|stream: pageview.v1 (T425624)]] [14:46:11] T425624: Relative Trending - Flink app for page_view - https://phabricator.wikimedia.org/T425624 [14:46:34] (03CR) 10Kamila Součková: "I think this is the cause of the repeated stunnel unhappiness on deploy2003." [puppet] - 10https://gerrit.wikimedia.org/r/1310589 (https://phabricator.wikimedia.org/T432108) (owner: 10Kamila Součková) [14:48:08] !log javiermonton@deploy2003 javiermonton: Backport for [[gerrit:1310587|stream: pageview.v1 (T425624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [14:49:25] !log javiermonton@deploy2003 javiermonton: Continuing with deployment [14:50:12] (03CR) 10Filippo Giunchedi: Pometheus nodes: adjust message on server_depool key (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1310588 (https://phabricator.wikimedia.org/T327300) (owner: 10Cathal Mooney) [14:53:30] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1015.eqiad.wmnet with OS trixie [14:53:43] !log javiermonton@deploy2003 Finished scap sync-world: Backport for [[gerrit:1310587|stream: pageview.v1 (T425624)]] (duration: 07m 35s) [14:53:44] (03PS2) 10Kamila Součková: modules/rsync: add pidfile to stunnel.conf [puppet] - 10https://gerrit.wikimedia.org/r/1310589 (https://phabricator.wikimedia.org/T432108) [14:53:46] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12119947 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-fe1015.eqiad.wmnet with OS trixie completed... [14:53:47] T425624: Relative Trending - Flink app for page_view - https://phabricator.wikimedia.org/T425624 [14:53:51] (03CR) 10Kamila Součková: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1310589 (https://phabricator.wikimedia.org/T432108) (owner: 10Kamila Součková) [14:54:32] !log Finished populateSitesTable for isvwiki (T429939) [14:54:35] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:54:35] T429939: ✨Add Wikidata support for isvwiki (🏆) - https://phabricator.wikimedia.org/T429939 [14:55:00] (03PS3) 10Kamila Součková: modules/rsync: add pidfile to stunnel.conf [puppet] - 10https://gerrit.wikimedia.org/r/1310589 (https://phabricator.wikimedia.org/T432108) [14:57:13] Hi, I have finished running the maintenance script, thanks! [14:57:20] (03PS2) 10Cathal Mooney: Pometheus nodes: adjust message on server_depool key [puppet] - 10https://gerrit.wikimedia.org/r/1310588 (https://phabricator.wikimedia.org/T327300) [14:57:29] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2015.codfw.wmnet with OS trixie [14:57:42] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12119987 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-fe2015.codfw.wmnet with OS trixie completed... [14:57:55] (03CR) 10Cathal Mooney: "@fgiunchedi@wikimedia.org thanks for the feedback, does the updated message make sense?" [puppet] - 10https://gerrit.wikimedia.org/r/1310588 (https://phabricator.wikimedia.org/T327300) (owner: 10Cathal Mooney) [14:59:56] !log dancy@deploy2003 Installing scap version "4.274.1" for 3 host(s) [15:00:02] (03PS4) 10Blake: mw-pretrain: Basic Ingress configuration. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309641 (https://phabricator.wikimedia.org/T427668) [15:00:05] jelto, arnoldokoth, mutante, and arnaudb: #bothumor I � Unicode. All rise for SRE Collaboration Services office hours deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260714T1500). [15:00:20] !log cmooney@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: repooling after rack b5 maintenance [15:00:24] !log cmooney@cumin1003 START - Cookbook sre.mysql.pool pool db2177: repooling after rack b5 maintenance [15:00:34] (03CR) 10Blake: mw-pretrain: Basic Ingress configuration. (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309641 (https://phabricator.wikimedia.org/T427668) (owner: 10Blake) [15:01:55] !log dancy@deploy2003 Installation of scap version "4.274.1" completed for 3 hosts [15:04:35] (03CR) 10BCornwall: [C:03+1] Revert "Failover url-dowloaders in eqiad and codfw" [dns] - 10https://gerrit.wikimedia.org/r/1310578 (owner: 10Elukey) [15:09:32] (03CR) 10BryanDavis: webservice-runner: default to 8000 only if PORT and TOOL_WEB_PORT has no value (031 comment) [docker-images/toollabs-images] - 10https://gerrit.wikimedia.org/r/1310212 (https://phabricator.wikimedia.org/T432078) (owner: 10Raymond Ndibe) [15:10:43] FIRING: [2x] JobUnavailable: Reduced availability for job atlas_exporter in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [15:12:21] 10ops-drmrs, 06SRE, 06Infrastructure-Foundations, 10netops: Move TATA IP Transit circuit 1243318 cross-connect from rack B12 to rack B13 - https://phabricator.wikimedia.org/T431605#12120098 (10cmooney) 05Open→03Declined As it turns out dc-ops report that the panel in B13 is also full. So we have o... [15:12:52] (03CR) 10Filippo Giunchedi: Pometheus nodes: adjust message on server_depool key (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1310588 (https://phabricator.wikimedia.org/T327300) (owner: 10Cathal Mooney) [15:17:14] (03PS2) 10Effie Mouzeli: Update to upstream v2026.07.06.00 [debs/mcrouter] - 10https://gerrit.wikimedia.org/r/1281942 (https://phabricator.wikimedia.org/T425255) [15:17:14] (03PS1) 10Effie Mouzeli: mcrouter: add helper script to get release commits [debs/mcrouter] - 10https://gerrit.wikimedia.org/r/1310591 (https://phabricator.wikimedia.org/T425255) [15:18:43] (03CR) 10Effie Mouzeli: "done!" [debs/mcrouter] - 10https://gerrit.wikimedia.org/r/1281942 (https://phabricator.wikimedia.org/T425255) (owner: 10Effie Mouzeli) [15:18:49] (03PS1) 10Kamila Součková: Apply title-related policies when selecting the name of the entity [software/hiddenparma/deploy] - 10https://gerrit.wikimedia.org/r/1310592 [15:19:30] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-fe1014.eqiad.wmnet with OS trixie [15:19:47] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12120137 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-fe1014.eqiad.wmnet with OS trixie [15:19:47] (03CR) 10Scott French: [C:03+1] Apply title-related policies when selecting the name of the entity [software/hiddenparma/deploy] - 10https://gerrit.wikimedia.org/r/1310592 (owner: 10Kamila Součková) [15:20:03] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host ms-fe2014.codfw.wmnet with OS trixie [15:20:18] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12120139 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host ms-fe2014.codfw.wmnet with OS trixie [15:20:21] !log mvernon@cumin2003 START - Cookbook sre.hosts.move-vlan for host ms-fe2014 [15:20:25] (03CR) 10Kamila Součková: [C:03+2] Apply title-related policies when selecting the name of the entity [software/hiddenparma/deploy] - 10https://gerrit.wikimedia.org/r/1310592 (owner: 10Kamila Součková) [15:20:29] !log mvernon@cumin2003 START - Cookbook sre.dns.netbox [15:20:31] (03CR) 10Kamila Součková: [V:03+2 C:03+2] Apply title-related policies when selecting the name of the entity [software/hiddenparma/deploy] - 10https://gerrit.wikimedia.org/r/1310592 (owner: 10Kamila Součková) [15:22:45] !log kamila@cumin1003 START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" [15:22:48] !log kamila@cumin1003 START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 [15:23:40] !log kamila@cumin1003 END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003 [15:23:42] !log kamila@cumin1003 END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003" [15:26:35] mvernon@cumin2003 reimage (PID 1449315) is awaiting input [15:28:24] !log mvernon@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" [15:28:28] !log mvernon@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003" [15:28:29] !log mvernon@cumin2003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [15:28:29] !log mvernon@cumin2003 START - Cookbook sre.dns.wipe-cache ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [15:28:32] !log mvernon@cumin2003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [15:28:34] !log mvernon@cumin2003 START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2014 [15:28:47] !log mvernon@cumin2003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2014 [15:28:48] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2014 [15:32:00] (03CR) 10BCornwall: [C:03+1] "Verified on a call" [puppet] - 10https://gerrit.wikimedia.org/r/1310133 (owner: 10Ssingh) [15:32:15] (03PS1) 10Krinkle: Set $wgMathInternalRestbaseURL explicitly [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310596 (https://phabricator.wikimedia.org/T349582) [15:32:32] 10SRE-SLO: Add alerts to detect missing sli error ratio rate metrics - https://phabricator.wikimedia.org/T432140 (10tappof) 03NEW [15:36:02] !log mvernon@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1014.eqiad.wmnet with reason: host reimage [15:38:10] (03PS2) 10Gerrit maintenance bot: mariadb: Promote db2241 to x3 master [puppet] - 10https://gerrit.wikimedia.org/r/1307084 (https://phabricator.wikimedia.org/T430925) [15:38:41] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1014.eqiad.wmnet with reason: host reimage [15:40:43] FIRING: [2x] JobUnavailable: Reduced availability for job atlas_exporter in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [15:42:22] FIRING: CoreRouterInterfaceDown: Core router interface down - cr3-ulsfo:et-0/0/0 (Transport: Hurricane Electric (dc4841.sfo1) {#changeme_ulsfo_he_cct}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr3-ulsfo:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [15:45:18] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2014.codfw.wmnet with reason: host reimage [15:45:43] FIRING: [2x] JobUnavailable: Reduced availability for job atlas_exporter in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [15:45:52] !log cmooney@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2177: repooling after rack b5 maintenance [15:45:56] !log cmooney@cumin1003 START - Cookbook sre.mysql.pool pool db2178: repooling after rack b5 maintenance [15:48:07] (03PS4) 10DLynch: Add script to get constructive edits for all wikis [puppet] - 10https://gerrit.wikimedia.org/r/1272633 (https://phabricator.wikimedia.org/T428490) (owner: 10Clare Ming) [15:49:36] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2014.codfw.wmnet with reason: host reimage [15:52:45] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12120345 (10MatthewVernon) [15:56:53] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1014.eqiad.wmnet with OS trixie [15:57:09] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12120357 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-fe1014.eqiad.wmnet with OS trixie completed... [15:57:32] (03PS1) 10Scott French: php8.3: Rebuild to pick up new PHP packages (8.3.32) [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1310598 (https://phabricator.wikimedia.org/T431099) [15:57:45] FIRING: [3x] WidespreadPuppetFailure: Puppet has failed in drmrs - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [15:58:24] that doesn't sound good [16:00:05] jhathaway and rzl: Time to do the Puppet request window deploy. Don't look at me like that. You signed up for it. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260714T1600). [16:00:05] No Gerrit patches in the queue for this window AFAICS. [16:01:56] PROBLEM - jenkins_service_running on contint1003 is CRITICAL: PROCS CRITICAL: 2 processes with regex args .*/bin/java .*-jar /usr/share/java/jenkins.war https://wikitech.wikimedia.org/wiki/Jenkins [16:02:45] FIRING: [4x] WidespreadPuppetFailure: Puppet has failed in drmrs - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [16:02:56] RECOVERY - jenkins_service_running on contint1003 is OK: PROCS OK: 1 process with regex args .*/bin/java .*-jar /usr/share/java/jenkins.war https://wikitech.wikimedia.org/wiki/Jenkins [16:03:56] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2014.codfw.wmnet with OS trixie [16:04:08] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12120393 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-fe2014.codfw.wmnet with OS trixie completed... [16:05:58] PROBLEM - Swift https backend on ms-fe2014 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/Swift [16:07:01] !log mvernon@cumin1003 START - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies rolling restart_daemons on A:thanos-fe [16:07:31] !log sukhe@puppetserver1001 conftool action : set/pooled=no; selector: name=cp4039.ulsfo.wmnet [16:07:40] FIRING: SystemdUnitFailed: dump_proxy_ranges.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [16:07:45] FIRING: [5x] WidespreadPuppetFailure: Puppet has failed in drmrs - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [16:07:48] RECOVERY - Swift https backend on ms-fe2014 is OK: HTTP OK: HTTP/1.1 200 OK - 572 bytes in 0.194 second response time https://wikitech.wikimedia.org/wiki/Swift [16:09:50] !log mvernon@cumin2003 START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw [16:10:09] !log mvernon@cumin1003 END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies (exit_code=0) rolling restart_daemons on A:thanos-fe [16:10:43] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:12:42] (03CR) 10Effie Mouzeli: [C:03+1] php8.3: Rebuild to pick up new PHP packages (8.3.32) [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1310598 (https://phabricator.wikimedia.org/T431099) (owner: 10Scott French) [16:15:43] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:16:29] (03CR) 10BCornwall: [C:03+1] ip_reputation_vendors: add basic safeguard to spur downloader [puppet] - 10https://gerrit.wikimedia.org/r/1307752 (owner: 10Fabfur) [16:16:30] (03PS1) 10Btullis: istio: expose a TLS passthrough port for PostgreSQL on dse-k8s [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310601 (https://phabricator.wikimedia.org/T432104) [16:16:33] (03PS1) 10Btullis: cloudnative-pg-cluster: omit empty backup encryption from Cluster manifests [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310602 (https://phabricator.wikimedia.org/T432104) [16:16:37] (03PS1) 10Btullis: cloudnative-pg-cluster: add SNI-based TLS passthrough ingress [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310603 (https://phabricator.wikimedia.org/T432104) [16:16:40] PROBLEM - HAProxy HTTPS wikiworkshop.org ECDSA on cp4039 is CRITICAL: SSL CRITICAL - failed to connect or SSL handshake:Connection refused https://wikitech.wikimedia.org/wiki/HTTPS [16:17:10] !log mvernon@cumin2003 conftool action : set/pooled=inactive; selector: name=ms-fe2014.codfw.wmnet [16:17:16] PROBLEM - HAProxy HTTPS wikipedia25.org ECDSA on cp4039 is CRITICAL: SSL CRITICAL - failed to connect or SSL handshake:Connection refused https://wikitech.wikimedia.org/wiki/HTTPS [16:17:16] PROBLEM - HAProxy HTTPS wikipedia.org ECDSA on cp4039 is CRITICAL: SSL CRITICAL - failed to connect or SSL handshake:Connection refused https://wikitech.wikimedia.org/wiki/HTTPS [16:17:18] PROBLEM - haproxy process on cp4039 is CRITICAL: PROCS CRITICAL: 0 processes with command name haproxy https://wikitech.wikimedia.org/wiki/HAProxy [16:18:05] !log mvernon@cumin2003 END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw [16:18:19] !log mvernon@cumin2003 conftool action : set/pooled=yes; selector: name=ms-fe2014.codfw.wmnet [16:18:59] (03CR) 10BCornwall: [C:03+2] ip_reputation_vendors: add basic safeguard to spur downloader [puppet] - 10https://gerrit.wikimedia.org/r/1307752 (owner: 10Fabfur) [16:21:26] (03CR) 10BCornwall: [C:03+2] ip_reputation_vendors: add basic safeguard to spur downloader (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1307752 (owner: 10Fabfur) [16:28:17] !log blake@deploy2003 helmfile [codfw] START helmfile.d/services/mw-pretrain: apply [16:28:42] !log blake@deploy2003 helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply [16:28:48] !log blake@deploy2003 helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply [16:29:06] !log blake@deploy2003 helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply [16:29:29] jouncebot: nowandnext [16:29:29] For the next 0 hour(s) and 30 minute(s): Puppet request window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260714T1600) [16:29:29] In 0 hour(s) and 30 minute(s): MediaWiki infrastructure (UTC late) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260714T1700) [16:29:58] since the puppet window is empty; I am going to use this window to reboot CI server and maybe Gerrit too [16:31:24] !log cmooney@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: repooling after rack b5 maintenance [16:31:28] !log cmooney@cumin1003 START - Cookbook sre.mysql.pool pool db2188: repooling after rack b5 maintenance [16:32:22] !log contint1003 - main CI server - rebooting [16:32:23] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [16:33:21] !log dzahn@cumin2002 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on contint1003.wikimedia.org with reason: reboot [16:35:57] (03PS2) 10Btullis: cloudnative-pg-cluster: omit empty backup encryption from Cluster manifests [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310602 (https://phabricator.wikimedia.org/T424493) [16:36:00] (03PS2) 10Btullis: cloudnative-pg-cluster: add SNI-based TLS passthrough ingress [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310603 (https://phabricator.wikimedia.org/T432104) [16:36:10] (03CR) 10Dreamy Jazz: "Seems fine once security review is complete" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1307751 (https://phabricator.wikimedia.org/T431023) (owner: 10Kosta Harlan) [16:42:25] RESOLVED: SystemdUnitFailed: dump_proxy_ranges.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [16:42:45] FIRING: [5x] WidespreadPuppetFailure: Puppet has failed in drmrs - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [16:43:16] RECOVERY - HAProxy HTTPS wikipedia25.org ECDSA on cp4039 is OK: SSL OK - Certificate wikipedia25.org contains all required SANs:Certificate wikipedia25.org (ECDSA) valid until 2026-10-04 05:54:30 +0000 (expires in 81 days) https://wikitech.wikimedia.org/wiki/HTTPS [16:43:16] RECOVERY - HAProxy HTTPS wikipedia.org ECDSA on cp4039 is OK: SSL OK - Certificate *.wikipedia.org contains all required SANs:Certificate *.wikipedia.org (ECDSA) valid until 2026-09-04 20:06:43 +0000 (expires in 52 days) https://wikitech.wikimedia.org/wiki/HTTPS [16:43:18] RECOVERY - haproxy process on cp4039 is OK: PROCS OK: 2 processes with command name haproxy https://wikitech.wikimedia.org/wiki/HAProxy [16:43:40] RECOVERY - HAProxy HTTPS wikiworkshop.org ECDSA on cp4039 is OK: SSL OK - Certificate wikiworkshop.org contains all required SANs:Certificate wikiworkshop.org (ECDSA) valid until 2026-09-10 03:24:50 +0000 (expires in 57 days) https://wikitech.wikimedia.org/wiki/HTTPS [16:44:21] !log sudo cumin -b31 "A:cp" "run-puppet-agent" [16:44:22] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [16:47:24] !log sukhe@puppetserver1001 conftool action : set/pooled=yes; selector: name=cp4039.ulsfo.wmnet [16:47:45] FIRING: [5x] WidespreadPuppetFailure: Puppet has failed in drmrs - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [16:50:04] !log pool cp2046 [16:50:05] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [16:50:08] RECOVERY - HAProxy HTTPS wikipedia25.org ECDSA on cp7001 is OK: SSL OK - Certificate wikipedia25.org contains all required SANs:Certificate wikipedia25.org (ECDSA) valid until 2026-10-04 05:54:30 +0000 (expires in 81 days) https://wikitech.wikimedia.org/wiki/HTTPS [16:50:08] RECOVERY - HAProxy HTTPS wikiworkshop.org ECDSA on cp7001 is OK: SSL OK - Certificate wikiworkshop.org contains all required SANs:Certificate wikiworkshop.org (ECDSA) valid until 2026-09-10 03:24:50 +0000 (expires in 57 days) https://wikitech.wikimedia.org/wiki/HTTPS [16:50:08] RECOVERY - HAProxy HTTPS wikipedia.org ECDSA on cp7001 is OK: SSL OK - Certificate *.wikipedia.org contains all required SANs:Certificate *.wikipedia.org (ECDSA) valid until 2026-08-16 10:23:55 +0000 (expires in 32 days) https://wikitech.wikimedia.org/wiki/HTTPS [16:50:20] RECOVERY - haproxy process on cp7001 is OK: PROCS OK: 2 processes with command name haproxy https://wikitech.wikimedia.org/wiki/HAProxy [16:50:26] coool [16:52:45] RESOLVED: [5x] WidespreadPuppetFailure: Puppet has failed in drmrs - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [16:55:01] (03PS2) 10FNegri: tox.ini: allow running a single test [cookbooks] - 10https://gerrit.wikimedia.org/r/1307091 [16:55:01] (03PS20) 10FNegri: sre.mysql.multiinstance_reboot: new cookbook [cookbooks] - 10https://gerrit.wikimedia.org/r/1290806 (https://phabricator.wikimedia.org/T420203) [16:57:19] !log reprepro include php8.3_8.3.32-1+wmf12u2 into component/php83 for bookworm-wikimedia [16:57:20] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [16:57:28] (03PS21) 10FNegri: sre.mysql.multiinstance_reboot: new cookbook [cookbooks] - 10https://gerrit.wikimedia.org/r/1290806 (https://phabricator.wikimedia.org/T420203) [16:59:06] (03CR) 10Dzahn: "Can this be tested somehow before merging it which affects a lot of machines? Maybe by disabling puppet, adding this line manually and the" [puppet] - 10https://gerrit.wikimedia.org/r/1310589 (https://phabricator.wikimedia.org/T432108) (owner: 10Kamila Součková) [16:59:45] (03CR) 10Dzahn: "could also be done on a bunch of other machines that sync between each other, for example releases*" [puppet] - 10https://gerrit.wikimedia.org/r/1310589 (https://phabricator.wikimedia.org/T432108) (owner: 10Kamila Součková) [17:00:04] swfrench-wmf: May I have your attention please! MediaWiki infrastructure (UTC late). (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260714T1700) [17:00:12] o/ [17:00:23] I'll be starting work on the infra window shortly [17:00:35] (03PS22) 10FNegri: sre.mysql.multiinstance_reboot: new cookbook [cookbooks] - 10https://gerrit.wikimedia.org/r/1290806 (https://phabricator.wikimedia.org/T420203) [17:01:26] jouncebot: now [17:01:26] For the next 0 hour(s) and 58 minute(s): MediaWiki infrastructure (UTC late) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260714T1700) [17:01:35] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp7001.magru.wmnet [17:02:54] (03CR) 10Scott French: [V:03+2] "Build locally:" [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1310598 (https://phabricator.wikimedia.org/T431099) (owner: 10Scott French) [17:03:04] (03CR) 10Scott French: [V:03+2] "Thanks for the review!" [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1310598 (https://phabricator.wikimedia.org/T431099) (owner: 10Scott French) [17:03:05] (03CR) 10Scott French: [V:03+2 C:03+2] php8.3: Rebuild to pick up new PHP packages (8.3.32) [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1310598 (https://phabricator.wikimedia.org/T431099) (owner: 10Scott French) [17:04:40] (03CR) 10Dzahn: "well, you fixed the tests and CI upvotes you. that seems good. I can't decide though if you really want to merge this all at once or break" [puppet] - 10https://gerrit.wikimedia.org/r/1305986 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [17:12:55] !log swfrench@deploy2003 Started scap sync-world: Deployment to pick up new production image [17:13:16] (03PS1) 10Ladsgroup: swift: Migrate storage of results to asyncio [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1310618 (https://phabricator.wikimedia.org/T431767) [17:16:39] (03PS1) 10AOkoth: site: revert phab1005 to insetup [puppet] - 10https://gerrit.wikimedia.org/r/1310621 (https://phabricator.wikimedia.org/T377889) [17:16:55] !log cmooney@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: repooling after rack b5 maintenance [17:16:59] !log cmooney@cumin1003 START - Cookbook sre.mysql.pool pool es2035: repooling after rack b5 maintenance [17:17:02] !log cmooney@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: repooling after rack b5 maintenance [17:18:18] RECOVERY - HAProxy HTTPS measure-eqiad.wikimedia.org ECDSA on cp7009 is OK: SSL OK - Certificate measure-eqiad.wikimedia.org contains all required SANs:Certificate measure-eqiad.wikimedia.org (ECDSA) valid until 2026-10-03 14:53:29 +0000 (expires in 80 days) https://wikitech.wikimedia.org/wiki/HTTPS [17:18:18] RECOVERY - HAProxy HTTPS upload.wikimedia.org ECDSA on cp7009 is OK: SSL OK - Certificate upload.wikimedia.org contains all required SANs:Certificate upload.wikimedia.org (ECDSA) valid until 2026-09-10 12:23:38 +0000 (expires in 57 days) https://wikitech.wikimedia.org/wiki/HTTPS [17:18:20] RECOVERY - haproxy process on cp7009 is OK: PROCS OK: 2 processes with command name haproxy https://wikitech.wikimedia.org/wiki/HAProxy [17:19:26] PROBLEM - Check if ntpsec.service has been restarted after /etc/ntpsec/ntp.conf was changed on dns7002 is CRITICAL: CRITICAL: Service ntpsec.service has not been restarted after /etc/ntpsec/ntp.conf was changed (gt 2h). https://wikitech.wikimedia.org/wiki/NTP%23Monitoring [17:21:53] (03CR) 10BCornwall: [C:04-1] P:cache::varnish::frontend configure varnish for thumbnail (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1306896 (https://phabricator.wikimedia.org/T427465) (owner: 10Slyngshede) [17:24:05] * swfrench-wmf taps imaginary watch and glares at docker push [17:27:35] (03PS1) 10Ladsgroup: images: Move loading of memcached key to asyncio [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1310623 (https://phabricator.wikimedia.org/T431767) [17:29:50] !log swfrench@deploy2003 swfrench: Deployment to pick up new production image synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [17:29:50] (03CR) 10Ssingh: [C:03+2] admin: add SSH key for sukhe [puppet] - 10https://gerrit.wikimedia.org/r/1310133 (owner: 10Ssingh) [17:32:17] !log swfrench@deploy2003 swfrench: Continuing with deployment [17:32:37] (03PS4) 10Ladsgroup: P:cache::varnish::frontend configure varnish for thumbnail [puppet] - 10https://gerrit.wikimedia.org/r/1306896 (https://phabricator.wikimedia.org/T427465) (owner: 10Slyngshede) [17:32:51] (03CR) 10CDobbins: P:cache::varnish::frontend configure varnish for thumbnail (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1306896 (https://phabricator.wikimedia.org/T427465) (owner: 10Slyngshede) [17:33:23] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp7009.magru.wmnet [17:33:37] RESOLVED: CoreRouterInterfaceDown: Core router interface down - cr3-ulsfo:et-0/0/0 (Transport: Hurricane Electric (dc4841.sfo1) {#changeme_ulsfo_he_cct}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr3-ulsfo:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [17:33:48] (03CR) 10Ladsgroup: P:cache::varnish::frontend configure varnish for thumbnail (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1306896 (https://phabricator.wikimedia.org/T427465) (owner: 10Slyngshede) [17:42:43] (03CR) 10BCornwall: varnish: add tests for text thumbnail config (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1306911 (https://phabricator.wikimedia.org/T427465) (owner: 10CDobbins) [17:42:59] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp7002.magru.wmnet [17:44:14] !log swfrench@deploy2003 Finished scap sync-world: Deployment to pick up new production image (duration: 31m 44s) [17:49:12] alright, I believe I'm done with today's infra window [17:50:38] FIRING: CertAlmostExpired: gNMI TLS certificate for lsw1-f6-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [17:52:41] (03CR) 10Federico Ceratto: [C:03+2] sre.mysql: split pool/depool [cookbooks] - 10https://gerrit.wikimedia.org/r/1295480 (https://phabricator.wikimedia.org/T422361) (owner: 10Federico Ceratto) [17:54:25] FIRING: [12x] BFDdown: BFD session down between cr1-codfw and 2620:0:860:104:10:192:48:14 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [17:57:24] (03Merged) 10jenkins-bot: sre.mysql: split pool/depool [cookbooks] - 10https://gerrit.wikimedia.org/r/1295480 (https://phabricator.wikimedia.org/T422361) (owner: 10Federico Ceratto) [17:57:46] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, July 14 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item" [core] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1310625 (https://phabricator.wikimedia.org/T432042) (owner: 10SBassett) [18:00:05] jeena and hashar: Deploy window MediaWiki train - Utc-7+Utc-0 Version (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260714T1800) [18:00:33] o/ [18:04:55] (03PS1) 10TrainBranchBot: group0 to 1.47.0-wmf.11 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310627 (https://phabricator.wikimedia.org/T430830) [18:04:58] (03CR) 10TrainBranchBot: [C:03+2] "Initiated by jhuneidi@deploy2003" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310627 (https://phabricator.wikimedia.org/T430830) (owner: 10TrainBranchBot) [18:05:52] (03Merged) 10jenkins-bot: group0 to 1.47.0-wmf.11 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310627 (https://phabricator.wikimedia.org/T430830) (owner: 10TrainBranchBot) [18:06:45] FIRING: [2x] WidespreadPuppetFailure: Puppet has failed in drmrs - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [18:07:07] again? [18:07:37] Could not evaluate: Could not retrieve file metadata for puppet:///volatile/ip_reputation_vendors/proxy.mmdb: Error 500 on SERVER: Server Error: Permission denied - /srv/puppet_fileserver/volatile/ip_reputation_vendors/proxy.mmdb [18:07:54] ah come on [18:08:25] permission denied? [18:08:36] (03CR) 10Scott French: [C:03+1] "This looks good to me!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309641 (https://phabricator.wikimedia.org/T427668) (owner: 10Blake) [18:11:45] FIRING: [4x] WidespreadPuppetFailure: Puppet has failed in drmrs - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [18:12:52] (03CR) 10Kareid: [C:03+1] Test Kitchen UI: Deploying v1.4.8 release to production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310139 (https://phabricator.wikimedia.org/T428984) (owner: 10Santiago Faci) [18:17:37] Deployment is still underway but I see a spike of errors and I suspect a rollback will be needed [18:18:29] !log jhuneidi@deploy2003 rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.11 refs T430830 [18:18:33] T430830: 1.47.0-wmf.11 deployment blockers - https://phabricator.wikimedia.org/T430830 [18:19:27] (03PS1) 10Bking: deployment-server: install python commandline parser [puppet] - 10https://gerrit.wikimedia.org/r/1310630 (https://phabricator.wikimedia.org/T431506) [18:19:46] the errors look like they are from wmf 10 though and not 11 [18:19:48] (03PS2) 10Bking: deployment-server: install python commandline parser [puppet] - 10https://gerrit.wikimedia.org/r/1310630 (https://phabricator.wikimedia.org/T431506) [18:20:31] (03PS1) 10Ssingh: ip_reputation: fix permissions for fetch_spur_proxy [puppet] - 10https://gerrit.wikimedia.org/r/1310631 [18:20:43] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1310630 (https://phabricator.wikimedia.org/T431506) (owner: 10Bking) [18:20:52] (03PS2) 10Ssingh: ip_reputation: fix permissions for fetch_spur_proxy [puppet] - 10https://gerrit.wikimedia.org/r/1310631 [18:21:39] (03CR) 10Ssingh: [V:03+1] "PCC SUCCESS (NOOP 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9000/console" [puppet] - 10https://gerrit.wikimedia.org/r/1310631 (owner: 10Ssingh) [18:25:40] (03CR) 10JHathaway: [C:03+1] ip_reputation: fix permissions for fetch_spur_proxy [puppet] - 10https://gerrit.wikimedia.org/r/1310631 (owner: 10Ssingh) [18:26:45] FIRING: [5x] WidespreadPuppetFailure: Puppet has failed in drmrs - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [18:27:07] (03CR) 10Ssingh: [V:03+1 C:03+2] ip_reputation: fix permissions for fetch_spur_proxy [puppet] - 10https://gerrit.wikimedia.org/r/1310631 (owner: 10Ssingh) [18:30:43] FIRING: [3x] JobUnavailable: Reduced availability for job atlas_exporter in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [18:30:48] It seems fine now so I will leave it [18:30:59] jeena: interesting, is it all mediawikiwiki? this looks like maybe a serialization compatibility issue that's resolved when the release is fully on .11? [18:31:45] as in, the warnings are only emitted while mediawikiwiki was running on a mix of .10 and .11 [18:33:06] looking closer, they were actually all warnings (deprecation and foreach() argument must be of type array|object, null given), just weirdly a spike of thousands during deployment [18:33:20] we won't be fully on .11 until thurs [18:33:25] https://gerrit.wikimedia.org/r/c/mediawiki/core/+/1309286 [18:34:03] (03CR) 10BCornwall: "I think we should be porting over the existing tests that we already have for thumbs in the upload cluster - they're pretty comprehensive " [puppet] - 10https://gerrit.wikimedia.org/r/1306911 (https://phabricator.wikimedia.org/T427465) (owner: 10CDobbins) [18:34:09] I'm wondering if the namespacing change is doing something odd for class serialization between .10 and .11 [18:35:05] the `PHP Deprecated: Creation of dynamic property ...` is a pretty common side-effect of de-serialization getting something it doesn't expect [18:35:11] hmm maybe [18:35:44] (03CR) 10BCornwall: "Marking unresolved. Stupid gerrit." [puppet] - 10https://gerrit.wikimedia.org/r/1306911 (https://phabricator.wikimedia.org/T427465) (owner: 10CDobbins) [18:37:06] in any case, might be worth having MediaWiki Eng take a look at whether https://gerrit.wikimedia.org/r/c/mediawiki/core/+/1309286 is doing something odd here, even if the incompatibility is only transient while the deployment is in flight [18:37:42] Yeah, that sounds good. Thanks for your insight! [18:37:57] no problem! hope it's not wildly off :) [18:38:08] !log rotating phabricator-gerrit bot token (its-phabricator) [18:38:09] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [18:41:55] jeena: ah, yeah I think I see what's maybe going on here ... the patch contains compatibility classes that ensure cached .10-serialized classes are forward-compatible with being deserialized on .11. however, there's nothing to ensure the other direction: that .10 will know what to do when it sees (and tries to deserialize) a .11-serialized class that uses the new namespace and has been written to cache. [18:42:00] (03CR) 10BCornwall: [C:04-1] "In fact, I think the tests should be in the parent CR, not this one. Keep it atomic. I'll get to work integrating the existing tests." [puppet] - 10https://gerrit.wikimedia.org/r/1306911 (https://phabricator.wikimedia.org/T427465) (owner: 10CDobbins) [18:42:02] 10ops-drmrs, 06DC-Ops, 06Traffic: hw troubleshooting: Memory failure for cp6008.drmrs.wmnet - https://phabricator.wikimedia.org/T431651#12121189 (10ssingh) Hi DC-Ops. Just following up on this in case it was missed, thanks! This is not urgent but will be good to know how long it may take. [18:42:24] I see [18:43:10] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp7010.magru.wmnet [18:43:39] Well I guess this will be resolved on thurs then 😅 [18:44:31] (03CR) 10BCornwall: [C:04-1] "Yes, we should port over the existing upload tests to cover text." [puppet] - 10https://gerrit.wikimedia.org/r/1306896 (https://phabricator.wikimedia.org/T427465) (owner: 10Slyngshede) [18:49:00] heh, indeed. the one thing I'm not clear on is whether this actually "breaks" anything while that incompatibility is live - i.e., there are no "hard" errors emitted, but I also don't know to what extent the affected code works correctly. [18:51:06] (03PS1) 10Jforrester: wikifunctions: Define the virtual-wikifunctions-usage table [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310634 [18:52:26] Although the warnings only showed up during the deployment so that's why I assumed it was fine [18:56:10] (03PS2) 10Jforrester: wikifunctions: Define the virtual-wikifunctions-usage table [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310634 (https://phabricator.wikimedia.org/T390557) [18:57:43] (03PS5) 10Cathal Mooney: Nokia SR-Linux: add export policy for unicast IBGP clusters [homer/public] - 10https://gerrit.wikimedia.org/r/1309643 (https://phabricator.wikimedia.org/T423430) [18:58:57] jeena: Argh, yeah, I guess ideally we'd have first shipped a shim in wmf.10 that was forward-compatible. Meh. [18:59:29] It *should* be OK as long as .11-serialised classes aren't pushed into .10 code somehow? [19:00:07] jeena: Also, OK for me to deploy 1310634 to fix that fatal? (Sorry! Thought I'd already configured the new DB virtual address.) [19:00:50] James_F: Thanks for taking a look! And yes I was just about to ask you if I was supposed to review that change 😆. It's fine to deploy now [19:01:00] Ack. [19:01:02] (03CR) 10Cathal Mooney: Nokia SR-Linux: add export policy for unicast IBGP clusters (031 comment) [homer/public] - 10https://gerrit.wikimedia.org/r/1309643 (https://phabricator.wikimedia.org/T423430) (owner: 10Cathal Mooney) [19:01:21] (03CR) 10TrainBranchBot: [C:03+2] "Approved by jforrester@deploy2003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310634 (https://phabricator.wikimedia.org/T390557) (owner: 10Jforrester) [19:02:18] (03Merged) 10jenkins-bot: wikifunctions: Define the virtual-wikifunctions-usage table [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310634 (https://phabricator.wikimedia.org/T390557) (owner: 10Jforrester) [19:02:44] !log jforrester@deploy2003 Started scap sync-world: Backport for [[gerrit:1310634|wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] [19:02:47] T390557: Display the local and cross-wiki pages on which a Function is used, so that Wikifunctions users can see the impact of their changes - https://phabricator.wikimedia.org/T390557 [19:04:54] !log jforrester@deploy2003 jforrester: Backport for [[gerrit:1310634|wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [19:07:00] !log jforrester@deploy2003 jforrester: Continuing with deployment [19:09:09] (03CR) 10Ryan Kemper: [C:04-1] "let's stick this in profile::kubernetes::deployment_server next to the existing ensure_packages(['istioctl', 'kubetail']), since global_co" [puppet] - 10https://gerrit.wikimedia.org/r/1310630 (https://phabricator.wikimedia.org/T431506) (owner: 10Bking) [19:11:17] !log jforrester@deploy2003 Finished scap sync-world: Backport for [[gerrit:1310634|wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] (duration: 08m 33s) [19:11:21] T390557: Display the local and cross-wiki pages on which a Function is used, so that Wikifunctions users can see the impact of their changes - https://phabricator.wikimedia.org/T390557 [19:12:53] jeena: Should be fixed. Thanks! [19:13:01] Thanks James_F ! [19:15:18] (03PS1) 10Lerickson: Bump the proxy to version 0.4.0. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310635 (https://phabricator.wikimedia.org/T429407) [19:16:45] FIRING: [5x] WidespreadPuppetFailure: Puppet has failed in drmrs - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [19:17:35] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, July 14 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309865 (https://phabricator.wikimedia.org/T431094) (owner: 10Priyankar22) [19:19:36] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp7003.magru.wmnet [19:21:45] FIRING: [5x] WidespreadPuppetFailure: Puppet has failed in drmrs - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [19:21:56] two hosts failing which should resolve soon [19:21:59] nothing to worry [19:22:09] I guess it is "widespread" :P [19:24:46] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp7011.magru.wmnet [19:26:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.38% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [19:26:45] RESOLVED: [5x] WidespreadPuppetFailure: Puppet has failed in drmrs - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [19:28:59] (03Abandoned) 10CDobbins: varnish: add tests for text thumbnail config [puppet] - 10https://gerrit.wikimedia.org/r/1306911 (https://phabricator.wikimedia.org/T427465) (owner: 10CDobbins) [19:31:07] the high load seems to correlate with s4 (commons) activity increase [19:31:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.35% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [19:34:18] probably a well meant bot, but a bit on the high side [19:34:58] (03CR) 10Gmodena: [C:03+1] Bump the proxy to version 0.4.0. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310635 (https://phabricator.wikimedia.org/T429407) (owner: 10Lerickson) [19:35:35] nah, it wouldn't be a bot because it would be api, not mw web [19:40:33] (03CR) 10Scott French: [C:03+1] "One additional thing to mention:" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309641 (https://phabricator.wikimedia.org/T427668) (owner: 10Blake) [19:43:52] (03CR) 10Lerickson: [C:03+2] Bump the proxy to version 0.4.0. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310635 (https://phabricator.wikimedia.org/T429407) (owner: 10Lerickson) [19:46:03] (03Merged) 10jenkins-bot: Bump the proxy to version 0.4.0. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310635 (https://phabricator.wikimedia.org/T429407) (owner: 10Lerickson) [19:48:10] I think I figured out, organic traffic, mostly fr, es. Lots of referers from "Lamine Yamal", "Oyarzabal" :-) [19:51:35] (03PS2) 10AOkoth: site: revert phab1005 to insetup [puppet] - 10https://gerrit.wikimedia.org/r/1310621 (https://phabricator.wikimedia.org/T377889) [19:51:36] and being uncached it is probably logged in actions, so all good [19:54:42] (03CR) 10AOkoth: [C:03+2] site: revert phab1005 to insetup [puppet] - 10https://gerrit.wikimedia.org/r/1310621 (https://phabricator.wikimedia.org/T377889) (owner: 10AOkoth) [19:56:43] (03CR) 10Cathal Mooney: "So I picked a random peer from a small ASN on cr3-eqsin to make some checks against. I deactivated the peer, then added the same config a" [homer/public] - 10https://gerrit.wikimedia.org/r/1309707 (https://phabricator.wikimedia.org/T431849) (owner: 10Cathal Mooney) [19:57:01] !log aokoth@cumin1003 START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet [19:58:15] (03PS1) 10Bking: stat hosts: Set up S3 config for wikidata platform team. [puppet] - 10https://gerrit.wikimedia.org/r/1310643 (https://phabricator.wikimedia.org/T432167) [19:58:34] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1310643 (https://phabricator.wikimedia.org/T432167) (owner: 10Bking) [20:00:05] RoanKattouw, urbanecm, TheresNoTime, kindrobot, and cjming: Your horoscope predicts another UTC late backport window deploy. May Zuul be (nice) with you. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260714T2000). [20:00:05] sbassett and anzx: A patch you scheduled for UTC late backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [20:00:28] o/ [20:00:44] !log aokoth@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet [20:00:59] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp7004.magru.wmnet [20:01:50] anzx do you need a deployer? [20:01:56] yes [20:02:41] okay I can deploy for you [20:02:48] !log aokoth@cumin1003 START - Cookbook sre.hosts.reimage for host phab1005.eqiad.wmnet with OS trixie [20:03:03] (03CR) 10TrainBranchBot: [C:03+2] "Approved by jhuneidi@deploy2003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309865 (https://phabricator.wikimedia.org/T431094) (owner: 10Priyankar22) [20:03:36] o/ [20:03:43] (03PS2) 10Bking: stat hosts: Set up S3 config for wikidata platform team. [puppet] - 10https://gerrit.wikimedia.org/r/1310643 (https://phabricator.wikimedia.org/T432167) [20:04:01] I can go second, the config patch (not mine) should be quick and easy [20:04:01] (03Merged) 10jenkins-bot: thwiki: Change to Wikipedia 25 logo [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309865 (https://phabricator.wikimedia.org/T431094) (owner: 10Priyankar22) [20:04:17] !log jhuneidi@deploy2003 Started scap sync-world: Backport for [[gerrit:1309865|thwiki: Change to Wikipedia 25 logo (T431094)]] [20:04:20] T431094: thwiki: change to Wikipedia 25 logo - https://phabricator.wikimedia.org/T431094 [20:05:53] (03PS3) 10Bking: stat hosts: Set up S3 config for wikidata platform team. [puppet] - 10https://gerrit.wikimedia.org/r/1310643 (https://phabricator.wikimedia.org/T432167) [20:06:01] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1310643 (https://phabricator.wikimedia.org/T432167) (owner: 10Bking) [20:06:14] !log jhuneidi@deploy2003 jhuneidi, priyankar22: Backport for [[gerrit:1309865|thwiki: Change to Wikipedia 25 logo (T431094)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [20:06:26] checking [20:06:29] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp7012.magru.wmnet [20:07:27] jeena: looks good, ok to sync [20:07:37] thanks anzx [20:07:42] !log jhuneidi@deploy2003 jhuneidi, priyankar22: Continuing with deployment [20:11:59] !log jhuneidi@deploy2003 Finished scap sync-world: Backport for [[gerrit:1309865|thwiki: Change to Wikipedia 25 logo (T431094)]] (duration: 07m 42s) [20:12:02] T431094: thwiki: change to Wikipedia 25 logo - https://phabricator.wikimedia.org/T431094 [20:12:41] jeena: please above to purge images https://www.irccloud.com/pastebin/Lk3TTxC2/ [20:13:02] okay just a moment please [20:15:41] jeena: is it ok if I start my spiderpig deploy now? or did you want to wait until you run those maint scripts? [20:16:00] I think it's fine to deploy now sbassett [20:16:05] ok, thanks [20:16:18] (03CR) 10TrainBranchBot: [C:03+2] "Approved by sbassett@deploy2003 using scap backport" [core] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1310625 (https://phabricator.wikimedia.org/T432042) (owner: 10SBassett) [20:16:31] anzx: done [20:16:47] jeena: thanks for deploying [20:17:12] you're welcome! [20:17:23] jeena: https://wikitech.wikimedia.org/wiki/Maintenance_scripts#Input_on_stdin [20:19:40] (03PS4) 10Bking: stat hosts: Set up S3 config for wikidata platform team. [puppet] - 10https://gerrit.wikimedia.org/r/1310643 (https://phabricator.wikimedia.org/T432167) [20:20:55] !log aokoth@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage [20:21:28] (03Merged) 10jenkins-bot: Add logging for various re-authentication methods [core] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1310625 (https://phabricator.wikimedia.org/T432042) (owner: 10SBassett) [20:21:45] !log sbassett@deploy2003 Started scap sync-world: Backport for [[gerrit:1310625|Add logging for various re-authentication methods (T432042)]] [20:21:48] T432042: Add logging for various re-authentication methods (AuthPopup vs DataStash flow) - https://phabricator.wikimedia.org/T432042 [20:23:37] !log aokoth@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage [20:23:39] !log sbassett@deploy2003 sbassett: Backport for [[gerrit:1310625|Add logging for various re-authentication methods (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [20:24:14] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1310643 (https://phabricator.wikimedia.org/T432167) (owner: 10Bking) [20:24:15] !log sbassett@deploy2003 sbassett: Continuing with deployment [20:28:32] !log sbassett@deploy2003 Finished scap sync-world: Backport for [[gerrit:1310625|Add logging for various re-authentication methods (T432042)]] (duration: 06m 47s) [20:28:35] T432042: Add logging for various re-authentication methods (AuthPopup vs DataStash flow) - https://phabricator.wikimedia.org/T432042 [20:29:14] ok, I should be good [20:34:16] (03PS1) 10Bking: Add secret for wikidata_platform thanos-swift [labs/private] - 10https://gerrit.wikimedia.org/r/1310652 (https://phabricator.wikimedia.org/T432167) [20:36:22] (03PS2) 10Bking: Add secret for wikidata_platform thanos-swift [labs/private] - 10https://gerrit.wikimedia.org/r/1310652 (https://phabricator.wikimedia.org/T432167) [20:36:50] (03CR) 10Bking: [C:03+2] Add secret for wikidata_platform thanos-swift [labs/private] - 10https://gerrit.wikimedia.org/r/1310652 (https://phabricator.wikimedia.org/T432167) (owner: 10Bking) [20:37:00] (03CR) 10Bking: [V:03+2 C:03+2] Add secret for wikidata_platform thanos-swift [labs/private] - 10https://gerrit.wikimedia.org/r/1310652 (https://phabricator.wikimedia.org/T432167) (owner: 10Bking) [20:37:17] (03PS1) 10Aude: Enable responsive Vector toolbar on beta cluster and testwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310653 (https://phabricator.wikimedia.org/T429518) [20:38:25] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, July 15 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployc" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310653 (https://phabricator.wikimedia.org/T429518) (owner: 10Aude) [20:39:07] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1310643 (https://phabricator.wikimedia.org/T432167) (owner: 10Bking) [20:40:02] (03PS5) 10Bking: stat hosts: Set up S3 config for wikidata platform team. [puppet] - 10https://gerrit.wikimedia.org/r/1310643 (https://phabricator.wikimedia.org/T432167) [20:40:43] FIRING: [3x] JobUnavailable: Reduced availability for job atlas_exporter in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [20:41:45] !log aokoth@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host phab1005.eqiad.wmnet with OS trixie [20:42:23] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp7005.magru.wmnet [20:42:44] (03CR) 10Aude: [C:04-2] Enable responsive Vector toolbar on beta cluster and testwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310653 (https://phabricator.wikimedia.org/T429518) (owner: 10Aude) [20:42:46] (03PS1) 10Aaron Schulz: Remove redundant fetchGoogleCloudVisionAnnotations entry from _jobqueue.yaml [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310654 [20:43:07] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1310643 (https://phabricator.wikimedia.org/T432167) (owner: 10Bking) [20:43:42] (03Abandoned) 10Aude: Enable responsive Vector toolbar on beta cluster and testwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1310653 (https://phabricator.wikimedia.org/T429518) (owner: 10Aude) [20:45:43] FIRING: [3x] JobUnavailable: Reduced availability for job atlas_exporter in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [20:48:12] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp7013.magru.wmnet [20:50:32] (03CR) 10CDobbins: "Thanks for the +1, @sbassett@wikimedia.org. Can this be merged tomorrow in its current state?" [puppet] - 10https://gerrit.wikimedia.org/r/1309263 (owner: 10CDobbins) [20:51:36] (03CR) 10SBassett: [C:03+1] varnish: update CSP report-only header [puppet] - 10https://gerrit.wikimedia.org/r/1308772 (owner: 10CDobbins) [20:53:32] (03CR) 10SBassett: [C:03+1] "Yes, this LGTM now. Thanks." [puppet] - 10https://gerrit.wikimedia.org/r/1309263 (owner: 10CDobbins) [20:55:31] (03PS1) 10JHathaway: WIP: Puppet 8: Replace legacy facts in misc kafka [puppet] - 10https://gerrit.wikimedia.org/r/1310655 (https://phabricator.wikimedia.org/T372666) [20:55:45] !log arlolra@deploy2003 helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply [20:55:48] (03PS1) 10Dzahn: site: add zuul1004/2004 - test commit to test gerrit bot [puppet] - 10https://gerrit.wikimedia.org/r/1310656 (https://phabricator.wikimedia.org/T432168) [20:56:17] !log arlolra@deploy2003 helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply [20:56:18] !log arlolra@deploy2003 helmfile [codfw] START helmfile.d/services/mw-parsoid: apply [20:56:37] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1310655 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [20:56:45] !log arlolra@deploy2003 helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply [20:58:24] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in role kafka [puppet] - 10https://gerrit.wikimedia.org/r/1308759 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [20:59:19] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in role wdqs [puppet] - 10https://gerrit.wikimedia.org/r/1308757 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:00:05] Deploy window Readers deployment window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260714T2100) [21:00:10] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in role druid [puppet] - 10https://gerrit.wikimedia.org/r/1308756 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:00:50] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in role analytics_test_cluster [puppet] - 10https://gerrit.wikimedia.org/r/1308755 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:01:22] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in profile query_service [puppet] - 10https://gerrit.wikimedia.org/r/1308752 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:03:29] (03PS2) 10Dzahn: site: add zuul1004/2004 - test commit to test gerrit bot [puppet] - 10https://gerrit.wikimedia.org/r/1310656 [21:03:36] (03PS3) 10Dzahn: site: add zuul1004/2004 - test commit to test gerrit bot [puppet] - 10https://gerrit.wikimedia.org/r/1310656 (https://phabricator.wikimedia.org/T432168) [21:08:06] jouncebot: now [21:08:06] For the next 0 hour(s) and 51 minute(s): Readers deployment window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260714T2100) [21:08:38] is a deployment happening? [21:09:56] jhathaway: still in the middle of merging stuff? [21:10:16] I'm done for the moment mutante [21:10:32] thanks, I am going to do a gerrit reboot [21:10:38] to fix 2 separate things [21:10:44] sounds good, thanks for the heads up [21:11:16] !log dzahn@cumin2002 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on gerrit2003.wikimedia.org with reason: reboot [21:11:39] !log gerrit2003 (gerrit.wikimedia.org) - reboot for maintenance [21:11:40] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [21:12:47] 06SRE, 06Infrastructure-Foundations: gerrit.w.o is not included in https://config-master.wikimedia.org/known_hosts - https://phabricator.wikimedia.org/T340947#12121648 (10Dzahn) a:03Dzahn [21:13:11] !log dzahn@cumin2002 DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:15:00 on gerrit.wikimedia.org with reason: reboot [21:15:46] FIRING: [4x] GerritHAProxyBackendUnavailable: Gerrit backend is unavilable for tcp-proxy (HAProxy) gerrit_ssh - https://wikitech.wikimedia.org/wiki/Gerrit/Operations#GerritHAProxyBackendUnavailable - grafana.wikimedia.org/d/459365f6-df37-48d6-8142-82b22c1875e7/gerrit-tcp-proxy?viewPanel=panel-15 - https://alerts.wikimedia.org/?q=alertname%3DGerritHAProxyBackendUnavailable [21:15:57] FIRING: [2x] GerritHAProxyServiceUnavailable: Gerrit tcp-proxy (HAProxy) service gerrit_ssh is DOWN in codfw - https://wikitech.wikimedia.org/wiki/Gerrit/Operations#GerritHAProxyServiceUnavailable - grafana.wikimedia.org/d/459365f6-df37-48d6-8142-82b22c1875e7/gerrit-tcp-proxy?viewPanel=panel-15 - https://alerts.wikimedia.org/?q=alertname%3DGerritHAProxyServiceUnavailable [21:16:20] (Any estimate on when gerrit might be back after it's reboot?) [21:16:48] 5 to 10 minutes [21:16:57] well.. 5.. since I see it at login now [21:17:08] Thanks, off to touch grass for 5 mins :D [21:17:12] Dreamy_Jazz: 0 seconds :P [21:17:21] lol [21:17:50] Hello! Is Gerritt down atm? [21:17:57] not anymore [21:18:01] ^ [21:18:02] it was until seconds ago [21:18:27] lolll nice [21:18:30] thank you [21:18:39] :) [21:19:13] it was a server reboot, not an attack [21:19:37] 😮‍💨 ok cool! [21:20:46] RESOLVED: [4x] GerritHAProxyBackendUnavailable: Gerrit backend is unavilable for tcp-proxy (HAProxy) gerrit_ssh - https://wikitech.wikimedia.org/wiki/Gerrit/Operations#GerritHAProxyBackendUnavailable - grafana.wikimedia.org/d/459365f6-df37-48d6-8142-82b22c1875e7/gerrit-tcp-proxy?viewPanel=panel-15 - https://alerts.wikimedia.org/?q=alertname%3DGerritHAProxyBackendUnavailable [21:20:57] (03PS4) 10Dzahn: site: add zuul1004/2004 - test commit to test gerrit bot [puppet] - 10https://gerrit.wikimedia.org/r/1310656 [21:20:57] RESOLVED: [2x] GerritHAProxyServiceUnavailable: Gerrit tcp-proxy (HAProxy) service gerrit_ssh is DOWN in codfw - https://wikitech.wikimedia.org/wiki/Gerrit/Operations#GerritHAProxyServiceUnavailable - grafana.wikimedia.org/d/459365f6-df37-48d6-8142-82b22c1875e7/gerrit-tcp-proxy?viewPanel=panel-15 - https://alerts.wikimedia.org/?q=alertname%3DGerritHAProxyServiceUnavailable [21:21:01] (03PS5) 10Dzahn: site: add zuul1004/2004 - test commit to test gerrit bot [puppet] - 10https://gerrit.wikimedia.org/r/1310656 (https://phabricator.wikimedia.org/T432168) [21:21:15] noiice [21:21:32] sorry for the inconvenince. we try to limit those but sometimes just have to. [21:22:25] NW thanks a lot! [21:23:35] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp7006.magru.wmnet [21:29:02] 10ops-eqiad, 06SRE, 06collaboration-services, 06DC-Ops: Repurpose ganeti102[3456] for Zuul migration - https://phabricator.wikimedia.org/T427353#12121773 (10Dzahn) Hi @VRiley-WMF let me know if there are still questions around this. [21:29:12] (03PS6) 10Dzahn: site: add zuul1004/2004 [puppet] - 10https://gerrit.wikimedia.org/r/1310656 (https://phabricator.wikimedia.org/T432168) [21:29:40] (03PS7) 10Dzahn: site: add zuul1004/2004 [puppet] - 10https://gerrit.wikimedia.org/r/1310656 (https://phabricator.wikimedia.org/T427353) [21:29:48] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp7014.magru.wmnet [21:35:51] (03PS2) 10JHathaway: Puppet 8: Replace legacy facts in role analytics_cluster [puppet] - 10https://gerrit.wikimedia.org/r/1308754 (https://phabricator.wikimedia.org/T372666) [21:36:04] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308754 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:36:16] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in profile presto [puppet] - 10https://gerrit.wikimedia.org/r/1308751 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:36:35] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in profile opensearch [puppet] - 10https://gerrit.wikimedia.org/r/1308750 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:37:29] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in profile kafkatee [puppet] - 10https://gerrit.wikimedia.org/r/1308749 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:37:51] (03PS1) 10Dzahn: site: limit regex for zuul physical machines to 4-7 range [puppet] - 10https://gerrit.wikimedia.org/r/1310667 (https://phabricator.wikimedia.org/T427353) [21:37:52] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in profile kafka [puppet] - 10https://gerrit.wikimedia.org/r/1308748 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:38:30] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in profile hive [puppet] - 10https://gerrit.wikimedia.org/r/1308747 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:38:40] (03CR) 10Dzahn: [C:03+2] site: limit regex for zuul physical machines to 4-7 range [puppet] - 10https://gerrit.wikimedia.org/r/1310667 (https://phabricator.wikimedia.org/T427353) (owner: 10Dzahn) [21:39:08] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in profile elasticsearch [puppet] - 10https://gerrit.wikimedia.org/r/1308732 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:39:19] m-m-m-m-multi kill [21:39:33] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in profile druid [puppet] - 10https://gerrit.wikimedia.org/r/1308730 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:39:54] jhathaway: shit it all please :) [21:39:59] ship! :) lol [21:40:10] lol [21:40:47] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in profile bigtop [puppet] - 10https://gerrit.wikimedia.org/r/1308727 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:41:21] mutante: can I merge your regex change to site? [21:42:02] jhathaway: that's what I meant by the weird phrasing :) [21:42:05] (03Abandoned) 10JHathaway: WIP: Puppet 8: Replace legacy facts in misc kafka [puppet] - 10https://gerrit.wikimedia.org/r/1310655 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:42:09] yes [21:42:11] ah! [21:43:15] 10ops-eqiad, 06SRE, 06collaboration-services, 06DC-Ops, 13Patch-For-Review: Repurpose ganeti102[3456] for Zuul migration - https://phabricator.wikimedia.org/T427353#12121853 (10Dzahn) I made sure also in site.pp there is only 4 through 7 - no more 8 and 9 appearing there. [21:43:49] jhathaway: every time I get the "multiple" warning at puppet-merge in my mind it says https://www.urbandictionary.com/define.php?term=M-m-m-monster+Kill [21:44:12] well, it was "multi" first and then "monster" next [21:44:23] ha, memories indeed [21:48:16] (03CR) 10Scott French: [C:03+1] "Thanks, Effie!" [puppet] - 10https://gerrit.wikimedia.org/r/1310103 (https://phabricator.wikimedia.org/T428909) (owner: 10Effie Mouzeli) [21:48:26] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in role analytics_cluster [puppet] - 10https://gerrit.wikimedia.org/r/1308754 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:49:39] (03CR) 10JHathaway: "@taavi@wikimedia.org any thoughts on this approach?" [puppet] - 10https://gerrit.wikimedia.org/r/1308177 (https://phabricator.wikimedia.org/T427799) (owner: 10JHathaway) [21:50:01] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in profile analytics [puppet] - 10https://gerrit.wikimedia.org/r/1308725 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:50:31] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in module zookeeper [puppet] - 10https://gerrit.wikimedia.org/r/1308724 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:50:38] FIRING: CertAlmostExpired: gNMI TLS certificate for lsw1-f6-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [21:50:48] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in module statistics [puppet] - 10https://gerrit.wikimedia.org/r/1308722 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:51:20] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in module opensearch [puppet] - 10https://gerrit.wikimedia.org/r/1308721 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:51:42] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in module kafkatee [puppet] - 10https://gerrit.wikimedia.org/r/1308720 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:52:07] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in module elasticsearch [puppet] - 10https://gerrit.wikimedia.org/r/1308719 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:52:33] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in module confluent [puppet] - 10https://gerrit.wikimedia.org/r/1308718 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:53:01] 06SRE, 06Infrastructure-Foundations: gerrit.w.o is not included in https://config-master.wikimedia.org/known_hosts - https://phabricator.wikimedia.org/T340947#12121881 (10Dzahn) In the context of the linked patch above we got back to this and it's actually resolved. - gerrit was already in https://config-mast... [21:53:09] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in module bigtop [puppet] - 10https://gerrit.wikimedia.org/r/1308717 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:54:03] 06SRE, 06collaboration-services, 10Gerrit, 06Infrastructure-Foundations: gerrit.w.o is not included in https://config-master.wikimedia.org/known_hosts - https://phabricator.wikimedia.org/T340947#12121883 (10Dzahn) [21:54:12] 06SRE, 06collaboration-services, 10Gerrit, 06Infrastructure-Foundations: gerrit.w.o is not included in https://config-master.wikimedia.org/known_hosts - https://phabricator.wikimedia.org/T340947#12121884 (10Dzahn) 05Open→03Resolved [21:54:25] FIRING: [12x] BFDdown: BFD session down between cr1-codfw and 2620:0:860:104:10:192:48:14 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [21:57:35] (03PS6) 10JHathaway: kafka_config: migrate functions to the 4.x API [puppet] - 10https://gerrit.wikimedia.org/r/1308760 (https://phabricator.wikimedia.org/T372666) [22:01:56] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308760 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [22:03:23] (03CR) 10Effie Mouzeli: trafficserver: Remove XWD routing for /w/rest.php mw-debug backend (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1310103 (https://phabricator.wikimedia.org/T428909) (owner: 10Effie Mouzeli) [22:04:46] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp7007.magru.wmnet [22:09:32] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp7015.magru.wmnet [22:46:01] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp7008.magru.wmnet [22:46:01] !log sukhe@cumin1003 END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_magru [22:51:22] !log sukhe@cumin1003 cookbooks.sre.cdn.roll-reboot finished rebooting cp7016.magru.wmnet [22:51:22] !log sukhe@cumin1003 END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_magru [22:59:57] (03CR) 10Scott French: [C:03+1] trafficserver: Remove XWD routing for /w/rest.php mw-debug backend (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1310103 (https://phabricator.wikimedia.org/T428909) (owner: 10Effie Mouzeli) [23:21:53] (03CR) 10Santiago Faci: [C:03+2] Test Kitchen UI: Deploying v1.4.8 release to production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310139 (https://phabricator.wikimedia.org/T428984) (owner: 10Santiago Faci) [23:23:59] (03Merged) 10jenkins-bot: Test Kitchen UI: Deploying v1.4.8 release to production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310139 (https://phabricator.wikimedia.org/T428984) (owner: 10Santiago Faci) [23:30:15] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1021.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [23:30:21] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1021.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [23:32:15] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [23:32:21] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [23:42:03] (03PS1) 10TrainBranchBot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1310678 [23:42:03] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1310678 (owner: 10TrainBranchBot) [23:49:42] (03Merged) 10jenkins-bot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1310678 (owner: 10TrainBranchBot)