[00:01:49] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Repooling after maintenance db1264 (T433990)', diff saved to https://phabricator.wikimedia.org/P95884 and previous config saved to /var/cache/conftool/dbconfig/20260805-000148-ladsgroup.json [00:01:54] T433990: Optimize echo tables in x1 - https://phabricator.wikimedia.org/T433990 [00:07:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [00:12:37] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95885 and previous config saved to /var/cache/conftool/dbconfig/20260805-001235-ladsgroup.json [00:18:01] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [00:23:23] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95886 and previous config saved to /var/cache/conftool/dbconfig/20260805-002322-ladsgroup.json [00:34:09] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Repooling after maintenance db1264 (T433990)', diff saved to https://phabricator.wikimedia.org/P95887 and previous config saved to /var/cache/conftool/dbconfig/20260805-003408-ladsgroup.json [00:34:14] T433990: Optimize echo tables in x1 - https://phabricator.wikimedia.org/T433990 [00:34:25] !log ladsgroup@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on dbstore1009.eqiad.wmnet with reason: Maintenance [01:12:24] (03PS1) 10TrainBranchBot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1321179 [01:12:24] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1321179 (owner: 10TrainBranchBot) [01:23:18] (03Merged) 10jenkins-bot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1321179 (owner: 10TrainBranchBot) [01:29:43] !log ladsgroup@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2186.codfw.wmnet with reason: Maintenance [01:30:30] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Depooling db2186 (T433990)', diff saved to https://phabricator.wikimedia.org/P95888 and previous config saved to /var/cache/conftool/dbconfig/20260805-013029-ladsgroup.json [01:30:34] T433990: Optimize echo tables in x1 - https://phabricator.wikimedia.org/T433990 [02:00:52] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Repooling after maintenance db2186 (T433990)', diff saved to https://phabricator.wikimedia.org/P95889 and previous config saved to /var/cache/conftool/dbconfig/20260805-020051-ladsgroup.json [02:00:57] T433990: Optimize echo tables in x1 - https://phabricator.wikimedia.org/T433990 [02:01:11] !log mwpresync@deploy1003 Started scap build-images: Publishing wmf/next image [02:07:49] !log mwpresync@deploy1003 Finished scap build-images: Publishing wmf/next image (duration: 06m 37s) [02:08:01] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:11:38] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95890 and previous config saved to /var/cache/conftool/dbconfig/20260805-021137-ladsgroup.json [02:13:57] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:18:01] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:22:24] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95891 and previous config saved to /var/cache/conftool/dbconfig/20260805-022223-ladsgroup.json [02:33:11] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Repooling after maintenance db2186 (T433990)', diff saved to https://phabricator.wikimedia.org/P95892 and previous config saved to /var/cache/conftool/dbconfig/20260805-023310-ladsgroup.json [02:33:16] T433990: Optimize echo tables in x1 - https://phabricator.wikimedia.org/T433990 [02:33:27] !log ladsgroup@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2196.codfw.wmnet with reason: Maintenance [02:34:14] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Depooling db2196 (T433990)', diff saved to https://phabricator.wikimedia.org/P95893 and previous config saved to /var/cache/conftool/dbconfig/20260805-023413-ladsgroup.json [02:56:47] (03PS1) 10Daimona Eaytoy: Rename ce_worklist_articles table to ce_invitation_list_articles [extensions/CampaignEvents] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1321223 (https://phabricator.wikimedia.org/T426102) [02:57:04] (03PS1) 10Daimona Eaytoy: Rename ce_worklist_articles table to ce_invitation_list_articles [extensions/CampaignEvents] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1321224 (https://phabricator.wikimedia.org/T426102) [02:58:26] (03PS1) 10Andrew Bogott: backy2: increase max read rate to 120MB/second [puppet] - 10https://gerrit.wikimedia.org/r/1321225 (https://phabricator.wikimedia.org/T434014) [03:01:06] (03CR) 10Andrew Bogott: [C:03+2] backy2: increase max read rate to 120MB/second [puppet] - 10https://gerrit.wikimedia.org/r/1321225 (https://phabricator.wikimedia.org/T434014) (owner: 10Andrew Bogott) [03:08:16] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Repooling after maintenance db2196 (T433990)', diff saved to https://phabricator.wikimedia.org/P95894 and previous config saved to /var/cache/conftool/dbconfig/20260805-030815-ladsgroup.json [03:08:20] T433990: Optimize echo tables in x1 - https://phabricator.wikimedia.org/T433990 [03:19:03] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95895 and previous config saved to /var/cache/conftool/dbconfig/20260805-031902-ladsgroup.json [03:29:50] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95896 and previous config saved to /var/cache/conftool/dbconfig/20260805-032948-ladsgroup.json [03:40:37] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Repooling after maintenance db2196 (T433990)', diff saved to https://phabricator.wikimedia.org/P95897 and previous config saved to /var/cache/conftool/dbconfig/20260805-034036-ladsgroup.json [03:40:42] T433990: Optimize echo tables in x1 - https://phabricator.wikimedia.org/T433990 [03:40:53] !log ladsgroup@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2197.codfw.wmnet with reason: Maintenance [04:07:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [04:30:41] !log ladsgroup@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2201.codfw.wmnet with reason: Maintenance [05:02:55] (03CR) 10KineticPelagic: [C:03+1] Remove boilerplate language from wmf-rest and wmf-math API modules [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319906 (https://phabricator.wikimedia.org/T433736) (owner: 10ArielGlenn) [05:13:40] FIRING: KubernetesRsyslogDown: rsyslog on wikikube-worker1147:9105 is missing kubernetes logs - https://wikitech.wikimedia.org/wiki/Kubernetes/Logging#Common_issues - https://grafana.wikimedia.org/d/OagQjQmnk?var-server=wikikube-worker1147 - https://alerts.wikimedia.org/?q=alertname%3DKubernetesRsyslogDown [05:18:40] RESOLVED: KubernetesRsyslogDown: rsyslog on wikikube-worker1147:9105 is missing kubernetes logs - https://wikitech.wikimedia.org/wiki/Kubernetes/Logging#Common_issues - https://grafana.wikimedia.org/d/OagQjQmnk?var-server=wikikube-worker1147 - https://alerts.wikimedia.org/?q=alertname%3DKubernetesRsyslogDown [05:18:54] !log ladsgroup@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2215.codfw.wmnet with reason: Maintenance [05:19:40] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Depooling db2215 (T433990)', diff saved to https://phabricator.wikimedia.org/P95898 and previous config saved to /var/cache/conftool/dbconfig/20260805-051939-ladsgroup.json [05:19:45] T433990: Optimize echo tables in x1 - https://phabricator.wikimedia.org/T433990 [05:49:19] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Repooling after maintenance db2215 (T433990)', diff saved to https://phabricator.wikimedia.org/P95899 and previous config saved to /var/cache/conftool/dbconfig/20260805-054918-ladsgroup.json [05:49:23] T433990: Optimize echo tables in x1 - https://phabricator.wikimedia.org/T433990 [06:00:04] Deploy window MediaWiki infrastructure (UTC early) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260805T0600) [06:00:06] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95900 and previous config saved to /var/cache/conftool/dbconfig/20260805-060004-ladsgroup.json [06:10:52] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95901 and previous config saved to /var/cache/conftool/dbconfig/20260805-061051-ladsgroup.json [06:18:01] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [06:21:39] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Repooling after maintenance db2215 (T433990)', diff saved to https://phabricator.wikimedia.org/P95902 and previous config saved to /var/cache/conftool/dbconfig/20260805-062137-ladsgroup.json [06:21:44] T433990: Optimize echo tables in x1 - https://phabricator.wikimedia.org/T433990 [06:21:55] !log ladsgroup@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2231.codfw.wmnet with reason: Maintenance [06:22:41] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Depooling db2231 (T433990)', diff saved to https://phabricator.wikimedia.org/P95903 and previous config saved to /var/cache/conftool/dbconfig/20260805-062240-ladsgroup.json [06:37:22] (03PS1) 10Kevin Bazira: ml-services: update qwen36-27b and qwen3-14b isvcs to support tool calling [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321438 (https://phabricator.wikimedia.org/T433977) [06:40:38] (03CR) 10Kevin Bazira: [C:03+2] ml-services: update qwen36-27b and qwen3-14b isvcs to support tool calling [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321438 (https://phabricator.wikimedia.org/T433977) (owner: 10Kevin Bazira) [06:43:03] (03Merged) 10jenkins-bot: ml-services: update qwen36-27b and qwen3-14b isvcs to support tool calling [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321438 (https://phabricator.wikimedia.org/T433977) (owner: 10Kevin Bazira) [06:45:01] !log kevinbazira@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . [06:51:41] (03PS1) 10Brouberol: admin_ng/dse-k8s-eqiad/mediawiki-dumps-legacy: increase cpu limit to 800 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321449 [06:52:07] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Repooling after maintenance db2231 (T433990)', diff saved to https://phabricator.wikimedia.org/P95904 and previous config saved to /var/cache/conftool/dbconfig/20260805-065206-ladsgroup.json [06:52:13] T433990: Optimize echo tables in x1 - https://phabricator.wikimedia.org/T433990 [06:55:17] (03PS2) 10Brouberol: admin_ng/dse-k8s-eqiad/mediawiki-dumps-legacy: increase cpu limit to 800 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321449 [07:00:04] Amir1, urbanecm, and awight: #bothumor Q:Why did functions stop calling each other? A:They had arguments. Rise for UTC morning backport window . (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260805T0700). [07:00:04] No Gerrit patches in the queue for this window AFAICS. [07:02:54] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95905 and previous config saved to /var/cache/conftool/dbconfig/20260805-070253-ladsgroup.json [07:07:46] (03PS1) 10Marostegui: m1 proxy: Test db1164 connectivity [puppet] - 10https://gerrit.wikimedia.org/r/1321463 (https://phabricator.wikimedia.org/T434043) [07:09:15] (03CR) 10Arnaudb: [C:03+2] gerrit: concatenation typo on dns wipe-cache [cookbooks] - 10https://gerrit.wikimedia.org/r/1320921 (owner: 10Arnaudb) [07:09:25] (03CR) 10Marostegui: [C:03+2] m1 proxy: Test db1164 connectivity [puppet] - 10https://gerrit.wikimedia.org/r/1321463 (https://phabricator.wikimedia.org/T434043) (owner: 10Marostegui) [07:11:11] (03PS1) 10Ayounsi: Drain eqsin-eqiad HE transport [homer/public] - 10https://gerrit.wikimedia.org/r/1321465 [07:13:41] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95906 and previous config saved to /var/cache/conftool/dbconfig/20260805-071340-ladsgroup.json [07:13:50] (03PS1) 10Marostegui: Revert "m1 proxy: Test db1164 connectivity" [puppet] - 10https://gerrit.wikimedia.org/r/1321466 [07:13:52] (03Merged) 10jenkins-bot: gerrit: concatenation typo on dns wipe-cache [cookbooks] - 10https://gerrit.wikimedia.org/r/1320921 (owner: 10Arnaudb) [07:16:27] (03CR) 10Marostegui: [C:03+2] Revert "m1 proxy: Test db1164 connectivity" [puppet] - 10https://gerrit.wikimedia.org/r/1321466 (owner: 10Marostegui) [07:18:55] (03CR) 10Slyngshede: [C:03+2] Geo-map: August update for Meta [dns] - 10https://gerrit.wikimedia.org/r/1320821 (owner: 10Slyngshede) [07:19:04] !log slyngshede@dns1004 START - running authdns-update [07:20:04] (03CR) 10Ayounsi: [C:03+2] Drain eqsin-eqiad HE transport [homer/public] - 10https://gerrit.wikimedia.org/r/1321465 (owner: 10Ayounsi) [07:21:13] !log slyngshede@dns1004 END - running authdns-update [07:24:27] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Repooling after maintenance db2231 (T433990)', diff saved to https://phabricator.wikimedia.org/P95908 and previous config saved to /var/cache/conftool/dbconfig/20260805-072426-ladsgroup.json [07:24:32] T433990: Optimize echo tables in x1 - https://phabricator.wikimedia.org/T433990 [07:24:34] (03Merged) 10jenkins-bot: Drain eqsin-eqiad HE transport [homer/public] - 10https://gerrit.wikimedia.org/r/1321465 (owner: 10Ayounsi) [07:24:44] !log ladsgroup@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2249.codfw.wmnet with reason: Maintenance [07:25:31] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Depooling db2249 (T433990)', diff saved to https://phabricator.wikimedia.org/P95909 and previous config saved to /var/cache/conftool/dbconfig/20260805-072529-ladsgroup.json [07:32:58] (03PS1) 10Jcrespo: mariadb: Switchover backup statistics for backup services [puppet] - 10https://gerrit.wikimedia.org/r/1321468 (https://phabricator.wikimedia.org/T434043) [07:33:13] (03PS2) 10Jcrespo: mariadb: Switchover backup statistics for backup services [puppet] - 10https://gerrit.wikimedia.org/r/1321468 (https://phabricator.wikimedia.org/T434043) [07:39:29] (03PS3) 10Jcrespo: mariadb: Switchover backup statistics primary db for dbbackups [puppet] - 10https://gerrit.wikimedia.org/r/1321468 (https://phabricator.wikimedia.org/T434043) [07:40:13] (03PS4) 10Jcrespo: mariadb: Switchover backup statistics primary db for dbbackups [puppet] - 10https://gerrit.wikimedia.org/r/1321468 (https://phabricator.wikimedia.org/T434043) [07:53:02] !log marostegui@cumin1003 conftool action : set/pooled=no; selector: name=clouddb1013.eqiad.wmnet,service=s1 [07:53:14] !log Depool clouddb1013:s1 T409557 [07:53:19] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [07:53:19] T409557: Productionize new clouddb* hosts (clouddb1022-1033) - https://phabricator.wikimedia.org/T409557 [07:53:44] !log Depool clouddb1013:s1 T434048 [07:53:48] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [07:53:48] T434048: Decommission clouddb1013-clouddb1020 - https://phabricator.wikimedia.org/T434048 [07:54:20] !log marostegui@cumin1003 conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s2 [07:54:23] !log marostegui@cumin1003 conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s7 [07:54:33] !log Depool clouddb1014 (s2,s7) T434048 [07:54:36] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [07:56:51] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Repooling after maintenance db2249 (T433990)', diff saved to https://phabricator.wikimedia.org/P95910 and previous config saved to /var/cache/conftool/dbconfig/20260805-075650-ladsgroup.json [07:56:55] T433990: Optimize echo tables in x1 - https://phabricator.wikimedia.org/T433990 [07:57:42] !log marostegui@cumin1003 conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4 [07:57:46] !log marostegui@cumin1003 conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6 [07:57:55] !log Depool clouddb1015 (s4,s6) T434048 [07:57:59] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [07:59:03] !log marostegui@cumin1003 conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5 [07:59:07] !log marostegui@cumin1003 conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8 [07:59:19] !log Depool clouddb1016 (s5,s8) T434048 [07:59:22] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [07:59:23] T434048: Decommission clouddb1013-clouddb1020 - https://phabricator.wikimedia.org/T434048 [08:00:05] jnuche and jeena: Time to snap out of that daydream and deploy MediaWiki train - Utc-0+Utc-7 Version. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260805T0800). [08:00:33] morning, train will happen in 5m [08:01:24] !log marostegui@cumin1003 conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1 [08:01:30] !log Depool clouddb1017 (s1) T434048 [08:01:34] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [08:01:47] !log marostegui@cumin1003 conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s2 [08:01:50] !log marostegui@cumin1003 conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s7 [08:01:58] !log Depool clouddb1018 (s2,s7) T434048 [08:02:02] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [08:02:10] !log marostegui@cumin1003 conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s5 [08:02:14] !log marostegui@cumin1003 conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s8 [08:02:21] !log Depool clouddb1020 (s5,s8) T434048 [08:02:24] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [08:05:50] !log jynus@cumin1003 START - Cookbook sre.hosts.decommission for hosts db1150.eqiad.wmnet [08:06:31] (03PS1) 10TrainBranchBot: group1 to 1.47.0-wmf.14 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1321513 (https://phabricator.wikimedia.org/T430833) [08:06:34] (03CR) 10TrainBranchBot: [C:03+2] "Initiated by jnuche@deploy1003" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1321513 (https://phabricator.wikimedia.org/T430833) (owner: 10TrainBranchBot) [08:07:38] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95911 and previous config saved to /var/cache/conftool/dbconfig/20260805-080737-ladsgroup.json [08:07:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [08:08:38] (03Merged) 10jenkins-bot: group1 to 1.47.0-wmf.14 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1321513 (https://phabricator.wikimedia.org/T430833) (owner: 10TrainBranchBot) [08:11:02] !log jynus@cumin1003 START - Cookbook sre.dns.netbox [08:14:10] !log jayme@cumin1003 START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1155.eqiad.wmnet [08:14:14] !log jayme@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1155.eqiad.wmnet [08:14:20] (03PS2) 10Jcrespo: mediabackups: Update versitygw systemd unit file to SIGHUP on reload [puppet] - 10https://gerrit.wikimedia.org/r/1321123 (https://phabricator.wikimedia.org/T430023) [08:14:50] !log jayme@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1155.eqiad.wmnet [08:15:06] !log jnuche@deploy1003 rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.14 refs T430833 [08:15:11] T430833: 1.47.0-wmf.14 deployment blockers - https://phabricator.wikimedia.org/T430833 [08:15:22] (03CR) 10CI reject: [V:04-1] mediabackups: Update versitygw systemd unit file to SIGHUP on reload [puppet] - 10https://gerrit.wikimedia.org/r/1321123 (https://phabricator.wikimedia.org/T430023) (owner: 10Jcrespo) [08:15:42] !log jynus@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" [08:15:45] (03PS3) 10Jcrespo: mediabackups: Update versitygw systemd unit file to SIGHUP on reload [puppet] - 10https://gerrit.wikimedia.org/r/1321123 (https://phabricator.wikimedia.org/T430023) [08:15:49] (03CR) 10Jcrespo: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1321123 (https://phabricator.wikimedia.org/T430023) (owner: 10Jcrespo) [08:17:03] !log jynus@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" [08:17:03] !log jynus@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [08:17:05] !log jynus@cumin1003 END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1150.eqiad.wmnet [08:17:41] !log jynus@cumin1003 START - Cookbook sre.hosts.decommission for hosts db1171.eqiad.wmnet [08:17:50] jayme@cumin1003 renumber-node (PID 731443) is awaiting input [08:18:18] (03CR) 10Jcrespo: [C:03+2] mariadb: Remove all references on puppet to db1150 & db1171 [puppet] - 10https://gerrit.wikimedia.org/r/1320923 (https://phabricator.wikimedia.org/T433825) (owner: 10Jcrespo) [08:18:24] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95912 and previous config saved to /var/cache/conftool/dbconfig/20260805-081823-ladsgroup.json [08:18:24] (03PS3) 10Jcrespo: mariadb: Remove all references on puppet to db1150 & db1171 [puppet] - 10https://gerrit.wikimedia.org/r/1320923 (https://phabricator.wikimedia.org/T433825) [08:21:17] (03CR) 10Jcrespo: [C:03+2] mariadb: Remove all references on puppet to db1150 & db1171 [puppet] - 10https://gerrit.wikimedia.org/r/1320923 (https://phabricator.wikimedia.org/T433825) (owner: 10Jcrespo) [08:22:47] !log jynus@cumin1003 START - Cookbook sre.dns.netbox [08:24:08] 10ops-eqiad, 06Data-Persistence, 10Data-Persistence-Backup, 10database-backups, and 3 others: decommission db1150.eqiad.wmnet - https://phabricator.wikimedia.org/T433825#12186238 (10jcrespo) [08:24:24] 10ops-eqiad, 06Data-Persistence, 10Data-Persistence-Backup, 10database-backups, and 3 others: decommission db1150.eqiad.wmnet - https://phabricator.wikimedia.org/T433825#12186245 (10jcrespo) a:05jcrespo→03None [08:24:57] 10ops-eqiad, 06Data-Persistence, 10Data-Persistence-Backup, 10database-backups, and 3 others: decommission db1150.eqiad.wmnet - https://phabricator.wikimedia.org/T433825#12186247 (10jcrespo) This is ready for dcops to unrack or repurpose. [08:25:39] (03CR) 10Effie Mouzeli: [C:03+2] memcached: implement refresh_certs [puppet] - 10https://gerrit.wikimedia.org/r/1320809 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [08:27:25] (03PS5) 10Effie Mouzeli: memcached: reload_certs on certificate renewal [puppet] - 10https://gerrit.wikimedia.org/r/1320940 (https://phabricator.wikimedia.org/T353511) [08:27:40] !log jynus@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" [08:29:07] (03CR) 10Fabfur: mediabackups: Update versitygw systemd unit file to SIGHUP on reload (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1321123 (https://phabricator.wikimedia.org/T430023) (owner: 10Jcrespo) [08:29:10] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Repooling after maintenance db2249 (T433990)', diff saved to https://phabricator.wikimedia.org/P95913 and previous config saved to /var/cache/conftool/dbconfig/20260805-082908-ladsgroup.json [08:29:14] T433990: Optimize echo tables in x1 - https://phabricator.wikimedia.org/T433990 [08:29:54] !log jynus@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" [08:29:54] !log jynus@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [08:29:55] !log jynus@cumin1003 END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1171.eqiad.wmnet [08:31:04] 10ops-eqiad, 06Data-Persistence, 10Data-Persistence-Backup, 10database-backups, and 2 others: decommission db1171.eqiad.wmnet - https://phabricator.wikimedia.org/T433826#12186266 (10jcrespo) a:05jcrespo→03None [08:31:45] 10ops-eqiad, 06Data-Persistence, 10Data-Persistence-Backup, 10database-backups, and 2 others: decommission db1171.eqiad.wmnet - https://phabricator.wikimedia.org/T433826#12186274 (10jcrespo) This is ready for dcops unracking or repurpose. [08:32:06] (03CR) 10Klausman: [C:03+1] "Nit: minor typo in commitmsg: "git"" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321449 (owner: 10Brouberol) [08:33:20] (03PS4) 10Jcrespo: bacula: Remove last references to backup1003 & backup2003 [puppet] - 10https://gerrit.wikimedia.org/r/1320928 (https://phabricator.wikimedia.org/T420506) [08:34:58] (03PS3) 10Brouberol: admin_ng/dse-k8s-eqiad/mediawiki-dumps-legacy: increase cpu limit to 800 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321449 [08:35:53] (03CR) 10Fabfur: mediabackups: Update versitygw systemd unit file to SIGHUP on reload (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1321123 (https://phabricator.wikimedia.org/T430023) (owner: 10Jcrespo) [08:37:54] !log jayme@cumin1003 START - Cookbook sre.hosts.reimage for host wikikube-worker1155.eqiad.wmnet with OS trixie [08:38:24] !log jayme@cumin1003 START - Cookbook sre.hosts.move-vlan for host wikikube-worker1155 [08:41:10] (03CR) 10Effie Mouzeli: [C:03+1] hieradata: Prepare conf1007 for reimage to bookworm [puppet] - 10https://gerrit.wikimedia.org/r/1321107 (https://phabricator.wikimedia.org/T428495) (owner: 10Scott French) [08:41:26] jayme@cumin1003 renumber-node (PID 731443) is awaiting input [08:42:48] (03CR) 10Slyngshede: [C:04-1] "Very nice, the -1 is only because I'd like to retain the ability to disable Redis." [puppet] - 10https://gerrit.wikimedia.org/r/1320817 (https://phabricator.wikimedia.org/T273950) (owner: 10Majavah) [08:43:12] (03CR) 10Effie Mouzeli: [C:03+1] hieradata: Temporarily move etcd replication to conf1007 [puppet] - 10https://gerrit.wikimedia.org/r/1321108 (https://phabricator.wikimedia.org/T428495) (owner: 10Scott French) [08:44:51] (03CR) 10Brouberol: [C:03+2] admin_ng/dse-k8s-eqiad/mediawiki-dumps-legacy: increase cpu limit to 800 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321449 (owner: 10Brouberol) [08:47:04] 10ops-eqiad, 06SRE, 06Data-Persistence, 10Data-Persistence-Backup, and 3 others: decommission db1150.eqiad.wmnet - https://phabricator.wikimedia.org/T433825#12186311 (10jcrespo) p:05High→03Medium [08:47:08] 10ops-eqiad, 06Data-Persistence, 10Data-Persistence-Backup, 10database-backups, and 2 others: decommission db1171.eqiad.wmnet - https://phabricator.wikimedia.org/T433826#12186312 (10jcrespo) p:05High→03Medium [08:53:52] (03PS1) 10Blake: httpbb: Release v0.0.6 [software/httpbb] - 10https://gerrit.wikimedia.org/r/1321516 (https://phabricator.wikimedia.org/T434052) [08:54:29] (03PS1) 10Arnaudb: sre.gitlab.failover: check discovery CNAMEs and gitlab-ssh records [cookbooks] - 10https://gerrit.wikimedia.org/r/1320977 (https://phabricator.wikimedia.org/T430655) [08:54:30] (03CR) 10Arnaudb: "this has been tested with `test-cookbook --dry-run -c 1320978 sre.gitlab.failover --switch-from gitlab1004 --switch-to gitlab2002` → I tes" [cookbooks] - 10https://gerrit.wikimedia.org/r/1320977 (https://phabricator.wikimedia.org/T430655) (owner: 10Arnaudb) [08:54:30] (03CR) 10Effie Mouzeli: [C:03+2] api-gateway:: remove API Gateway specific ratelimiting logic #3 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319563 (https://phabricator.wikimedia.org/T428625) (owner: 10Aaron Schulz) [08:54:57] (03PS2) 10Arnaudb: sre.gitlab: fix rollback crash and minor pre-existing issues [cookbooks] - 10https://gerrit.wikimedia.org/r/1320978 (https://phabricator.wikimedia.org/T430655) [08:54:57] (03CR) 10Arnaudb: "as mentioned in the earlier change, this is the patch that has been tested with `test-cookbook --dry-run -c 1320978 sre.gitlab.failover --" [cookbooks] - 10https://gerrit.wikimedia.org/r/1320978 (https://phabricator.wikimedia.org/T430655) (owner: 10Arnaudb) [08:55:19] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-private-users for aramirezwmf - https://phabricator.wikimedia.org/T433999#12186337 (10Arnoldokoth) 05Open→03In progress [08:56:03] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-private-users for aramirezwmf - https://phabricator.wikimedia.org/T433999#12186343 (10Arnoldokoth) [08:56:25] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-private-users for aramirezwmf - https://phabricator.wikimedia.org/T433999#12186345 (10Arnoldokoth) [08:57:02] (03CR) 10Ayounsi: "No objection to put ulsfo back at #1 choice, but shouldn't eqsin be at #2?" [dns] - 10https://gerrit.wikimedia.org/r/1306727 (owner: 10BCornwall) [08:57:03] (03Merged) 10jenkins-bot: api-gateway:: remove API Gateway specific ratelimiting logic #3 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319563 (https://phabricator.wikimedia.org/T428625) (owner: 10Aaron Schulz) [08:57:42] (03CR) 10Slyngshede: [C:03+1] hieradata: Remove data for sso Cloud VPS project [puppet] - 10https://gerrit.wikimedia.org/r/1320815 (owner: 10Majavah) [08:59:44] 06SRE, 10SRE-Access-Requests: Requesting access to deployment for Chlod Alejandro - https://phabricator.wikimedia.org/T433791#12186349 (10Arnoldokoth) 05Open→03In progress [09:01:05] (03CR) 10Slyngshede: [C:04-1] "cloudinfra-idp-2 can't reach rdb2011.codfw.wmnet and shouldn't." [puppet] - 10https://gerrit.wikimedia.org/r/1320824 (https://phabricator.wikimedia.org/T433963) (owner: 10Majavah) [09:02:11] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-private-users for aramirezwmf - https://phabricator.wikimedia.org/T433999#12186351 (10Arnoldokoth) [09:02:38] (03CR) 10Majavah: [C:03+2] hieradata: Remove data for sso Cloud VPS project [puppet] - 10https://gerrit.wikimedia.org/r/1320815 (owner: 10Majavah) [09:03:24] (03CR) 10Majavah: [V:03+1] "This role controls cloudidp2001-dev.codfw.wmnet (cloudidp-dev.wikimedia.org), not cloudinfra-idp-2 (idp.wmcloud.org)?" [puppet] - 10https://gerrit.wikimedia.org/r/1320824 (https://phabricator.wikimedia.org/T433963) (owner: 10Majavah) [09:05:38] !log marostegui@cumin1003 START - Cookbook sre.mysql.depool depool db1152: Cloning [09:05:38] !log marostegui@cumin1003 START - Cookbook sre.mysql.parsercache [09:06:35] !log marostegui@cumin1003 END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) [09:06:35] !log marostegui@cumin1003 END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Cloning [09:06:40] 06SRE, 06Infrastructure-Foundations, 10netops: InboundInterfaceErrors alerts firing for Nokia switches on v25.10.1 - https://phabricator.wikimedia.org/T412733#12186371 (10ayounsi) Nokia 26.7.1 release notes have this in "resolved issues": > On the 7220 IXR-Dx, STP packets received on the mgmt0 management por... [09:08:35] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [09:08:44] !log fceratto@cumin1003 END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) [09:09:40] !log jayme@cumin1003 START - Cookbook sre.dns.netbox [09:09:41] !log marostegui@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: cloning [09:10:15] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [09:10:26] !log fceratto@cumin1003 END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) [09:11:22] (03PS1) 10Marostegui: mariadb: Productionize db1267 [puppet] - 10https://gerrit.wikimedia.org/r/1321518 (https://phabricator.wikimedia.org/T407942) [09:12:46] (03CR) 10Marostegui: [C:03+2] mariadb: Productionize db1267 [puppet] - 10https://gerrit.wikimedia.org/r/1321518 (https://phabricator.wikimedia.org/T407942) (owner: 10Marostegui) [09:14:03] !log jayme@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" [09:14:08] !log jayme@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003" [09:14:08] !log jayme@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [09:14:08] !log jayme@cumin1003 START - Cookbook sre.dns.wipe-cache wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [09:14:12] !log jayme@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [09:14:13] !log jayme@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1155 [09:17:16] !log push pfw policies - T434038 [09:17:19] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [09:17:19] (03PS1) 10AOkoth: admin: add aramirezwmf to analytics-privatedata [puppet] - 10https://gerrit.wikimedia.org/r/1321519 (https://phabricator.wikimedia.org/T433999) [09:18:01] RECOVERY - Backup freshness on backup1014 is OK: Fresh: 144 jobs https://wikitech.wikimedia.org/wiki/Bacula%23Monitoring [09:18:10] (03CR) 10CI reject: [V:04-1] admin: add aramirezwmf to analytics-privatedata [puppet] - 10https://gerrit.wikimedia.org/r/1321519 (https://phabricator.wikimedia.org/T433999) (owner: 10AOkoth) [09:19:40] (03PS2) 10AOkoth: admin: add aramirezwmf to analytics-privatedata [puppet] - 10https://gerrit.wikimedia.org/r/1321519 (https://phabricator.wikimedia.org/T433999) [09:20:31] (03PS36) 10Federico Ceratto: mysql: update replication source [cookbooks] - 10https://gerrit.wikimedia.org/r/1238368 (https://phabricator.wikimedia.org/T373436) [09:20:38] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [09:20:40] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to analytics-private-users for aramirezwmf - https://phabricator.wikimedia.org/T433999#12186395 (10Arnoldokoth) [09:20:53] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [09:22:02] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [09:22:43] !log fceratto@cumin1003 END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) [09:22:47] !log jiji@deploy1003 helmfile [codfw] START helmfile.d/services/rest-gateway: apply [09:22:50] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [09:23:12] !log jiji@deploy1003 helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply [09:23:15] !log fceratto@cumin1003 END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) [09:23:30] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [09:23:43] (03CR) 10Effie Mouzeli: [C:03+1] hieradata: Use component/zookeeper34 on all configcluster hosts [puppet] - 10https://gerrit.wikimedia.org/r/1321106 (https://phabricator.wikimedia.org/T428495) (owner: 10Scott French) [09:23:50] (03PS2) 10Majavah: P:idp::standalone: Remove profile [puppet] - 10https://gerrit.wikimedia.org/r/1320816 [09:23:50] (03PS2) 10Majavah: hieradata: Migrate cloudidp-dev to Redis ticket storage [puppet] - 10https://gerrit.wikimedia.org/r/1320824 (https://phabricator.wikimedia.org/T433963) [09:23:50] (03PS4) 10Majavah: P:idp: Remove support for memcached [puppet] - 10https://gerrit.wikimedia.org/r/1320817 (https://phabricator.wikimedia.org/T273950) [09:24:15] (03CR) 10Majavah: P:idp: Remove support for memcached (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1320817 (https://phabricator.wikimedia.org/T273950) (owner: 10Majavah) [09:24:40] !log fceratto@cumin1003 END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) [09:24:47] !log jiji@deploy1003 helmfile [eqiad] START helmfile.d/services/rest-gateway: apply [09:25:06] !log jiji@deploy1003 helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply [09:27:04] !log jayme@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1155 [09:27:04] !log jayme@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1155 [09:29:06] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for iberker - https://phabricator.wikimedia.org/T433258#12186410 (10Arnoldokoth) [09:29:37] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for iberker - https://phabricator.wikimedia.org/T433258#12186412 (10Arnoldokoth) a:05IBerker-WMF→03SCherukuwada [09:30:22] 06SRE, 10SRE-Access-Requests: Requesting access to deployment for Chlod Alejandro - https://phabricator.wikimedia.org/T433791#12186419 (10Arnoldokoth) Adding @KFrancis to verify NDA. [09:31:40] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [09:31:53] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [09:32:13] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [09:32:25] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [09:37:46] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for iberker - https://phabricator.wikimedia.org/T433258#12186460 (10Arnoldokoth) 05Open→03In progress [09:39:17] (03CR) 10Arnaudb: [C:03+1] "looks good to me!" [puppet] - 10https://gerrit.wikimedia.org/r/1321519 (https://phabricator.wikimedia.org/T433999) (owner: 10AOkoth) [09:40:59] !log jayme@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage [09:43:58] !log jayme@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage [09:43:58] !log marostegui@cumin1003 START - Cookbook sre.mysql.pool pool db1152: after cloning [09:43:59] !log marostegui@cumin1003 START - Cookbook sre.mysql.parsercache [09:43:59] !log marostegui@cumin1003 END (FAIL) - Cookbook sre.mysql.parsercache (exit_code=99) [09:44:00] !log marostegui@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: after cloning [09:48:36] o/ I'll be deploying this changeprop patch shortly https://gerrit.wikimedia.org/r/c/operations/deployment-charts/+/1320762 [09:49:02] (03PS1) 10Mvolz: datahub-next: remove duplicate key [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321520 [09:52:13] !log marostegui@cumin1003 dbctl commit (dc=all): 'Pool back ms1', diff saved to https://phabricator.wikimedia.org/P95918 and previous config saved to /var/cache/conftool/dbconfig/20260805-095212-marostegui.json [10:00:04] Deploy window MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260805T1000) [10:04:00] !log jayme@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1155.eqiad.wmnet with OS trixie [10:04:12] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [10:04:22] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [10:04:41] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [10:04:51] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [10:04:58] (03PS1) 10Mvolz: kartotherian: remove duplicate key [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321523 [10:05:03] !log aikochou@deploy1003 helmfile [eqiad] START helmfile.d/services/changeprop: sync [10:05:10] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [10:05:24] !log aikochou@deploy1003 helmfile [eqiad] DONE helmfile.d/services/changeprop: sync [10:05:41] !log fceratto@cumin1003 END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) [10:07:10] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [10:07:25] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [10:08:29] (03CR) 10Effie Mouzeli: [C:03+1] mediawiki: Disable legacy `short_urls` on vhosts where it does not work [puppet] - 10https://gerrit.wikimedia.org/r/1290104 (https://phabricator.wikimedia.org/T107188) (owner: 10Krinkle) [10:09:14] !log aikochou@deploy1003 helmfile [codfw] START helmfile.d/services/changeprop: sync [10:09:24] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [10:09:37] !log aikochou@deploy1003 helmfile [codfw] DONE helmfile.d/services/changeprop: sync [10:09:52] !log fceratto@cumin1003 END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) [10:10:14] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [10:10:39] !log fceratto@cumin1003 END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) [10:11:08] (03PS1) 10Mvolz: citoid: update version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321524 [10:11:32] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [10:11:45] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [10:12:18] (03PS1) 10Cathal Mooney: eqsin: make new Arelion codfw primary transport [homer/public] - 10https://gerrit.wikimedia.org/r/1321525 (https://phabricator.wikimedia.org/T424839) [10:14:22] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [10:14:41] !log fceratto@cumin1003 END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) [10:14:57] !log jayme@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1155.eqiad.wmnet [10:14:59] !log jayme@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1155.eqiad.wmnet [10:15:00] !log jayme@cumin1003 END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1155.eqiad.wmnet [10:16:26] (03CR) 10Ayounsi: eqsin: make new Arelion codfw primary transport (031 comment) [homer/public] - 10https://gerrit.wikimedia.org/r/1321525 (https://phabricator.wikimedia.org/T424839) (owner: 10Cathal Mooney) [10:16:45] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [10:17:05] !log fceratto@cumin1003 END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) [10:18:01] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [10:22:56] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [10:23:10] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [10:23:19] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [10:23:31] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [10:23:38] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [10:23:50] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [10:23:56] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [10:24:04] PROBLEM - orchestrator resolve cache non-FQDNs on dborch1002 is CRITICAL: CRITICAL: 1 non-FQDN entries in orchestrator resolve cache: https://wikitech.wikimedia.org/wiki/Orchestrator [10:24:08] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [10:24:28] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [10:24:38] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [10:24:45] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [10:24:54] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [10:27:15] (03PS2) 10Cathal Mooney: eqsin: make new Arelion codfw primary transport [homer/public] - 10https://gerrit.wikimedia.org/r/1321525 (https://phabricator.wikimedia.org/T424839) [10:28:14] (03CR) 10Cathal Mooney: eqsin: make new Arelion codfw primary transport (031 comment) [homer/public] - 10https://gerrit.wikimedia.org/r/1321525 (https://phabricator.wikimedia.org/T424839) (owner: 10Cathal Mooney) [10:28:57] (03PS5) 10Blake: Update to v1.39.0 [debs/envoyproxy] (v1.39) - 10https://gerrit.wikimedia.org/r/1321526 (https://phabricator.wikimedia.org/T421418) [10:30:35] 06SRE, 06Infrastructure-Foundations, 10netops: InboundInterfaceErrors alerts firing for Nokia switches on v25.10.1 - https://phabricator.wikimedia.org/T412733#12186604 (10cmooney) For the record does look to be fixed on lswtest-d8-eqiad, which we upgraded a few weeks ago {F97254377 width=400} [10:32:14] (03CR) 10Ayounsi: [C:03+1] eqsin: make new Arelion codfw primary transport [homer/public] - 10https://gerrit.wikimedia.org/r/1321525 (https://phabricator.wikimedia.org/T424839) (owner: 10Cathal Mooney) [10:32:55] (03CR) 10Cathal Mooney: [C:03+2] eqsin: make new Arelion codfw primary transport [homer/public] - 10https://gerrit.wikimedia.org/r/1321525 (https://phabricator.wikimedia.org/T424839) (owner: 10Cathal Mooney) [10:34:20] (03Merged) 10jenkins-bot: eqsin: make new Arelion codfw primary transport [homer/public] - 10https://gerrit.wikimedia.org/r/1321525 (https://phabricator.wikimedia.org/T424839) (owner: 10Cathal Mooney) [10:46:47] (03PS1) 10Majavah: P:toolforge::k8s::haproxy: Track per-IP connection count [puppet] - 10https://gerrit.wikimedia.org/r/1321530 [10:48:20] (03PS2) 10Majavah: P:toolforge::k8s::haproxy: Track per-IP connection count [puppet] - 10https://gerrit.wikimedia.org/r/1321530 [10:51:29] (03CR) 10David Caro: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1321530 (owner: 10Majavah) [10:52:05] (03CR) 10Majavah: [C:03+2] P:toolforge::k8s::haproxy: Track per-IP connection count [puppet] - 10https://gerrit.wikimedia.org/r/1321530 (owner: 10Majavah) [10:52:32] (03PS1) 10Marostegui: clouddb1013-clouddb1020: Disable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1321531 (https://phabricator.wikimedia.org/T434048) [10:53:23] (03CR) 10Marostegui: "FYI" [puppet] - 10https://gerrit.wikimedia.org/r/1321531 (https://phabricator.wikimedia.org/T434048) (owner: 10Marostegui) [10:53:25] (03CR) 10Marostegui: [C:03+2] clouddb1013-clouddb1020: Disable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1321531 (https://phabricator.wikimedia.org/T434048) (owner: 10Marostegui) [10:55:34] (03CR) 10Effie Mouzeli: memcached: reload_certs on certificate renewal (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1320940 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [10:55:37] (03CR) 10Effie Mouzeli: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1320940 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [10:58:47] (03PS1) 10Brouberol: airflow: don't inject newlines in single-line secrets [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321533 [11:00:05] mvolz: It is that lovely time of the day again! You are hereby commanded to deploy Services – Citoid / Zotero. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260805T1100). [11:00:09] (03CR) 10Trueg: [C:03+1] "nice" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321533 (owner: 10Brouberol) [11:02:16] (03CR) 10Blake: [C:03+1] memcached: reload_certs on certificate renewal [puppet] - 10https://gerrit.wikimedia.org/r/1320940 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [11:02:21] (03CR) 10Brouberol: [C:03+2] airflow: don't inject newlines in single-line secrets [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321533 (owner: 10Brouberol) [11:04:21] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply [11:04:45] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply [11:05:33] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [11:05:48] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [11:06:01] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [11:06:11] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [11:06:27] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [11:06:42] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [11:06:45] (03PS37) 10Federico Ceratto: mysql: update replication source [cookbooks] - 10https://gerrit.wikimedia.org/r/1238368 (https://phabricator.wikimedia.org/T373436) [11:07:02] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [11:07:13] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [11:07:33] (03Abandoned) 10Effie Mouzeli: restbase: remove mw-parsoid hardcoded listener [puppet] - 10https://gerrit.wikimedia.org/r/1265441 (owner: 10Effie Mouzeli) [11:08:00] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [11:08:12] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [11:08:57] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [11:09:09] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [11:10:29] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [11:10:41] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [11:11:58] (03CR) 10Kevin Bazira: [C:03+1] ml-services: declare ephemeral-storage for all ml-staging isvcs [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320919 (https://phabricator.wikimedia.org/T431089) (owner: 10Bartosz Wójtowicz) [11:13:02] !log mvolz@deploy1003 helmfile [staging] START helmfile.d/services/citoid: apply [11:13:06] !log mvolz@deploy1003 helmfile [staging] DONE helmfile.d/services/citoid: apply [11:13:22] (03CR) 10Mvolz: [C:03+2] citoid: update version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321524 (owner: 10Mvolz) [11:16:02] (03Merged) 10jenkins-bot: citoid: update version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321524 (owner: 10Mvolz) [11:16:04] (03CR) 10Effie Mouzeli: [C:03+2] memcached: reload_certs on certificate renewal [puppet] - 10https://gerrit.wikimedia.org/r/1320940 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [11:18:32] !log mvolz@deploy1003 helmfile [staging] START helmfile.d/services/citoid: apply [11:18:45] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [11:18:48] !log mvolz@deploy1003 helmfile [staging] DONE helmfile.d/services/citoid: apply [11:18:58] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [11:20:02] (03PS1) 10Effie Mouzeli: hieradata: use PKI on mc2046 [puppet] - 10https://gerrit.wikimedia.org/r/1321535 (https://phabricator.wikimedia.org/T353511) [11:20:32] (03PS2) 10Effie Mouzeli: hieradata: use PKI on mc2046 [puppet] - 10https://gerrit.wikimedia.org/r/1321535 (https://phabricator.wikimedia.org/T353511) [11:21:31] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [11:21:42] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [11:21:55] (03CR) 10Bartosz Wójtowicz: [C:03+2] ml-services: declare ephemeral-storage for all ml-staging isvcs [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320919 (https://phabricator.wikimedia.org/T431089) (owner: 10Bartosz Wójtowicz) [11:23:00] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [11:23:11] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [11:23:15] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [11:23:27] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [11:24:08] (03PS5) 10Krinkle: mediawiki: Disable legacy `short_urls` on vhosts where it does not work [puppet] - 10https://gerrit.wikimedia.org/r/1290104 (https://phabricator.wikimedia.org/T107188) [11:24:14] (03CR) 10Ladsgroup: [C:03+2] mediawiki: Disable legacy `short_urls` on vhosts where it does not work [puppet] - 10https://gerrit.wikimedia.org/r/1290104 (https://phabricator.wikimedia.org/T107188) (owner: 10Krinkle) [11:24:17] (03CR) 10Ladsgroup: [V:03+2 C:03+2] mediawiki: Disable legacy `short_urls` on vhosts where it does not work [puppet] - 10https://gerrit.wikimedia.org/r/1290104 (https://phabricator.wikimedia.org/T107188) (owner: 10Krinkle) [11:25:12] jouncebot: nowandnext [11:25:13] For the next 0 hour(s) and 34 minute(s): Services – Citoid / Zotero (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260805T1100) [11:25:13] In 1 hour(s) and 34 minute(s): UTC afternoon backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260805T1300) [11:25:17] (03CR) 10Effie Mouzeli: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1321535 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [11:27:54] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [11:28:07] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [11:28:08] (03Merged) 10jenkins-bot: ml-services: declare ephemeral-storage for all ml-staging isvcs [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320919 (https://phabricator.wikimedia.org/T431089) (owner: 10Bartosz Wójtowicz) [11:28:45] (03CR) 10Effie Mouzeli: "As far as logistics go, LGTM!" [puppet] - 10https://gerrit.wikimedia.org/r/1320824 (https://phabricator.wikimedia.org/T433963) (owner: 10Majavah) [11:28:55] (03CR) 10Ladsgroup: [V:03+2 C:03+2] "I'm getting this:" [puppet] - 10https://gerrit.wikimedia.org/r/1290104 (https://phabricator.wikimedia.org/T107188) (owner: 10Krinkle) [11:29:42] (03PS1) 10Ladsgroup: Revert "mediawiki: Disable legacy `short_urls` on vhosts where it does not work" [puppet] - 10https://gerrit.wikimedia.org/r/1321536 [11:29:49] (03CR) 10Ladsgroup: [V:03+2 C:03+2] Revert "mediawiki: Disable legacy `short_urls` on vhosts where it does not work" [puppet] - 10https://gerrit.wikimedia.org/r/1321536 (owner: 10Ladsgroup) [11:30:27] (03CR) 10Effie Mouzeli: [C:03+2] hieradata: use PKI on mc2046 [puppet] - 10https://gerrit.wikimedia.org/r/1321535 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [11:34:51] !log jiji@cumin1003 START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie [11:35:05] 10ops-codfw, 06SRE, 06DC-Ops, 06ServiceOps: Rerack mc-gp2006 and mc2046 - https://phabricator.wikimedia.org/T432716#12186868 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by jiji@cumin1003 for host mc2046.codfw.wmnet with OS trixie [11:35:13] !log jiji@cumin1003 START - Cookbook sre.hosts.move-vlan for host mc2046 [11:35:13] !log jiji@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 [11:38:27] (03PS38) 10Federico Ceratto: mysql: update replication source [cookbooks] - 10https://gerrit.wikimedia.org/r/1238368 (https://phabricator.wikimedia.org/T373436) [11:38:40] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [11:42:12] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [11:43:47] (03PS39) 10Federico Ceratto: mysql: update replication source [cookbooks] - 10https://gerrit.wikimedia.org/r/1238368 (https://phabricator.wikimedia.org/T373436) [11:43:56] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [11:44:09] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [11:44:27] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [11:44:33] claime: I'm deploying and in theory new logs should be showing up in staging, but I'm only seeing the old logs when I test the update. ideas why that might be? (the work locally... famous last years). Not entirely sure I should finish deploying? [11:44:39] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [11:44:49] they* work locally... famous last words* [11:45:15] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [11:45:27] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [11:45:44] Mvolz: hmm let me see [11:47:09] (03PS40) 10Federico Ceratto: mysql: update replication source [cookbooks] - 10https://gerrit.wikimedia.org/r/1238368 (https://phabricator.wikimedia.org/T373436) [11:47:11] Mvolz: I have logs on the kubectl logs cli, do you mean in logstash? [11:47:16] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [11:47:29] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [11:47:38] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [11:47:41] yeah in logstash, and there should be two log entries per request. [11:47:51] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [11:47:56] Mvolz: can you send me the logstash link you're using? [11:47:57] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [11:48:09] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [11:48:56] (03PS41) 10Federico Ceratto: mysql: update replication source [cookbooks] - 10https://gerrit.wikimedia.org/r/1238368 (https://phabricator.wikimedia.org/T373436) [11:48:59] claime: https://logstash.wikimedia.org/goto/96ab42b189950d956839982a3ad4e3d2 [11:49:34] so I know it looks like it's working but we added a new info level log in https://gerrit.wikimedia.org/r/c/mediawiki/services/citoid/+/1313140 and those aren't showing up. [11:49:51] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [11:50:06] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [11:50:41] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [11:50:53] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [11:51:49] claime: hmmm, maybe it's these weird logs? It says use a different option key? https://logstash.wikimedia.org/app/discover#/doc/logstash-*/logstash-k8s-1-7.0.0-1-2026.08.05?id=seGl0Z8Br9cHI7Petiwo [11:52:12] Mvolz: I see these new logs in the cli so they're being emitted [11:52:28] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [11:52:40] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [11:52:53] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [11:53:02] Mvolz: And I see one in logstash too, second INFO log down 11:23:30.804 [11:53:04] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [11:53:18] Ah no my bad that's another one [11:53:39] (03PS2) 10Trueg: WDQSv2: Qlever index rebuild container [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320946 (https://phabricator.wikimedia.org/T432627) [11:53:45] mine was a red herring too. [11:53:57] !log jiji@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage [11:54:05] yeah they're a couple "warning" logs from startup if I had to guess [11:54:32] just did a few more requests, nothing. here's how to get what should be them: https://logstash.wikimedia.org/goto/6d19b8be17d0747a86c2eb9db5c05ab4 [11:54:33] 06SRE, 10Maps, 06Traffic: Possibility to allow Wikimedia Maps usage on all Wikibase Cloud instances - https://phabricator.wikimedia.org/T429191#12186892 (10Anton.Kokh) @MSantos There are currently around 800K+ pages using geo coordinates across 140 Wikibases. 90% of these pages are concentrated in 5 of those... [11:55:03] (excludes the outgoing request logs) [11:55:39] FIRING: CoreBGPDown: Core BGP session down between cr3-ulsfo and cr1-eqiad (198.35.26.150) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=ulsfo&var-device=cr3-ulsfo:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr1-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [11:57:10] Mvolz: they're falling into dlq (dead letter queue) https://logstash.wikimedia.org/app/discover#/doc/19d32430-85d8-11eb-8ab2-63c7f3b019fc/dlq-default-1-1.0.0-1-2026.08.05?id=J5HF0Z8BsYGHbYxISGCJ [11:57:23] !log jiji@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage [11:57:29] Mvolz: There's something logstash doesn't like with them [11:58:23] tnx, I'll avoid finishing the deploy to avoid polluting logstash.... [11:58:25] Mvolz: Maybe a reserved field or something, I would flag that with #wikimedia-observability [11:59:02] !log jayme@cumin1003 START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1156.eqiad.wmnet [11:59:07] !log jayme@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1156.eqiad.wmnet [11:59:28] it could be the event keyword, I was trying to gradually move to ecs with that but it might be unhappy with where it lives, I'll ask there, thanks! [11:59:41] !log jayme@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1156.eqiad.wmnet [12:00:39] RESOLVED: CoreBGPDown: Core BGP session down between cr3-ulsfo and cr1-eqiad (198.35.26.150) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=ulsfo&var-device=cr3-ulsfo:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr1-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [12:01:01] !log jayme@cumin1003 START - Cookbook sre.hosts.reimage for host wikikube-worker1156.eqiad.wmnet with OS trixie [12:01:41] !log jayme@cumin1003 START - Cookbook sre.hosts.move-vlan for host wikikube-worker1156 [12:02:36] !log bwojtowicz@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . [12:04:17] !log bwojtowicz@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . [12:04:44] jayme@cumin1003 renumber-node (PID 889077) is awaiting input [12:04:54] !log jayme@cumin1003 START - Cookbook sre.dns.netbox [12:06:20] !log bwojtowicz@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . [12:07:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [12:08:41] (03CR) 10Federico Ceratto: mysql: update replication source (032 comments) [cookbooks] - 10https://gerrit.wikimedia.org/r/1238368 (https://phabricator.wikimedia.org/T373436) (owner: 10Federico Ceratto) [12:09:16] !log bwojtowicz@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . [12:11:06] jayme@cumin1003 renumber-node (PID 889077) is awaiting input [12:14:19] (03CR) 10CWilliams: mysql: update replication source (031 comment) [cookbooks] - 10https://gerrit.wikimedia.org/r/1238368 (https://phabricator.wikimedia.org/T373436) (owner: 10Federico Ceratto) [12:14:46] (03CR) 10Federico Ceratto: "Few tweaks making the process more robust; I ran tests on the testbed using `--wait-lag-timeout 20` and `--delete-heartbeat`." [cookbooks] - 10https://gerrit.wikimedia.org/r/1238368 (https://phabricator.wikimedia.org/T373436) (owner: 10Federico Ceratto) [12:14:46] !log jiji@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie [12:15:02] 10ops-codfw, 06SRE, 06DC-Ops, 06ServiceOps: Rerack mc-gp2006 and mc2046 - https://phabricator.wikimedia.org/T432716#12186916 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by jiji@cumin1003 for host mc2046.codfw.wmnet with OS trixie completed: - mc2046 (**PASS**) - Downtimed on Icing... [12:17:12] !log bwojtowicz@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . [12:17:39] FIRING: CoreBGPDown: Core BGP session down between cr3-ulsfo and cr1-codfw (198.35.26.152) - group Confed_codfw - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=ulsfo&var-device=cr3-ulsfo:9804&var-bgp_group=Confed_codfw&var-bgp_neighbor=cr1-codfw - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [12:18:36] (03PS18) 10Slyngshede: varnish: Configure text for thumbnails [puppet] - 10https://gerrit.wikimedia.org/r/1306896 (https://phabricator.wikimedia.org/T427465) [12:19:37] !log jayme@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" [12:19:42] !log jayme@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003" [12:19:43] !log jayme@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [12:19:43] !log jayme@cumin1003 START - Cookbook sre.dns.wipe-cache wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [12:19:48] !log jayme@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [12:19:48] !log jayme@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1156 [12:20:14] (03CR) 10Federico Ceratto: mysql: update replication source (031 comment) [cookbooks] - 10https://gerrit.wikimedia.org/r/1238368 (https://phabricator.wikimedia.org/T373436) (owner: 10Federico Ceratto) [12:22:09] !log bwojtowicz@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . [12:22:31] !log jayme@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1156 [12:22:31] !log jayme@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1156 [12:22:39] RESOLVED: CoreBGPDown: Core BGP session down between cr3-ulsfo and cr1-codfw (198.35.26.152) - group Confed_codfw - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=ulsfo&var-device=cr3-ulsfo:9804&var-bgp_group=Confed_codfw&var-bgp_neighbor=cr1-codfw - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [12:24:21] !log update bgp confed settings in eqsin [12:24:23] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [12:26:23] !log bwojtowicz@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . [12:26:56] (03PS19) 10Slyngshede: varnish: Configure text for thumbnails [puppet] - 10https://gerrit.wikimedia.org/r/1306896 (https://phabricator.wikimedia.org/T427465) [12:28:32] !log bwojtowicz@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . [12:30:03] !log bwojtowicz@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . [12:31:25] !log bwojtowicz@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . [12:32:15] (03PS20) 10Slyngshede: varnish: Configure text for thumbnails [puppet] - 10https://gerrit.wikimedia.org/r/1306896 (https://phabricator.wikimedia.org/T427465) [12:32:49] !log bwojtowicz@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' . [12:35:27] PROBLEM - Host wikikube-worker1154 is DOWN: PING CRITICAL - Packet loss = 50%, RTA = 2802.43 ms [12:35:41] RECOVERY - Host wikikube-worker1154 is UP: PING OK - Packet loss = 0%, RTA = 0.22 ms [12:38:19] !log jayme@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage [12:40:21] (03CR) 10Marostegui: [C:03+1] "This looks good, but of course not yet as the master isn't switched. I will do it on Tuesday next week most likely." [puppet] - 10https://gerrit.wikimedia.org/r/1321468 (https://phabricator.wikimedia.org/T434043) (owner: 10Jcrespo) [12:42:59] !log jayme@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage [12:43:50] (03CR) 10Ayounsi: [C:03+1] sre.network.tls: add SR-Linux v26 support [cookbooks] - 10https://gerrit.wikimedia.org/r/1319814 (https://phabricator.wikimedia.org/T433105) (owner: 10Ayounsi) [12:43:56] (03PS1) 10CWilliams: sre.mysql: Minor cleanup [cookbooks] - 10https://gerrit.wikimedia.org/r/1321543 (https://phabricator.wikimedia.org/T433044) [12:50:28] (03PS1) 10Dpogorzelski: ml-serve: fallback for GPU usage and power metrics [puppet] - 10https://gerrit.wikimedia.org/r/1321545 [12:50:31] (03CR) 10CWilliams: "I have sent you a message re 1321543. It seems to make more sense to incorporate that first as it removes some unnecessary code from this " [cookbooks] - 10https://gerrit.wikimedia.org/r/1238368 (https://phabricator.wikimedia.org/T373436) (owner: 10Federico Ceratto) [12:52:44] (03PS2) 10Dpogorzelski: ml-serve: fallback for GPU usage and power metrics [puppet] - 10https://gerrit.wikimedia.org/r/1321545 [12:53:38] (03CR) 10CI reject: [V:04-1] ml-serve: fallback for GPU usage and power metrics [puppet] - 10https://gerrit.wikimedia.org/r/1321545 (owner: 10Dpogorzelski) [12:55:07] (03CR) 10Federico Ceratto: [C:03+1] "LGTM" [cookbooks] - 10https://gerrit.wikimedia.org/r/1321543 (https://phabricator.wikimedia.org/T433044) (owner: 10CWilliams) [12:55:28] (03CR) 10CWilliams: [C:03+2] sre.mysql: Minor cleanup [cookbooks] - 10https://gerrit.wikimedia.org/r/1321543 (https://phabricator.wikimedia.org/T433044) (owner: 10CWilliams) [12:59:08] (03Merged) 10jenkins-bot: sre.mysql: Minor cleanup [cookbooks] - 10https://gerrit.wikimedia.org/r/1321543 (https://phabricator.wikimedia.org/T433044) (owner: 10CWilliams) [13:00:04] Lucas_WMDE, urbanecm, and TheresNoTime: How many deployers does it take to do UTC afternoon backport window deploy? (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260805T1300). [13:00:05] No Gerrit patches in the queue for this window AFAICS. [13:01:56] !log jayme@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1156.eqiad.wmnet with OS trixie [13:02:25] RESOLVED: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:07:38] (03CR) 10Bking: [C:03+2] cirrussearch: Enable performance governor on master-eligible hosts [puppet] - 10https://gerrit.wikimedia.org/r/1321102 (https://phabricator.wikimedia.org/T324335) (owner: 10Bking) [13:14:54] (03PS1) 10Cathal Mooney: ulsfo: reduce OSPF cost on new transport circuits [homer/public] - 10https://gerrit.wikimedia.org/r/1321546 (https://phabricator.wikimedia.org/T424839) [13:15:40] (03PS1) 10Reedy: wmf-config: OATHAuth config changes [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1321547 (https://phabricator.wikimedia.org/T428103) [13:15:40] !log jayme@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1156.eqiad.wmnet [13:15:41] !log jayme@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1156.eqiad.wmnet [13:15:43] !log jayme@cumin1003 END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1156.eqiad.wmnet [13:15:59] (03Abandoned) 10Reedy: Set $wmgOATHAuthRequire2FAForAll = true for various private wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1299644 (https://phabricator.wikimedia.org/T428103) (owner: 10Reedy) [13:16:04] (03Abandoned) 10Reedy: Set $wmgOATHAuthRequire2FAForAll = true for all private wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1299645 (https://phabricator.wikimedia.org/T428103) (owner: 10Reedy) [13:16:44] (03PS21) 10Slyngshede: varnish: Configure text for thumbnails [puppet] - 10https://gerrit.wikimedia.org/r/1306896 (https://phabricator.wikimedia.org/T427465) [13:17:04] (03PS42) 10Federico Ceratto: mysql: update replication source [cookbooks] - 10https://gerrit.wikimedia.org/r/1238368 (https://phabricator.wikimedia.org/T373436) [13:21:21] (03CR) 10Ayounsi: [C:03+1] ulsfo: reduce OSPF cost on new transport circuits [homer/public] - 10https://gerrit.wikimedia.org/r/1321546 (https://phabricator.wikimedia.org/T424839) (owner: 10Cathal Mooney) [13:21:47] (03CR) 10Cathal Mooney: [C:03+2] ulsfo: reduce OSPF cost on new transport circuits [homer/public] - 10https://gerrit.wikimedia.org/r/1321546 (https://phabricator.wikimedia.org/T424839) (owner: 10Cathal Mooney) [13:21:55] !log jayme@cumin1003 START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1157.eqiad.wmnet [13:21:59] !log jayme@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1157.eqiad.wmnet [13:22:33] !log jayme@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1157.eqiad.wmnet [13:22:47] !log jayme@cumin1003 START - Cookbook sre.hosts.reimage for host wikikube-worker1157.eqiad.wmnet with OS trixie [13:23:15] !log jayme@cumin1003 START - Cookbook sre.hosts.move-vlan for host wikikube-worker1157 [13:23:15] RECOVERY - orchestrator resolve cache non-FQDNs on dborch1002 is OK: OK: all orchestrator resolve cache entries are FQDNs https://wikitech.wikimedia.org/wiki/Orchestrator [13:23:28] (03Merged) 10jenkins-bot: ulsfo: reduce OSPF cost on new transport circuits [homer/public] - 10https://gerrit.wikimedia.org/r/1321546 (https://phabricator.wikimedia.org/T424839) (owner: 10Cathal Mooney) [13:24:49] (03PS3) 10AOkoth: admin: add aramirezwmf to analytics-privatedata [puppet] - 10https://gerrit.wikimedia.org/r/1321519 (https://phabricator.wikimedia.org/T433999) [13:24:53] (03CR) 10Slyngshede: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1320824 (https://phabricator.wikimedia.org/T433963) (owner: 10Majavah) [13:25:53] jouncebot: nownandnext [13:25:54] (03CR) 10Slyngshede: [C:03+1] "Aaah, sorry. Yes, that's fine then. Memory backend is probably fine, but using Redis doesn't hurt." [puppet] - 10https://gerrit.wikimedia.org/r/1320824 (https://phabricator.wikimedia.org/T433963) (owner: 10Majavah) [13:25:57] jouncebot: nownandnext [13:26:01] ffs [13:26:04] jouncebot: nowandnext [13:26:04] For the next 0 hour(s) and 33 minute(s): UTC afternoon backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260805T1300) [13:26:04] In 0 hour(s) and 33 minute(s): Wikifunctions Services UTC Afternoon (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260805T1400) [13:26:18] jayme@cumin1003 renumber-node (PID 951977) is awaiting input [13:26:37] (03CR) 10Jcrespo: [C:04-1] "Blocked on master switch." [puppet] - 10https://gerrit.wikimedia.org/r/1321468 (https://phabricator.wikimedia.org/T434043) (owner: 10Jcrespo) [13:27:34] (03CR) 10Reedy: [C:03+2] wmf-config: OATHAuth config changes [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1321547 (https://phabricator.wikimedia.org/T428103) (owner: 10Reedy) [13:29:58] (03Merged) 10jenkins-bot: wmf-config: OATHAuth config changes [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1321547 (https://phabricator.wikimedia.org/T428103) (owner: 10Reedy) [13:32:14] !log reedy@deploy1003 Started scap sync-world: Backport for [[gerrit:1321547|wmf-config: OATHAuth config changes (T428103)]] [13:32:18] (03CR) 10Federico Ceratto: "Ack, I rebased this CR again." [cookbooks] - 10https://gerrit.wikimedia.org/r/1238368 (https://phabricator.wikimedia.org/T373436) (owner: 10Federico Ceratto) [13:32:19] T428103: Enforce 2FA for all users on private wikis in WMF production - https://phabricator.wikimedia.org/T428103 [13:34:21] !log reedy@deploy1003 reedy: Backport for [[gerrit:1321547|wmf-config: OATHAuth config changes (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [13:34:53] (03PS5) 10CDanis: airflow: main: add cloudflare_radar connection [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320964 (https://phabricator.wikimedia.org/T433982) [13:35:03] !log reedy@deploy1003 reedy: Continuing with deployment [13:35:26] !log jayme@cumin1003 START - Cookbook sre.dns.netbox [13:35:44] (03CR) 10Krinkle: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1290104 (https://phabricator.wikimedia.org/T107188) (owner: 10Krinkle) [13:35:50] (03CR) 10JHathaway: memcached: reload_certs on certificate renewal (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1320940 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [13:39:14] !log reedy@deploy1003 Finished scap sync-world: Backport for [[gerrit:1321547|wmf-config: OATHAuth config changes (T428103)]] (duration: 07m 00s) [13:39:18] T428103: Enforce 2FA for all users on private wikis in WMF production - https://phabricator.wikimedia.org/T428103 [13:39:58] !log jayme@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" [13:40:02] !log jayme@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003" [13:40:02] !log jayme@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [13:40:03] !log jayme@cumin1003 START - Cookbook sre.dns.wipe-cache wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [13:40:07] !log jayme@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [13:40:07] !log jayme@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1157 [13:41:55] (03CR) 10Scott French: httpbb: Release v0.0.6 (031 comment) [software/httpbb] - 10https://gerrit.wikimedia.org/r/1321516 (https://phabricator.wikimedia.org/T434052) (owner: 10Blake) [13:42:57] !log jayme@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1157 [13:42:57] !log jayme@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1157 [13:46:04] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host search-loader1002.eqiad.wmnet with OS trixie [13:46:08] (03PS1) 10Brouberol: dse-k8s-eqiad/airflow-wikidata: allow the creation of extremely large pods [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321554 (https://phabricator.wikimedia.org/T422179) [13:47:22] (03PS3) 10TheDJ: Supported animated webp thumbnailing [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1270150 (https://phabricator.wikimedia.org/T290345) [13:47:24] (03PS2) 10Brouberol: dse-k8s-eqiad/airflow-wikidata: allow the creation of extremely large pods [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321554 (https://phabricator.wikimedia.org/T422179) [13:47:28] (03PS22) 10Slyngshede: varnish: Configure text for thumbnails [puppet] - 10https://gerrit.wikimedia.org/r/1306896 (https://phabricator.wikimedia.org/T427465) [13:47:44] (03PS3) 10Brouberol: dse-k8s-eqiad/airflow-wikidata: allow the creation of extremely large pods [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321554 (https://phabricator.wikimedia.org/T422179) [13:48:19] (03CR) 10Trueg: [C:03+1] "ty" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321554 (https://phabricator.wikimedia.org/T422179) (owner: 10Brouberol) [13:49:04] (03PS3) 10Dpogorzelski: ml-serve: fallback for GPU usage and power metrics [puppet] - 10https://gerrit.wikimedia.org/r/1321545 [13:49:52] (03CR) 10CI reject: [V:04-1] ml-serve: fallback for GPU usage and power metrics [puppet] - 10https://gerrit.wikimedia.org/r/1321545 (owner: 10Dpogorzelski) [13:53:57] FIRING: [2x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [13:54:00] !log brouberol@cumin1003 START - Cookbook sre.hosts.reimage for host an-test-worker1003.eqiad.wmnet with OS bookworm [13:57:10] (03PS2) 10Diegodlh: wmf-config: Set JSON content model for Web2Cit configuration pages [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1314816 (https://phabricator.wikimedia.org/T305571) [13:57:21] (03CR) 10Brouberol: [C:03+2] dse-k8s-eqiad/airflow-wikidata: allow the creation of extremely large pods [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321554 (https://phabricator.wikimedia.org/T422179) (owner: 10Brouberol) [13:57:53] !log bking@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage [13:58:28] !log jayme@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage [13:59:01] (03PS1) 10Jforrester: wikifunctions: Upgrade evaluators from 2026-07-27-212343 to 2026-08-04-215640 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321556 (https://phabricator.wikimedia.org/T373110) [13:59:02] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. [13:59:04] (03CR) 10Scott French: [C:03+2] hieradata: Use component/zookeeper34 on all configcluster hosts [puppet] - 10https://gerrit.wikimedia.org/r/1321106 (https://phabricator.wikimedia.org/T428495) (owner: 10Scott French) [13:59:23] (03PS1) 10Jforrester: wikifunctions: Upgrade orchestrator from 2026-07-29-180129 to 2026-08-04-203437 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321557 (https://phabricator.wikimedia.org/T400517) [14:00:05] Deploy window Wikifunctions Services UTC Afternoon (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260805T1400) [14:00:06] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. [14:01:11] !log bking@cumin2003 conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad [14:01:12] !log bking@cumin2003 conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=eqiad [14:01:12] !log bking@cumin2003 conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad [14:02:00] PROBLEM - Host wikikube-worker1157 is DOWN: PING CRITICAL - Packet loss = 100% [14:03:00] (03PS4) 10Dpogorzelski: ml-serve: fallback for GPU usage and power metrics [puppet] - 10https://gerrit.wikimedia.org/r/1321545 [14:03:14] (03CR) 10Jforrester: [C:03+2] wikifunctions: Upgrade evaluators from 2026-07-27-212343 to 2026-08-04-215640 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321556 (https://phabricator.wikimedia.org/T373110) (owner: 10Jforrester) [14:03:32] (03PS1) 10Brouberol: dse-k8s-eqiad/airflow-wikidata: increase namespace quotas [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321562 (https://phabricator.wikimedia.org/T422179) [14:03:34] (03CR) 10CWilliams: [C:03+1] "LGTM - I have not retested\, but the state of the test cluster and the log entries look as expected." [cookbooks] - 10https://gerrit.wikimedia.org/r/1238368 (https://phabricator.wikimedia.org/T373436) (owner: 10Federico Ceratto) [14:03:49] (03CR) 10CI reject: [V:04-1] ml-serve: fallback for GPU usage and power metrics [puppet] - 10https://gerrit.wikimedia.org/r/1321545 (owner: 10Dpogorzelski) [14:04:10] (03CR) 10Lerickson: [C:03+1] dse-k8s-eqiad/airflow-wikidata: increase namespace quotas [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321562 (https://phabricator.wikimedia.org/T422179) (owner: 10Brouberol) [14:04:32] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage [14:04:39] FIRING: [2x] TransitBGPDown: Transit BGP session down between cr2-esams and Hurricane Electric (193.239.116.14) - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DTransitBGPDown [14:04:49] !log bking@cumin2003 START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - T324335 [14:04:55] T324335: Index the OpenSearch server application logs in the central observability cluster, using ECS format - https://phabricator.wikimedia.org/T324335 [14:05:51] (03CR) 10TheDJ: [C:04-1] "This animates webp to a thumbnailed webp." [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1270150 (https://phabricator.wikimedia.org/T290345) (owner: 10TheDJ) [14:05:54] (03Merged) 10jenkins-bot: wikifunctions: Upgrade evaluators from 2026-07-27-212343 to 2026-08-04-215640 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321556 (https://phabricator.wikimedia.org/T373110) (owner: 10Jforrester) [14:06:37] (03CR) 10Federico Ceratto: [C:03+2] mysql: update replication source [cookbooks] - 10https://gerrit.wikimedia.org/r/1238368 (https://phabricator.wikimedia.org/T373436) (owner: 10Federico Ceratto) [14:07:03] RECOVERY - Host wikikube-worker1157 is UP: PING OK - Packet loss = 0%, RTA = 0.39 ms [14:08:01] !log jforrester@deploy1003 helmfile [staging] START helmfile.d/services/wikifunctions: apply [14:08:06] !log jayme@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage [14:08:53] !log jforrester@deploy1003 helmfile [staging] DONE helmfile.d/services/wikifunctions: apply [14:09:13] !log jforrester@deploy1003 helmfile [codfw] START helmfile.d/services/wikifunctions: apply [14:09:39] RESOLVED: [2x] TransitBGPDown: Transit BGP session down between cr2-esams and Hurricane Electric (193.239.116.14) - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DTransitBGPDown [14:09:44] !log jforrester@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply [14:09:50] !log jforrester@deploy1003 helmfile [eqiad] START helmfile.d/services/wikifunctions: apply [14:10:08] (03Merged) 10jenkins-bot: mysql: update replication source [cookbooks] - 10https://gerrit.wikimedia.org/r/1238368 (https://phabricator.wikimedia.org/T373436) (owner: 10Federico Ceratto) [14:10:18] !log jforrester@deploy1003 helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply [14:10:35] (03CR) 10Jforrester: [C:03+2] wikifunctions: Upgrade orchestrator from 2026-07-29-180129 to 2026-08-04-203437 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321557 (https://phabricator.wikimedia.org/T400517) (owner: 10Jforrester) [14:10:46] (03PS5) 10Dpogorzelski: ml-serve: fallback for GPU usage and power metrics [puppet] - 10https://gerrit.wikimedia.org/r/1321545 [14:11:08] !log brouberol@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage [14:13:10] (03Merged) 10jenkins-bot: wikifunctions: Upgrade orchestrator from 2026-07-29-180129 to 2026-08-04-203437 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321557 (https://phabricator.wikimedia.org/T400517) (owner: 10Jforrester) [14:13:32] !log jiji@cumin1003 START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie [14:13:42] !log jiji@cumin1003 START - Cookbook sre.hosts.move-vlan for host mc2046 [14:13:42] !log jiji@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 [14:13:45] 10ops-codfw, 06SRE, 06DC-Ops, 06ServiceOps: Rerack mc-gp2006 and mc2046 - https://phabricator.wikimedia.org/T432716#12187503 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by jiji@cumin1003 for host mc2046.codfw.wmnet with OS trixie [14:14:36] !log jforrester@deploy1003 helmfile [staging] START helmfile.d/services/wikifunctions: apply [14:14:44] (03PS1) 10Effie Mouzeli: mcrouter_wancache: replace mc2046's IP [puppet] - 10https://gerrit.wikimedia.org/r/1321564 (https://phabricator.wikimedia.org/T432716) [14:14:56] !log jforrester@deploy1003 helmfile [staging] DONE helmfile.d/services/wikifunctions: apply [14:14:59] (03CR) 10Brouberol: [C:03+2] dse-k8s-eqiad/airflow-wikidata: increase namespace quotas [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321562 (https://phabricator.wikimedia.org/T422179) (owner: 10Brouberol) [14:15:00] !log brouberol@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage [14:16:09] (03CR) 10TheDJ: "thinking about it again, that also shouldn't hurt it. It will simply be unused until core is (potentially switched over to use webp thumbn" [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1270150 (https://phabricator.wikimedia.org/T290345) (owner: 10TheDJ) [14:16:34] (03CR) 10Effie Mouzeli: [C:03+2] mcrouter_wancache: replace mc2046's IP [puppet] - 10https://gerrit.wikimedia.org/r/1321564 (https://phabricator.wikimedia.org/T432716) (owner: 10Effie Mouzeli) [14:18:01] FIRING: [2x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [14:18:17] !log jforrester@deploy1003 helmfile [codfw] START helmfile.d/services/wikifunctions: apply [14:18:58] !log jforrester@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply [14:19:14] !log jforrester@deploy1003 helmfile [eqiad] START helmfile.d/services/wikifunctions: apply [14:19:55] !log jforrester@deploy1003 helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply [14:22:07] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader1002.eqiad.wmnet with OS trixie [14:22:27] (03PS7) 10RLazarus: wikifunctions: Add mesh.restricted_listeners port to orchestrator [deployment-charts] - 10https://gerrit.wikimedia.org/r/1296071 (https://phabricator.wikimedia.org/T427863) [14:22:29] (03CR) 10Jforrester: [C:03+2] wikifunctions: Add mesh.restricted_listeners port to orchestrator [deployment-charts] - 10https://gerrit.wikimedia.org/r/1296071 (https://phabricator.wikimedia.org/T427863) (owner: 10RLazarus) [14:24:36] (03CR) 10Ssingh: [C:03+1] Temporarily point all etcd client SRV records to codfw [dns] - 10https://gerrit.wikimedia.org/r/1321090 (https://phabricator.wikimedia.org/T428495) (owner: 10Scott French) [14:24:44] (03CR) 10Ssingh: [C:03+1] hieradata: Temporarily point eqiad PyBals at codfw etcd [puppet] - 10https://gerrit.wikimedia.org/r/1321089 (https://phabricator.wikimedia.org/T428495) (owner: 10Scott French) [14:24:49] (03Merged) 10jenkins-bot: wikifunctions: Add mesh.restricted_listeners port to orchestrator [deployment-charts] - 10https://gerrit.wikimedia.org/r/1296071 (https://phabricator.wikimedia.org/T427863) (owner: 10RLazarus) [14:26:05] !log jforrester@deploy1003 helmfile [staging] START helmfile.d/services/wikifunctions: apply [14:26:40] !log jforrester@deploy1003 helmfile [staging] DONE helmfile.d/services/wikifunctions: apply [14:26:41] !log jiji@deploy1003 helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply [14:26:48] !log jiji@deploy1003 helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply [14:26:55] !log jforrester@deploy1003 helmfile [codfw] START helmfile.d/services/wikifunctions: apply [14:26:57] !log jiji@deploy1003 helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply [14:27:03] !log jiji@deploy1003 helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply [14:27:27] (03CR) 10Federico Ceratto: "This is pending final review." [cookbooks] - 10https://gerrit.wikimedia.org/r/1291952 (https://phabricator.wikimedia.org/T426613) (owner: 10Federico Ceratto) [14:27:38] !log jforrester@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply [14:27:43] !log jforrester@deploy1003 helmfile [eqiad] START helmfile.d/services/wikifunctions: apply [14:27:46] !log jayme@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1157.eqiad.wmnet with OS trixie [14:28:16] !log jforrester@deploy1003 helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply [14:29:15] (03PS7) 10RLazarus: function-evaluator: Add outgoing Envoy config and egress policy for callbacks [deployment-charts] - 10https://gerrit.wikimedia.org/r/1296072 (https://phabricator.wikimedia.org/T427863) [14:29:42] (03CR) 10Jforrester: [C:03+2] function-evaluator: Add outgoing Envoy config and egress policy for callbacks [deployment-charts] - 10https://gerrit.wikimedia.org/r/1296072 (https://phabricator.wikimedia.org/T427863) (owner: 10RLazarus) [14:30:05] Deploy window Wikifunctions Services UTC Afternoon (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260805T1400) [14:30:05] Deploy window Test Kitchen Experiment Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260805T1430) [14:30:45] (03Abandoned) 10Jforrester: Bump up the maximum concurrent calls (32 -> 512 in orchestrator, 100 -> 512 in evaluators). [deployment-charts] - 10https://gerrit.wikimedia.org/r/1318191 (https://phabricator.wikimedia.org/T433280) (owner: 10Cory Massaro) [14:31:46] jayme@cumin1003 renumber-node (PID 951977) is awaiting input [14:32:00] !log jiji@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage [14:32:44] (03Merged) 10jenkins-bot: function-evaluator: Add outgoing Envoy config and egress policy for callbacks [deployment-charts] - 10https://gerrit.wikimedia.org/r/1296072 (https://phabricator.wikimedia.org/T427863) (owner: 10RLazarus) [14:34:18] !log jforrester@deploy1003 helmfile [staging] START helmfile.d/services/wikifunctions: apply [14:34:27] !log jforrester@deploy1003 helmfile [staging] DONE helmfile.d/services/wikifunctions: apply [14:34:38] !log jforrester@deploy1003 helmfile [codfw] START helmfile.d/services/wikifunctions: apply [14:34:44] !log jforrester@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply [14:34:58] (03PS1) 10Majavah: hieradata: Update striker-toolsbeta to 2026-08-04-183239-production [puppet] - 10https://gerrit.wikimedia.org/r/1321567 [14:34:58] (03PS1) 10Majavah: hieradata: Update production Striker to 2026-08-04-183239-production [puppet] - 10https://gerrit.wikimedia.org/r/1321568 [14:35:27] (03CR) 10Cathal Mooney: [C:03+2] sre.network.tls: add SR-Linux v26 support [cookbooks] - 10https://gerrit.wikimedia.org/r/1319814 (https://phabricator.wikimedia.org/T433105) (owner: 10Ayounsi) [14:35:41] !log jforrester@deploy1003 helmfile [staging] START helmfile.d/services/wikifunctions: apply [14:35:45] !log jforrester@deploy1003 helmfile [staging] DONE helmfile.d/services/wikifunctions: apply [14:36:05] (03CR) 10Majavah: [C:03+2] hieradata: Update striker-toolsbeta to 2026-08-04-183239-production [puppet] - 10https://gerrit.wikimedia.org/r/1321567 (owner: 10Majavah) [14:36:26] (03CR) 10Scott French: "Thanks for the review!" [dns] - 10https://gerrit.wikimedia.org/r/1321090 (https://phabricator.wikimedia.org/T428495) (owner: 10Scott French) [14:37:10] (03CR) 10Scott French: [C:03+2] Temporarily point all etcd client SRV records to codfw [dns] - 10https://gerrit.wikimedia.org/r/1321090 (https://phabricator.wikimedia.org/T428495) (owner: 10Scott French) [14:37:27] !log jforrester@deploy1003 helmfile [staging] START helmfile.d/services/wikifunctions: apply [14:37:31] !log jforrester@deploy1003 helmfile [staging] DONE helmfile.d/services/wikifunctions: apply [14:37:39] (03PS4) 10Jforrester: wikifunctions: Double the number of evaluators from 2 to 4 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1271942 (https://phabricator.wikimedia.org/T419933) [14:37:41] !log swfrench@dns1004 START - running authdns-update [14:37:42] (03CR) 10Jforrester: [C:03+2] wikifunctions: Double the number of evaluators from 2 to 4 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1271942 (https://phabricator.wikimedia.org/T419933) (owner: 10Jforrester) [14:38:35] !log jiji@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage [14:39:25] (03Merged) 10jenkins-bot: sre.network.tls: add SR-Linux v26 support [cookbooks] - 10https://gerrit.wikimedia.org/r/1319814 (https://phabricator.wikimedia.org/T433105) (owner: 10Ayounsi) [14:39:32] (03PS1) 10Jforrester: wikifunctions: Rev the evaluator chart for I24c750f [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321570 (https://phabricator.wikimedia.org/T427863) [14:39:40] !log swfrench@dns1004 END - running authdns-update [14:39:41] (03CR) 10Jforrester: [C:03+2] wikifunctions: Rev the evaluator chart for I24c750f [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321570 (https://phabricator.wikimedia.org/T427863) (owner: 10Jforrester) [14:39:45] !log authdns update to direct eqiad-associated etcd clients to codfw - T428495 [14:39:51] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:39:52] T428495: Migrate conf* hosts away from bullseye - https://phabricator.wikimedia.org/T428495 [14:40:14] (03Merged) 10jenkins-bot: wikifunctions: Double the number of evaluators from 2 to 4 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1271942 (https://phabricator.wikimedia.org/T419933) (owner: 10Jforrester) [14:41:42] !log brouberol@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1003.eqiad.wmnet with OS bookworm [14:41:49] !log jforrester@deploy1003 helmfile [staging] START helmfile.d/services/wikifunctions: apply [14:41:52] jayme@cumin1003 renumber-node (PID 951977) is awaiting input [14:41:52] !log jforrester@deploy1003 helmfile [staging] DONE helmfile.d/services/wikifunctions: apply [14:41:58] !log jforrester@deploy1003 helmfile [codfw] START helmfile.d/services/wikifunctions: apply [14:42:16] !log jforrester@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply [14:42:21] !log jforrester@deploy1003 helmfile [eqiad] START helmfile.d/services/wikifunctions: apply [14:42:22] (03Merged) 10jenkins-bot: wikifunctions: Rev the evaluator chart for I24c750f [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321570 (https://phabricator.wikimedia.org/T427863) (owner: 10Jforrester) [14:42:38] !log jforrester@deploy1003 helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply [14:42:46] (03CR) 10Majavah: [C:03+2] hieradata: Update production Striker to 2026-08-04-183239-production [puppet] - 10https://gerrit.wikimedia.org/r/1321568 (owner: 10Majavah) [14:43:27] !log jayme@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1157.eqiad.wmnet [14:43:29] !log jayme@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1157.eqiad.wmnet [14:43:31] !log jayme@cumin1003 END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1157.eqiad.wmnet [14:43:38] !log jforrester@deploy1003 helmfile [staging] START helmfile.d/services/wikifunctions: apply [14:43:41] !log jayme@cumin1003 START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1159.eqiad.wmnet [14:43:45] !log jforrester@deploy1003 helmfile [staging] DONE helmfile.d/services/wikifunctions: apply [14:43:46] !log jayme@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1159.eqiad.wmnet [14:43:54] !log jforrester@deploy1003 helmfile [staging] START helmfile.d/services/wikifunctions: apply [14:44:19] !log jayme@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1159.eqiad.wmnet [14:44:55] !log jforrester@deploy1003 helmfile [staging] DONE helmfile.d/services/wikifunctions: apply [14:45:14] !log jforrester@deploy1003 helmfile [codfw] START helmfile.d/services/wikifunctions: apply [14:46:48] !log jayme@cumin1003 START - Cookbook sre.hosts.reimage for host wikikube-worker1159.eqiad.wmnet with OS trixie [14:46:49] !log jforrester@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply [14:46:54] !log jforrester@deploy1003 helmfile [eqiad] START helmfile.d/services/wikifunctions: apply [14:47:16] !log jayme@cumin1003 START - Cookbook sre.hosts.move-vlan for host wikikube-worker1159 [14:47:53] !log begin rolling restart of confd in drmrs, eqiad, esams, magru - T428495 [14:47:57] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:47:57] T428495: Migrate conf* hosts away from bullseye - https://phabricator.wikimedia.org/T428495 [14:48:14] !log jayme@cumin1003 START - Cookbook sre.dns.netbox [14:48:16] (03PS1) 10Cathal Mooney: Set standard metric for new easms and drmrs transports to codfw [homer/public] - 10https://gerrit.wikimedia.org/r/1321572 (https://phabricator.wikimedia.org/T424839) [14:49:12] (03PS1) 10JMeybohm: coredns: Rename coredns to coredns1.11 to support multiple versions [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1321573 (https://phabricator.wikimedia.org/T433590) [14:49:14] (03PS1) 10JMeybohm: Add coredns 1.12 [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1321574 (https://phabricator.wikimedia.org/T433590) [14:49:16] !log jforrester@deploy1003 helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply [14:52:33] !log jayme@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" [14:52:38] !log jayme@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003" [14:52:38] !log jayme@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [14:52:38] !log jayme@cumin1003 START - Cookbook sre.dns.wipe-cache wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [14:52:42] !log jayme@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [14:52:43] !log jayme@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1159 [14:54:18] !log restarted navtiming on webperf1003 - T428495 [14:54:21] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:54:22] T428495: Migrate conf* hosts away from bullseye - https://phabricator.wikimedia.org/T428495 [14:55:38] !log jiji@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie [14:55:50] 10ops-codfw, 06SRE, 06DC-Ops, 06ServiceOps: Rerack mc-gp2006 and mc2046 - https://phabricator.wikimedia.org/T432716#12187787 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by jiji@cumin1003 for host mc2046.codfw.wmnet with OS trixie completed: - mc2046 (**PASS**) - Downtimed on Icing... [14:59:24] (03CR) 10AOkoth: [C:03+2] admin: add aramirezwmf to analytics-privatedata [puppet] - 10https://gerrit.wikimedia.org/r/1321519 (https://phabricator.wikimedia.org/T433999) (owner: 10AOkoth) [15:00:07] !log begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - T428495 [15:00:11] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [15:00:12] T428495: Migrate conf* hosts away from bullseye - https://phabricator.wikimedia.org/T428495 [15:05:39] 10ops-codfw, 06SRE, 06DC-Ops, 06ServiceOps: Rerack mc-gp2006 and mc2046 - https://phabricator.wikimedia.org/T432716#12187841 (10jijiki) 05In progress→03Resolved a:03jijiki Thank you @Jhancock.wm and @cmooney for doing all the magic stuff, server is serving :) [15:10:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.42% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [15:10:43] !log jayme@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1159 [15:10:44] !log jayme@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1159 [15:18:15] (03PS2) 10Trueg: WDQSv2: Enable persistent updates on Qlever [deployment-charts] - 10https://gerrit.wikimedia.org/r/1314712 (https://phabricator.wikimedia.org/T432440) [15:19:16] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - search-https_9243: Servers cirrussearch1077.eqiad.wmnet, cirrussearch1103.eqiad.wmnet, cirrussearch1115.eqiad.wmnet, cirrussearch1090.eqiad.wmnet, cirrussearch1092.eqiad.wmnet, cirrussearch1110.eqiad.wmnet, cirrussearch1120.eqiad.wmnet, cirrussearch1085.eqiad.wmnet, cirrussearch1087.eqiad.wmnet, cirrussearch1079.eqiad.wmnet, cirrussearch1121.eqiad.wm [15:19:16] russearch1123.eqiad.wmnet, cirrussearch1113.eqiad.wmnet, cirrussearch1101.eqiad.wmnet, cirrussearch1108.eqiad.wmnet, cirrussearch1124.eqiad.wmnet, cirrussearch1088.eqiad.wmnet, cirrussearch1070.eqiad.wmnet, cirrussearch1086.eqiad.wmnet, cirrussearch1107.eqiad.wmnet, cirrussearch1091.eqiad.wmnet, cirrussearch1080.eqiad.wmnet, cirrussearch1089.eqiad.wmnet, cirrussearch1083.eqiad.wmnet, cirrussearch1112.eqiad.wmnet, cirrussearch1096.eqiad.wm [15:19:16] russearch1114.eqiad.wmnet, cirrussearch1116.eqiad.wmnet are marked down but pooled: search_9200: Servers cirrussearch1111.eqiad.wmnet, cirrussearch1117.eqiad.wmnet, cirrussearch1109.eqi https://wikitech.wikimedia.org/wiki/PyBal [15:19:26] PROBLEM - OpenSearch health check for shards on 9200 on cirrussearch1073 is CRITICAL: CRITICAL - elasticsearch http://localhost:9200/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9200): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:19:26] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1081 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:19:28] PROBLEM - OpenSearch health check for shards on 9200 on cirrussearch1094 is CRITICAL: CRITICAL - elasticsearch http://localhost:9200/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9200): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:19:28] PROBLEM - OpenSearch health check for shards on 9200 on cirrussearch1089 is CRITICAL: CRITICAL - elasticsearch http://localhost:9200/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9200): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:19:28] PROBLEM - OpenSearch health check for shards on 9200 on cirrussearch1095 is CRITICAL: CRITICAL - elasticsearch http://localhost:9200/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9200): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:19:28] PROBLEM - OpenSearch health check for shards on 9200 on cirrussearch1092 is CRITICAL: CRITICAL - elasticsearch http://localhost:9200/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9200): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:19:28] PROBLEM - OpenSearch health check for shards on 9200 on cirrussearch1077 is CRITICAL: CRITICAL - elasticsearch http://localhost:9200/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9200): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:19:28] PROBLEM - OpenSearch health check for shards on 9200 on cirrussearch1079 is CRITICAL: CRITICAL - elasticsearch http://localhost:9200/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9200): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:19:29] PROBLEM - OpenSearch health check for shards on 9200 on cirrussearch1080 is CRITICAL: CRITICAL - elasticsearch http://localhost:9200/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9200): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:19:29] PROBLEM - OpenSearch health check for shards on 9200 on cirrussearch1082 is CRITICAL: CRITICAL - elasticsearch http://localhost:9200/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9200): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:20:54] (03CR) 10Lerickson: [C:03+2] WDQSv2: Enable persistent updates on Qlever [deployment-charts] - 10https://gerrit.wikimedia.org/r/1314712 (https://phabricator.wikimedia.org/T432440) (owner: 10Trueg) [15:21:46] PROBLEM - Check unit status of push_cross_cluster_settings_9200 on cirrussearch1094 is CRITICAL: CRITICAL: Status of the systemd unit push_cross_cluster_settings_9200 https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [15:21:46] PROBLEM - Check unit status of push_cross_cluster_settings_9600 on cirrussearch1095 is CRITICAL: CRITICAL: Status of the systemd unit push_cross_cluster_settings_9600 https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [15:21:46] PROBLEM - Check unit status of push_cross_cluster_settings_9400 on cirrussearch1093 is CRITICAL: CRITICAL: Status of the systemd unit push_cross_cluster_settings_9400 https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [15:23:18] (03Merged) 10jenkins-bot: WDQSv2: Enable persistent updates on Qlever [deployment-charts] - 10https://gerrit.wikimedia.org/r/1314712 (https://phabricator.wikimedia.org/T432440) (owner: 10Trueg) [15:23:40] (03CR) 10CDobbins: [C:03+2] hieradata: Temporarily point eqiad PyBals at codfw etcd [puppet] - 10https://gerrit.wikimedia.org/r/1321089 (https://phabricator.wikimedia.org/T428495) (owner: 10Scott French) [15:23:57] FIRING: [2x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [15:25:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [15:26:29] (03PS1) 10Brouberol: airflow-dev: ensure a dev instance cannot send an email [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321580 (https://phabricator.wikimedia.org/T416596) [15:27:23] !log jayme@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage [15:27:27] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to analytics-private-users for aramirezwmf - https://phabricator.wikimedia.org/T433999#12187911 (10Arnoldokoth) I believe this should be good to go. @ARamirez_WMF Could you kindly test and confirm? [15:27:49] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to analytics-private-users for aramirezwmf - https://phabricator.wikimedia.org/T433999#12187912 (10Arnoldokoth) a:03ARamirez_WMF [15:28:19] 06SRE, 10SRE-swift-storage, 10MediaWiki-extensions-Score: Add cache key information to metadata json - https://phabricator.wikimedia.org/T257093#12187914 (10MatthewVernon) @TheDJ sorry, what do you mean by "this", in this context, please? [I don't think there's an action for swift-admin here, but I wanted to... [15:29:16] (03PS1) 10AOkoth: mariadb: add grants for phab1005 [puppet] - 10https://gerrit.wikimedia.org/r/1321581 (https://phabricator.wikimedia.org/T377889) [15:30:45] FIRING: CirrusStreamingUpdaterRateTooLow: CirrusSearch update rate from flink-app-consumer-search is critically low - https://wikitech.wikimedia.org/wiki/Search#Streaming_Updater - https://grafana.wikimedia.org/d/jKqki4MSk/cirrus-streaming-updater - https://alerts.wikimedia.org/?q=alertname%3DCirrusStreamingUpdaterRateTooLow [15:31:23] (03PS2) 10AOkoth: mariadb: add grants for phab1005 [puppet] - 10https://gerrit.wikimedia.org/r/1321581 (https://phabricator.wikimedia.org/T377889) [15:31:45] RECOVERY - Check unit status of push_cross_cluster_settings_9200 on cirrussearch1094 is OK: OK: Status of the systemd unit push_cross_cluster_settings_9200 https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [15:31:45] RECOVERY - Check unit status of push_cross_cluster_settings_9400 on cirrussearch1093 is OK: OK: Status of the systemd unit push_cross_cluster_settings_9400 https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [15:31:45] RECOVERY - Check unit status of push_cross_cluster_settings_9600 on cirrussearch1095 is OK: OK: Status of the systemd unit push_cross_cluster_settings_9600 https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [15:32:05] (03CR) 10Bartosz Dziewoński: [C:03+1] wmf-config: Set JSON content model for Web2Cit configuration pages [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1314816 (https://phabricator.wikimedia.org/T305571) (owner: 10Diegodlh) [15:32:12] !log bking@cumin2003 END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - T324335 [15:32:17] T324335: Index the OpenSearch server application logs in the central observability cluster, using ECS format - https://phabricator.wikimedia.org/T324335 [15:32:27] RECOVERY - OpenSearch health check for shards on 9200 on cirrussearch1100 is OK: OK - elasticsearch status production-search-eqiad: cluster_name: production-search-eqiad, status: red, timed_out: False, number_of_nodes: 4, number_of_data_nodes: 4, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 0, active_shards: 0, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, delayed_unassigned_shard [15:32:27] mber_of_pending_tasks: 64, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 6410, active_shards_percent_as_number: NaN https://wikitech.wikimedia.org/wiki/Search%23Administration [15:32:29] RECOVERY - OpenSearch health check for shards on 9200 on cirrussearch1081 is OK: OK - elasticsearch status production-search-eqiad: cluster_name: production-search-eqiad, status: red, timed_out: False, number_of_nodes: 4, number_of_data_nodes: 4, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 0, active_shards: 0, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, delayed_unassigned_shard [15:32:29] mber_of_pending_tasks: 64, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 6411, active_shards_percent_as_number: NaN https://wikitech.wikimedia.org/wiki/Search%23Administration [15:32:29] RECOVERY - OpenSearch health check for shards on 9200 on cirrussearch1094 is OK: OK - elasticsearch status production-search-eqiad: cluster_name: production-search-eqiad, status: red, timed_out: False, number_of_nodes: 4, number_of_data_nodes: 4, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 0, active_shards: 0, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, delayed_unassigned_shard [15:32:29] mber_of_pending_tasks: 64, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 6411, active_shards_percent_as_number: NaN https://wikitech.wikimedia.org/wiki/Search%23Administration [15:32:29] RECOVERY - OpenSearch health check for shards on 9200 on cirrussearch1068 is OK: OK - elasticsearch status production-search-eqiad: cluster_name: production-search-eqiad, status: red, timed_out: False, number_of_nodes: 4, number_of_data_nodes: 4, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 0, active_shards: 0, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, delayed_unassigned_shard [15:32:30] mber_of_pending_tasks: 64, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 7397, active_shards_percent_as_number: NaN https://wikitech.wikimedia.org/wiki/Search%23Administration [15:32:30] RECOVERY - OpenSearch health check for shards on 9200 on cirrussearch1122 is OK: OK - elasticsearch status production-search-eqiad: cluster_name: production-search-eqiad, status: red, timed_out: False, number_of_nodes: 4, number_of_data_nodes: 4, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 0, active_shards: 0, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, delayed_unassigned_shard [15:33:06] (03CR) 10CI reject: [V:04-1] wmf-config: Set JSON content model for Web2Cit configuration pages [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1314816 (https://phabricator.wikimedia.org/T305571) (owner: 10Diegodlh) [15:33:16] !log jayme@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage [15:33:34] can we do something about the opensearch alerts which have been very noisy and causing a bunch of flood kills for icinga-wm lately? [15:33:57] FIRING: [2x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [15:35:23] PROBLEM - OpenSearch health check for shards on 9200 on cirrussearch1094 is CRITICAL: CRITICAL - elasticsearch inactive shards 2580 threshold =0.15 breach: cluster_name: production-search-eqiad, status: yellow, timed_out: False, number_of_nodes: 55, number_of_data_nodes: 55, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1294, active_shards: 1294, relocating_shards: 0, initializing_shards: 0, unassigned_ [15:35:23] 2580, delayed_unassigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 33.402168301497156 https://wikitech.wikimedia.org/wiki/Search%23Administration [15:35:23] PROBLEM - OpenSearch health check for shards on 9200 on cirrussearch1068 is CRITICAL: CRITICAL - elasticsearch inactive shards 2580 threshold =0.15 breach: cluster_name: production-search-eqiad, status: yellow, timed_out: False, number_of_nodes: 55, number_of_data_nodes: 55, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1294, active_shards: 1294, relocating_shards: 0, initializing_shards: 0, unassigned_ [15:35:23] 2580, delayed_unassigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 33.402168301497156 https://wikitech.wikimedia.org/wiki/Search%23Administration [15:35:23] PROBLEM - OpenSearch health check for shards on 9200 on cirrussearch1086 is CRITICAL: CRITICAL - elasticsearch inactive shards 2580 threshold =0.15 breach: cluster_name: production-search-eqiad, status: yellow, timed_out: False, number_of_nodes: 55, number_of_data_nodes: 55, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1294, active_shards: 1294, relocating_shards: 0, initializing_shards: 0, unassigned_ [15:35:24] 2580, delayed_unassigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 33.402168301497156 https://wikitech.wikimedia.org/wiki/Search%23Administration [15:35:24] PROBLEM - OpenSearch health check for shards on 9200 on cirrussearch1079 is CRITICAL: CRITICAL - elasticsearch inactive shards 2580 threshold =0.15 breach: cluster_name: production-search-eqiad, status: yellow, timed_out: False, number_of_nodes: 55, number_of_data_nodes: 55, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1294, active_shards: 1294, relocating_shards: 0, initializing_shards: 0, unassigned_ [15:35:25] 2580, delayed_unassigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 33.402168301497156 https://wikitech.wikimedia.org/wiki/Search%23Administration [15:35:25] PROBLEM - OpenSearch health check for shards on 9200 on cirrussearch1080 is CRITICAL: CRITICAL - elasticsearch inactive shards 2580 threshold =0.15 breach: cluster_name: production-search-eqiad, status: yellow, timed_out: False, number_of_nodes: 55, number_of_data_nodes: 55, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1294, active_shards: 1294, relocating_shards: 0, initializing_shards: 0, unassigned_ [15:36:47] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [15:38:14] !log bking@cumin2003 START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - T324335 [15:38:18] T324335: Index the OpenSearch server application logs in the central observability cluster, using ECS format - https://phabricator.wikimedia.org/T324335 [15:39:39] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - search-psi-https_9643: Servers cirrussearch1086.eqiad.wmnet, cirrussearch1097.eqiad.wmnet, cirrussearch1087.eqiad.wmnet, cirrussearch1083.eqiad.wmnet, cirrussearch1079.eqiad.wmnet, cirrussearch1081.eqiad.wmnet, cirrussearch1121.eqiad.wmnet, cirrussearch1123.eqiad.wmnet, cirrussearch1073.eqiad.wmnet, cirrussearch1116.eqiad.wmnet, cirrussearch1101.eqia [15:39:39] cirrussearch1108.eqiad.wmnet, cirrussearch1085.eqiad.wmnet, cirrussearch1102.eqiad.wmnet, cirrussearch1088.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [15:40:41] RECOVERY - PyBal connections to etcd on lvs1020 is OK: OK: 116 connections established with conf2006.codfw.wmnet:4001 (min=116) https://wikitech.wikimedia.org/wiki/PyBal [15:41:15] (03PS3) 10Diegodlh: wmf-config: Set JSON content model for Web2Cit configuration pages [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1314816 (https://phabricator.wikimedia.org/T305571) [15:45:31] (03CR) 10Diegodlh: "Apparently some tests failed because of extra spaces in my code. I fixed them and updated the patch again. Sorry for that. I'm not sure ho" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1314816 (https://phabricator.wikimedia.org/T305571) (owner: 10Diegodlh) [15:45:41] PROBLEM - PyBal connections to etcd on lvs1017 is CRITICAL: CRITICAL: 0 connections established with conf2006.codfw.wmnet:4001 (min=12) https://wikitech.wikimedia.org/wiki/PyBal [15:45:45] RESOLVED: CirrusStreamingUpdaterRateTooLow: CirrusSearch update rate from flink-app-consumer-search is critically low - https://wikitech.wikimedia.org/wiki/Search#Streaming_Updater - https://grafana.wikimedia.org/d/jKqki4MSk/cirrus-streaming-updater - https://alerts.wikimedia.org/?q=alertname%3DCirrusStreamingUpdaterRateTooLow [15:46:06] (03PS1) 10Reedy: maintenance: Add require_once statements for AllUsers [extensions/OATHAuth] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1321583 (https://phabricator.wikimedia.org/T420792) [15:46:23] jouncebot: nowandnext [15:46:23] No deployments scheduled for the next 1 hour(s) and 13 minute(s) [15:46:24] In 1 hour(s) and 13 minute(s): MediaWiki infrastructure (UTC late) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260805T1700) [15:47:31] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch1078 is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: green, timed_out: False, number_of_nodes: 30, number_of_data_nodes: 30, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1741, active_shards: 5206, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, delaye [15:47:31] gned_shards: 0, number_of_pending_tasks: 38, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 1180, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [15:47:32] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch1073 is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: green, timed_out: False, number_of_nodes: 30, number_of_data_nodes: 30, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1741, active_shards: 5206, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, delaye [15:47:32] gned_shards: 0, number_of_pending_tasks: 38, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 1215, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [15:47:33] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch1092 is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: green, timed_out: False, number_of_nodes: 30, number_of_data_nodes: 30, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1741, active_shards: 5206, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, delaye [15:47:33] gned_shards: 0, number_of_pending_tasks: 38, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 1231, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [15:47:33] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch1079 is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: green, timed_out: False, number_of_nodes: 30, number_of_data_nodes: 30, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1741, active_shards: 5206, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, delaye [15:47:33] gned_shards: 0, number_of_pending_tasks: 40, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 1309, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [15:47:33] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch1099 is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: green, timed_out: False, number_of_nodes: 30, number_of_data_nodes: 30, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1741, active_shards: 5206, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, delaye [15:48:21] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [15:48:29] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch1095 is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: yellow, timed_out: False, number_of_nodes: 30, number_of_data_nodes: 30, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1741, active_shards: 4859, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 347, del [15:48:29] ssigned_shards: 347, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 93.33461390703035 https://wikitech.wikimedia.org/wiki/Search%23Administration [15:48:41] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [15:49:39] (03CR) 10Marostegui: [C:03+1] "You can deploy this anytime, if you connect thru the proxies this is already working." [puppet] - 10https://gerrit.wikimedia.org/r/1321581 (https://phabricator.wikimedia.org/T377889) (owner: 10AOkoth) [15:50:22] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, August 05 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320196 (https://phabricator.wikimedia.org/T428911) (owner: 10Jforrester) [15:50:56] (03CR) 10Brouberol: [C:03+1] airflow: main: add cloudflare_radar connection [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320964 (https://phabricator.wikimedia.org/T433982) (owner: 10CDanis) [15:51:08] (03CR) 10Brouberol: [C:03+1] "Thanks!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321520 (owner: 10Mvolz) [15:51:10] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, August 05 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311106 (owner: 10Jforrester) [15:51:14] (03CR) 10Reedy: [C:03+2] maintenance: Add require_once statements for AllUsers [extensions/OATHAuth] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1321583 (https://phabricator.wikimedia.org/T420792) (owner: 10Reedy) [15:51:39] 10ops-codfw, 06DC-Ops, 06ServiceOps, 10ServiceOps-Upgrades-Hardware: Power Supply - PS Redundancy - issue on wikikube-worker2187:9290 - https://phabricator.wikimedia.org/T433994#12188087 (10Jhancock.wm) 05Open→03Resolved a:03Jhancock.wm @Clement_Goubert this is fixed. did a firmware update first... [15:52:45] PROBLEM - Check unit status of push_cross_cluster_settings_9600 on cirrussearch1095 is CRITICAL: CRITICAL: Status of the systemd unit push_cross_cluster_settings_9600 https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [15:52:57] (03Merged) 10jenkins-bot: maintenance: Add require_once statements for AllUsers [extensions/OATHAuth] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1321583 (https://phabricator.wikimedia.org/T420792) (owner: 10Reedy) [15:52:57] (03PS1) 10Jforrester: abstractwiki: Add dag/ml/ig/ha languages to generation script [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1321585 [15:53:07] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, August 05 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1321585 (owner: 10Jforrester) [15:54:05] (03CR) 10Marostegui: "@cwilliams@wikimedia.org if you have time, let's review before the end of the week so we can close it and start fresh next week. Thanks!" [cookbooks] - 10https://gerrit.wikimedia.org/r/1291952 (https://phabricator.wikimedia.org/T426613) (owner: 10Federico Ceratto) [15:54:17] !log jayme@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1159.eqiad.wmnet with OS trixie [15:55:41] RECOVERY - PyBal connections to etcd on lvs1017 is OK: OK: 12 connections established with conf2006.codfw.wmnet:4001 (min=12) https://wikitech.wikimedia.org/wiki/PyBal [15:56:07] !log reedy@deploy1003 Started scap sync-world: Backport for [[gerrit:1321583|maintenance: Add require_once statements for AllUsers (T420792)]] [15:56:11] T420792: Allow 2FA to be enforced for all accounts on a private wiki - https://phabricator.wikimedia.org/T420792 [15:58:03] !log reedy@deploy1003 reedy: Backport for [[gerrit:1321583|maintenance: Add require_once statements for AllUsers (T420792)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [15:58:59] !log reedy@deploy1003 reedy: Continuing with deployment [16:01:46] (03CR) 10Klausman: [C:03+1] airflow-dev: ensure a dev instance cannot send an email [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321580 (https://phabricator.wikimedia.org/T416596) (owner: 10Brouberol) [16:02:01] (03CR) 10Brouberol: [C:03+2] airflow-dev: ensure a dev instance cannot send an email [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321580 (https://phabricator.wikimedia.org/T416596) (owner: 10Brouberol) [16:02:39] FIRING: [2x] CirrusSearchNodeIndexingNotIncreasing: OpenSearch instance cirrussearch1093-production-search-eqiad is not indexing - https://wikitech.wikimedia.org/wiki/Search/OpenSearch/Administration#Alerts/Dashboards - https://grafana.wikimedia.org/d/JLK3I_siz/elasticsearch-indexing?orgId=1&from=now-3d&to=now&viewPanel=57 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchNodeIndexingNotIncreasing [16:02:45] RECOVERY - Check unit status of push_cross_cluster_settings_9600 on cirrussearch1095 is OK: OK: Status of the systemd unit push_cross_cluster_settings_9600 https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [16:04:15] 06SRE, 06cloud-services-team, 10Cloud-VPS, 06tools-infrastructure-team, 13Patch-For-Review: Modernise memcached systemd unit / sync, and make it presentable - https://phabricator.wikimedia.org/T273950#12188151 (10jijiki) [16:05:18] !log reedy@deploy1003 Finished scap sync-world: Backport for [[gerrit:1321583|maintenance: Add require_once statements for AllUsers (T420792)]] (duration: 09m 11s) [16:05:23] T420792: Allow 2FA to be enforced for all accounts on a private wiki - https://phabricator.wikimedia.org/T420792 [16:05:27] 10ops-codfw, 10ops-drmrs, 10ops-eqdfw, 10ops-eqiad, and 6 others: Fix cables with placeholders names and planned status - https://phabricator.wikimedia.org/T432317#12188152 (10Jhancock.wm) [16:06:15] !log jayme@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1159.eqiad.wmnet [16:06:17] !log jayme@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1159.eqiad.wmnet [16:06:20] !log jayme@cumin1003 END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1159.eqiad.wmnet [16:08:10] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:09:00] 06SRE, 06cloud-services-team, 10Cloud-VPS, 06tools-infrastructure-team, 13Patch-For-Review: Modernise memcached systemd unit / sync, and make it presentable - https://phabricator.wikimedia.org/T273950#12188180 (10jijiki) [16:13:57] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:16:22] !log cdobbins@cumin1003 START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams and A:liberica (T428495) [16:16:23] RECOVERY - OpenSearch health check for shards on 9200 on cirrussearch1095 is OK: OK - elasticsearch status production-search-eqiad: cluster_name: production-search-eqiad, status: yellow, timed_out: False, number_of_nodes: 55, number_of_data_nodes: 55, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1294, active_shards: 3850, relocating_shards: 0, initializing_shards: 24, unassigned_shards: 0, delayed_unas [16:16:23] hards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 99.38048528652556 https://wikitech.wikimedia.org/wiki/Search%23Administration [16:16:23] RECOVERY - OpenSearch health check for shards on 9200 on cirrussearch1073 is OK: OK - elasticsearch status production-search-eqiad: cluster_name: production-search-eqiad, status: yellow, timed_out: False, number_of_nodes: 55, number_of_data_nodes: 55, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1294, active_shards: 3850, relocating_shards: 0, initializing_shards: 24, unassigned_shards: 0, delayed_unas [16:16:23] hards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 99.38048528652556 https://wikitech.wikimedia.org/wiki/Search%23Administration [16:16:23] RECOVERY - OpenSearch health check for shards on 9200 on cirrussearch1068 is OK: OK - elasticsearch status production-search-eqiad: cluster_name: production-search-eqiad, status: yellow, timed_out: False, number_of_nodes: 55, number_of_data_nodes: 55, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1294, active_shards: 3851, relocating_shards: 0, initializing_shards: 23, unassigned_shards: 0, delayed_unas [16:16:24] hards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 99.40629839958699 https://wikitech.wikimedia.org/wiki/Search%23Administration [16:16:24] RECOVERY - OpenSearch health check for shards on 9200 on cirrussearch1093 is OK: OK - elasticsearch status production-search-eqiad: cluster_name: production-search-eqiad, status: yellow, timed_out: False, number_of_nodes: 55, number_of_data_nodes: 55, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1294, active_shards: 3851, relocating_shards: 0, initializing_shards: 23, unassigned_shards: 0, delayed_unas [16:16:25] hards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 99.40629839958699 https://wikitech.wikimedia.org/wiki/Search%23Administration [16:16:25] RECOVERY - OpenSearch health check for shards on 9200 on cirrussearch1094 is OK: OK - elasticsearch status production-search-eqiad: cluster_name: production-search-eqiad, status: yellow, timed_out: False, number_of_nodes: 55, number_of_data_nodes: 55, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1294, active_shards: 3851, relocating_shards: 0, initializing_shards: 23, unassigned_shards: 0, delayed_unas [16:16:26] T428495: Migrate conf* hosts away from bullseye - https://phabricator.wikimedia.org/T428495 [16:17:39] RESOLVED: [2x] CirrusSearchNodeIndexingNotIncreasing: OpenSearch instance cirrussearch1093-production-search-eqiad is not indexing - https://wikitech.wikimedia.org/wiki/Search/OpenSearch/Administration#Alerts/Dashboards - https://grafana.wikimedia.org/d/JLK3I_siz/elasticsearch-indexing?orgId=1&from=now-3d&to=now&viewPanel=57 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchNodeIndexingNotIncreasing [16:17:58] (03PS1) 10Clare Ming: logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 [extensions/WikiLambda] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1321590 (https://phabricator.wikimedia.org/T433550) [16:18:14] (03PS1) 10Clare Ming: logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 [extensions/WikiLambda] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1321591 (https://phabricator.wikimedia.org/T433550) [16:18:16] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams and A:liberica (T428495) [16:18:32] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, August 05 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-" [extensions/WikiLambda] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1321590 (https://phabricator.wikimedia.org/T433550) (owner: 10Clare Ming) [16:18:44] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, August 05 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-" [extensions/WikiLambda] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1321591 (https://phabricator.wikimedia.org/T433550) (owner: 10Clare Ming) [16:18:58] !log cdobbins@cumin1003 START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs and A:liberica (T428495) [16:20:45] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs and A:liberica (T428495) [16:21:21] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - search-psi-https_9643: Servers cirrussearch1095.eqiad.wmnet, cirrussearch1111.eqiad.wmnet, cirrussearch1097.eqiad.wmnet, cirrussearch1122.eqiad.wmnet, cirrussearch1115.eqiad.wmnet, cirrussearch1117.eqiad.wmnet, cirrussearch1108.eqiad.wmnet, cirrussearch1079.eqiad.wmnet, cirrussearch1121.eqiad.wmnet, cirrussearch1123.eqiad.wmnet, cirrussearch1101.eqia [16:21:21] cirrussearch1078.eqiad.wmnet, cirrussearch1110.eqiad.wmnet, cirrussearch1085.eqiad.wmnet, cirrussearch1116.eqiad.wmnet are marked down but pooled: search-omega-https_9443: Servers cirrussearch1068.eqiad.wmnet, cirrussearch1096.eqiad.wmnet, cirrussearch1091.eqiad.wmnet, cirrussearch1098.eqiad.wmnet, cirrussearch1089.eqiad.wmnet, cirrussearch1094.eqiad.wmnet, cirrussearch1113.eqiad.wmnet, cirrussearch1125.eqiad.wmnet, cirrussearch1082.eqia [16:21:21] cirrussearch1100.eqiad.wmnet, cirrussearch1114.eqiad.wmnet, cirrussearch1124.eqiad.wmnet, cirrussearch1070.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [16:21:29] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1076 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:29] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1074 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:31] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1071 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:31] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1068 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:31] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1082 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:31] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1080 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:31] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1100 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:32] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1093 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:32] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1096 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:33] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1098 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:33] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1089 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:34] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1070 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:34] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1094 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:35] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1091 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:35] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1107 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:36] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1077 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:36] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1109 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:37] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1103 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:37] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1112 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:38] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1113 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:38] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1120 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:39] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1125 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:39] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1124 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:40] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1114 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:40] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1072 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:41] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1078 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:41] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1073 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:42] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1086 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:42] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1069 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:43] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1075 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:43] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1102 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:44] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1095 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:44] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1122 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:45] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1099 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:45] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1108 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:46] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1084 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:46] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1092 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:47] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1087 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:47] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1079 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:48] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1083 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:48] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1085 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:49] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1110 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:50] PROBLEM - ElasticSearch health check for shards on 9443 on search.svc.eqiad.wmnet is CRITICAL: CRITICAL - elasticsearch https://search.svc.eqiad.wmnet:9443/_cluster/health error while fetching: HTTPSConnectionPool(host=search.svc.eqiad.wmnet, port=9443): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:50] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - search-psi-https_9643: Servers cirrussearch1111.eqiad.wmnet, cirrussearch1097.eqiad.wmnet, cirrussearch1115.eqiad.wmnet, cirrussearch1117.eqiad.wmnet, cirrussearch1083.eqiad.wmnet, cirrussearch1079.eqiad.wmnet, cirrussearch1121.eqiad.wmnet, cirrussearch1123.eqiad.wmnet, cirrussearch1116.eqiad.wmnet, cirrussearch1101.eqiad.wmnet, cirrussearch1078.eqia [16:21:51] cirrussearch1110.eqiad.wmnet, cirrussearch1085.eqiad.wmnet, cirrussearch1099.eqiad.wmnet, cirrussearch1088.eqiad.wmnet are marked down but pooled: search-omega-https_9443: Servers cirrussearch1068.eqiad.wmnet, cirrussearch1091.eqiad.wmnet, cirrussearch1093.eqiad.wmnet, cirrussearch1080.eqiad.wmnet, cirrussearch1089.eqiad.wmnet, cirrussearch1094.eqiad.wmnet, cirrussearch1077.eqiad.wmnet, cirrussearch1103.eqiad.wmnet, cirrussearch1071.eqia [16:21:51] cirrussearch1120.eqiad.wmnet, cirrussearch1082.eqiad.wmnet, cirrussearch1096.eqiad.wmnet, cirrussearch1070.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [16:21:52] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1090 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:52] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1117 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:53] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1116 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:53] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1111 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:54] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1121 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:54] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1097 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:55] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1088 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:55] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1123 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:56] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1101 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:56] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch1115 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:21:57] PROBLEM - ElasticSearch health check for shards on 9643 on search.svc.eqiad.wmnet is CRITICAL: CRITICAL - elasticsearch https://search.svc.eqiad.wmnet:9643/_cluster/health error while fetching: HTTPSConnectionPool(host=search.svc.eqiad.wmnet, port=9643): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:23:16] !log cdobbins@cumin1003 START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru and A:liberica (T428495) [16:23:20] T428495: Migrate conf* hosts away from bullseye - https://phabricator.wikimedia.org/T428495 [16:23:23] (03CR) 10RLazarus: [C:03+1] Update to v1.39.0 [debs/envoyproxy] (v1.39) - 10https://gerrit.wikimedia.org/r/1321526 (https://phabricator.wikimedia.org/T421418) (owner: 10Blake) [16:24:15] 10ops-eqiad, 06DC-Ops, 06Data-Platform-SRE (2026-07-03 - 2026-07-31): Unresponsive management for an-worker1147.mgmt:22 - https://phabricator.wikimedia.org/T433348#12188295 (10VRiley-WMF) Thank you! I'm ready whenever they are ready to take it down. [16:25:18] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru and A:liberica (T428495) [16:25:21] 10ops-eqiad, 06SRE, 06Data-Persistence, 10Data-Persistence-Backup, and 3 others: decommission db1171.eqiad.wmnet - https://phabricator.wikimedia.org/T433826#12188300 (10VRiley-WMF) a:03VRiley-WMF [16:25:46] 10ops-eqiad, 06SRE, 06Data-Persistence, 10Data-Persistence-Backup, and 3 others: decommission db1150.eqiad.wmnet - https://phabricator.wikimedia.org/T433825#12188304 (10VRiley-WMF) a:03VRiley-WMF [16:31:41] hi folks, for awareness, will be rebooting kafka-main-eqiad brokers momentarily to pick up a kernel upgrade, please ping me if anything~ [16:33:08] (03PS1) 10Brouberol: airflow-dev: fix typo [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321599 (https://phabricator.wikimedia.org/T416596) [16:34:57] !log bking@cumin2003 END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - T324335 [16:35:01] T324335: Index the OpenSearch server application logs in the central observability cluster, using ECS format - https://phabricator.wikimedia.org/T324335 [16:35:11] (03PS2) 10Brouberol: airflow-dev: fix typo [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321599 (https://phabricator.wikimedia.org/T416596) [16:35:37] (03Abandoned) 10Clare Ming: logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 [extensions/WikiLambda] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1321590 (https://phabricator.wikimedia.org/T433550) (owner: 10Clare Ming) [16:36:21] RECOVERY - OpenSearch health check for shards on 9400 on cirrussearch1076 is OK: OK - elasticsearch status production-search-omega-eqiad: cluster_name: production-search-omega-eqiad, status: green, timed_out: False, number_of_nodes: 25, number_of_data_nodes: 25, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 5195, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, de [16:36:21] assigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [16:36:21] RECOVERY - OpenSearch health check for shards on 9400 on cirrussearch1074 is OK: OK - elasticsearch status production-search-omega-eqiad: cluster_name: production-search-omega-eqiad, status: green, timed_out: False, number_of_nodes: 25, number_of_data_nodes: 25, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 5195, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, de [16:36:21] assigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [16:36:21] RECOVERY - OpenSearch health check for shards on 9400 on cirrussearch1068 is OK: OK - elasticsearch status production-search-omega-eqiad: cluster_name: production-search-omega-eqiad, status: green, timed_out: False, number_of_nodes: 25, number_of_data_nodes: 25, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 5195, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, de [16:36:22] assigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [16:36:22] RECOVERY - OpenSearch health check for shards on 9400 on cirrussearch1071 is OK: OK - elasticsearch status production-search-omega-eqiad: cluster_name: production-search-omega-eqiad, status: green, timed_out: False, number_of_nodes: 25, number_of_data_nodes: 25, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 5195, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, de [16:36:23] assigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [16:36:23] RECOVERY - OpenSearch health check for shards on 9400 on cirrussearch1082 is OK: OK - elasticsearch status production-search-omega-eqiad: cluster_name: production-search-omega-eqiad, status: green, timed_out: False, number_of_nodes: 25, number_of_data_nodes: 25, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 5195, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, de [16:36:24] assigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [16:36:24] RECOVERY - OpenSearch health check for shards on 9400 on cirrussearch1080 is OK: OK - elasticsearch status production-search-omega-eqiad: cluster_name: production-search-omega-eqiad, status: green, timed_out: False, number_of_nodes: 25, number_of_data_nodes: 25, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 5195, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, de [16:36:25] assigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [16:36:25] RECOVERY - OpenSearch health check for shards on 9400 on cirrussearch1100 is OK: OK - elasticsearch status production-search-omega-eqiad: cluster_name: production-search-omega-eqiad, status: green, timed_out: False, number_of_nodes: 25, number_of_data_nodes: 25, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 5195, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, de [16:38:25] (03CR) 10Ladsgroup: [V:03+2 C:03+2] "You need to set the Hosts: footer so it only runs on those hosts. deploy1003.eqiad.wmnet would be enough. Otherwise it tries to run it on " [puppet] - 10https://gerrit.wikimedia.org/r/1290104 (https://phabricator.wikimedia.org/T107188) (owner: 10Krinkle) [16:39:21] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [16:39:41] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [16:40:30] jasmine@cumin2002 roll-restart-reboot-brokers (PID 87462) is awaiting input [16:40:40] !log bking@cumin2003 START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - T324335 [16:40:42] !log jasmine@cumin2002 START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-eqiad [16:40:44] T324335: Index the OpenSearch server application logs in the central observability cluster, using ECS format - https://phabricator.wikimedia.org/T324335 [16:41:21] !log cgoubert@cumin2003 START - Cookbook sre.hosts.remove-downtime for wikikube-worker2187.codfw.wmnet [16:41:22] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2187.codfw.wmnet [16:41:27] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch1078 is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: green, timed_out: False, number_of_nodes: 30, number_of_data_nodes: 30, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1741, active_shards: 5206, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, delaye [16:41:27] gned_shards: 0, number_of_pending_tasks: 1, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [16:41:27] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch1072 is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: green, timed_out: False, number_of_nodes: 30, number_of_data_nodes: 30, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1741, active_shards: 5206, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, delaye [16:41:27] gned_shards: 0, number_of_pending_tasks: 1, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [16:41:27] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch1069 is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: green, timed_out: False, number_of_nodes: 30, number_of_data_nodes: 30, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1741, active_shards: 5206, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, delaye [16:41:28] gned_shards: 0, number_of_pending_tasks: 1, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [16:41:28] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch1073 is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: green, timed_out: False, number_of_nodes: 30, number_of_data_nodes: 30, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1741, active_shards: 5206, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, delaye [16:41:29] gned_shards: 0, number_of_pending_tasks: 1, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [16:41:29] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch1086 is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: green, timed_out: False, number_of_nodes: 30, number_of_data_nodes: 30, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1741, active_shards: 5206, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, delaye [16:41:30] gned_shards: 0, number_of_pending_tasks: 1, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [16:41:31] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch1102 is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: green, timed_out: False, number_of_nodes: 30, number_of_data_nodes: 30, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1741, active_shards: 5206, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, delaye [16:41:31] gned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [16:41:31] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch1122 is OK: OK - elasticsearch status production-search-psi-eqiad: cluster_name: production-search-psi-eqiad, status: green, timed_out: False, number_of_nodes: 30, number_of_data_nodes: 30, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1741, active_shards: 5206, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, delaye [16:41:36] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2187.codfw.wmnet [16:41:38] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2187.codfw.wmnet [16:41:41] 10ops-codfw, 06DC-Ops, 06ServiceOps, 10ServiceOps-Upgrades-Hardware: Power Supply - PS Redundancy - issue on wikikube-worker2187:9290 - https://phabricator.wikimedia.org/T433994#12188444 (10ops-monitoring-bot) Cookbook cookbooks.sre.k8s.pool-depool-node started by cgoubert@cumin2003 pool for host wikik... [16:41:43] (03PS3) 10AOkoth: mariadb: add grants for phab1005 [puppet] - 10https://gerrit.wikimedia.org/r/1321581 (https://phabricator.wikimedia.org/T377889) [16:41:54] 10ops-codfw, 06DC-Ops, 06ServiceOps, 10ServiceOps-Upgrades-Hardware: Power Supply - PS Redundancy - issue on wikikube-worker2187:9290 - https://phabricator.wikimedia.org/T433994#12188446 (10Clement_Goubert) Thanks, removed downtime and repooled. [16:43:51] (03CR) 10Dzahn: [C:03+2] disable the banner about the SSH host key change [software/gerrit] (deploy/wmf/stable-3.10) - 10https://gerrit.wikimedia.org/r/1314995 (https://phabricator.wikimedia.org/T240266) (owner: 10Dzahn) [16:44:04] (03CR) 10David Caro: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1321058 (owner: 10Andrew Bogott) [16:44:20] (03CR) 10Brouberol: [C:03+2] airflow-dev: fix typo [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321599 (https://phabricator.wikimedia.org/T416596) (owner: 10Brouberol) [16:44:21] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - search-omega-https_9443: Servers cirrussearch1107.eqiad.wmnet, cirrussearch1089.eqiad.wmnet, cirrussearch1109.eqiad.wmnet, cirrussearch1077.eqiad.wmnet, cirrussearch1113.eqiad.wmnet, cirrussearch1125.eqiad.wmnet, cirrussearch1071.eqiad.wmnet, cirrussearch1103.eqiad.wmnet, cirrussearch1112.eqiad.wmnet, cirrussearch1096.eqiad.wmnet, cirrussearch1114.eq [16:44:21] t, cirrussearch1124.eqiad.wmnet, cirrussearch1070.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [16:44:29] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1076 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:44:29] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1074 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:44:31] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1068 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:44:31] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1080 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:44:31] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1082 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:44:31] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1071 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:44:33] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1094 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:44:33] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1091 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:44:33] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1070 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:44:33] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1093 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:44:33] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1077 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:44:33] (03Merged) 10jenkins-bot: disable the banner about the SSH host key change [software/gerrit] (deploy/wmf/stable-3.10) - 10https://gerrit.wikimedia.org/r/1314995 (https://phabricator.wikimedia.org/T240266) (owner: 10Dzahn) [16:44:34] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1089 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:44:34] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1109 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:44:35] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1096 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:44:35] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1103 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:44:36] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1107 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:44:36] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1113 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:44:37] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1124 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:44:37] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1112 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:44:38] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1120 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:44:38] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1125 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:44:39] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1114 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:44:41] PROBLEM - ElasticSearch health check for shards on 9443 on search.svc.eqiad.wmnet is CRITICAL: CRITICAL - elasticsearch https://search.svc.eqiad.wmnet:9443/_cluster/health error while fetching: HTTPSConnectionPool(host=search.svc.eqiad.wmnet, port=9443): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:44:41] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - search-omega-https_9443: Servers cirrussearch1107.eqiad.wmnet, cirrussearch1082.eqiad.wmnet, cirrussearch1109.eqiad.wmnet, cirrussearch1077.eqiad.wmnet, cirrussearch1125.eqiad.wmnet, cirrussearch1113.eqiad.wmnet, cirrussearch1103.eqiad.wmnet, cirrussearch1120.eqiad.wmnet, cirrussearch1112.eqiad.wmnet, cirrussearch1096.eqiad.wmnet, cirrussearch1114.eq [16:44:41] t, cirrussearch1124.eqiad.wmnet, cirrussearch1070.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [16:48:00] (03CR) 10Dzahn: [C:03+2] "Ahmon, remind me what we did to enable it?" [software/gerrit] (deploy/wmf/stable-3.10) - 10https://gerrit.wikimedia.org/r/1314995 (https://phabricator.wikimedia.org/T240266) (owner: 10Dzahn) [16:48:54] (03PS2) 10Blake: mw-pretrain: Add a jobrunner values.yaml. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321561 (https://phabricator.wikimedia.org/T427668) [16:49:00] (03CR) 10BCornwall: "Some of the feedback was that eqsin had surprisingly poor latency, so I'm not sure about the numbers here. @cdobbins@wikimedia.org ?" [dns] - 10https://gerrit.wikimedia.org/r/1306727 (owner: 10BCornwall) [16:49:23] 06SRE, 10SRE-Access-Requests, 06Release-Engineering-Team: Requesting access to Jenkins server for PWangai-WMF - https://phabricator.wikimedia.org/T433563#12188481 (10brennen) > Cool! in that case I think the simplest path forward is we just add people to "contint-users" Discussed with team in RelEng weekly,... [16:51:33] RECOVERY - OpenSearch health check for shards on 9400 on cirrussearch1094 is OK: OK - elasticsearch status production-search-omega-eqiad: cluster_name: production-search-omega-eqiad, status: green, timed_out: False, number_of_nodes: 25, number_of_data_nodes: 25, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 5195, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, de [16:51:33] assigned_shards: 0, number_of_pending_tasks: 50, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 3617, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [16:51:33] RECOVERY - OpenSearch health check for shards on 9400 on cirrussearch1113 is OK: OK - elasticsearch status production-search-omega-eqiad: cluster_name: production-search-omega-eqiad, status: green, timed_out: False, number_of_nodes: 25, number_of_data_nodes: 25, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 5195, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, de [16:51:33] assigned_shards: 0, number_of_pending_tasks: 50, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 3599, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [16:51:33] RECOVERY - OpenSearch health check for shards on 9400 on cirrussearch1125 is OK: OK - elasticsearch status production-search-omega-eqiad: cluster_name: production-search-omega-eqiad, status: green, timed_out: False, number_of_nodes: 25, number_of_data_nodes: 25, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 5195, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, de [16:52:21] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [16:52:23] RECOVERY - OpenSearch health check for shards on 9400 on cirrussearch1068 is OK: OK - elasticsearch status production-search-omega-eqiad: cluster_name: production-search-omega-eqiad, status: green, timed_out: False, number_of_nodes: 25, number_of_data_nodes: 25, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 5195, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, de [16:52:23] assigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [16:52:23] RECOVERY - OpenSearch health check for shards on 9400 on cirrussearch1076 is OK: OK - elasticsearch status production-search-omega-eqiad: cluster_name: production-search-omega-eqiad, status: green, timed_out: False, number_of_nodes: 25, number_of_data_nodes: 25, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 5195, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, de [16:52:23] assigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [16:52:23] RECOVERY - OpenSearch health check for shards on 9400 on cirrussearch1080 is OK: OK - elasticsearch status production-search-omega-eqiad: cluster_name: production-search-omega-eqiad, status: green, timed_out: False, number_of_nodes: 25, number_of_data_nodes: 25, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 5195, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, de [16:52:24] assigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [16:52:24] RECOVERY - OpenSearch health check for shards on 9400 on cirrussearch1074 is OK: OK - elasticsearch status production-search-omega-eqiad: cluster_name: production-search-omega-eqiad, status: green, timed_out: False, number_of_nodes: 25, number_of_data_nodes: 25, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 5195, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, de [16:52:25] assigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [16:52:25] RECOVERY - OpenSearch health check for shards on 9400 on cirrussearch1082 is OK: OK - elasticsearch status production-search-omega-eqiad: cluster_name: production-search-omega-eqiad, status: green, timed_out: False, number_of_nodes: 25, number_of_data_nodes: 25, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 5195, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, de [16:52:26] assigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [16:52:26] RECOVERY - OpenSearch health check for shards on 9400 on cirrussearch1071 is OK: OK - elasticsearch status production-search-omega-eqiad: cluster_name: production-search-omega-eqiad, status: green, timed_out: False, number_of_nodes: 25, number_of_data_nodes: 25, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 5195, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, de [16:52:27] assigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [16:53:49] !log bking@cumin2003 END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - T324335 [16:53:54] T324335: Index the OpenSearch server application logs in the central observability cluster, using ECS format - https://phabricator.wikimedia.org/T324335 [16:55:43] !log bking@cumin2003 conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad [16:55:44] !log bking@cumin2003 conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad [16:55:44] !log bking@cumin2003 conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad [16:56:54] (03PS1) 10Dzahn: admin: upgrade pwangai from ldap_only to contint-users shell [puppet] - 10https://gerrit.wikimedia.org/r/1321601 (https://phabricator.wikimedia.org/T433563) [16:57:16] (03CR) 10CI reject: [V:04-1] admin: upgrade pwangai from ldap_only to contint-users shell [puppet] - 10https://gerrit.wikimedia.org/r/1321601 (https://phabricator.wikimedia.org/T433563) (owner: 10Dzahn) [16:57:29] !log aokoth@deploy1003 Started deploy [phabricator/deployment@e2ebca5]: Deploy Phab [16:58:04] !log aokoth@deploy1003 Finished deploy [phabricator/deployment@e2ebca5]: Deploy Phab (duration: 00m 34s) [16:58:10] 06SRE, 10SRE-Access-Requests, 06Release-Engineering-Team, 13Patch-For-Review: Requesting access to Jenkins server for PWangai-WMF - https://phabricator.wikimedia.org/T433563#12188521 (10Dzahn) [16:59:24] (03PS2) 10Dzahn: admin: upgrade pwangai from ldap_only to contint-users shell [puppet] - 10https://gerrit.wikimedia.org/r/1321601 (https://phabricator.wikimedia.org/T433563) [16:59:31] (03CR) 10Andrew Bogott: [C:03+2] wmcs-backup: don't error out on a few failed snapshot removals [puppet] - 10https://gerrit.wikimedia.org/r/1321058 (owner: 10Andrew Bogott) [16:59:57] (03CR) 10Andrew Bogott: [C:03+2] wmcs-backup: don't error out on a few failed snapshot removals (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1321058 (owner: 10Andrew Bogott) [17:00:05] Deploy window MediaWiki infrastructure (UTC late) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260805T1700) [17:01:49] 06SRE, 10SRE-Access-Requests, 06Release-Engineering-Team, 13Patch-For-Review: Requesting access to Jenkins server for PWangai-WMF - https://phabricator.wikimedia.org/T433563#12188543 (10Dzahn) [17:02:47] 06SRE, 10SRE-Access-Requests, 06Release-Engineering-Team, 13Patch-For-Review: Requesting access to Jenkins server for PWangai-WMF - https://phabricator.wikimedia.org/T433563#12188552 (10Dzahn) @dancy You are an approver for the `contint-users` group. Approved? [17:05:57] 06SRE, 10SRE-Access-Requests, 06Release-Engineering-Team, 13Patch-For-Review: Requesting access to Jenkins server for PWangai-WMF - https://phabricator.wikimedia.org/T433563#12188566 (10dancy) >>! In T433563#12188543, @Dzahn wrote: > @dancy You are an approver for the `contint-users` group. Approved? Appr... [17:07:17] 06SRE, 10SRE-Access-Requests, 06Release-Engineering-Team, 13Patch-For-Review: Requesting access to Jenkins server for PWangai-WMF - https://phabricator.wikimedia.org/T433563#12188572 (10Dzahn) [17:09:20] 06SRE, 10LDAP-Access-Requests: Grant Access to ciadmin for Vaughn Walters - https://phabricator.wikimedia.org/T433615#12188598 (10dancy) This request has my approval. [17:12:05] !log LDAP - added vwalters to group ciadmin - T433615 [17:12:08] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [17:12:09] T433615: Grant Access to ciadmin for Vaughn Walters - https://phabricator.wikimedia.org/T433615 [17:12:48] 06SRE, 10LDAP-Access-Requests: Grant Access to ciadmin for Vaughn Walters - https://phabricator.wikimedia.org/T433615#12188637 (10Dzahn) 05In progress→03Resolved a:05Lferreira→03Dzahn thanks all. this is done. I added Vaughn to the ciadmin LDAP group. [17:14:37] 10ops-eqiad, 06DC-Ops: Alert for device ps1-e2-eqiad.mgmt.eqiad.wmnet - PDU sensor over limit - https://phabricator.wikimedia.org/T434118 (10phaultfinder) 03NEW [17:14:58] (03CR) 10AOkoth: [C:03+1] admin: upgrade pwangai from ldap_only to contint-users shell [puppet] - 10https://gerrit.wikimedia.org/r/1321601 (https://phabricator.wikimedia.org/T433563) (owner: 10Dzahn) [17:15:12] (03CR) 10Ssingh: "Yeah good call, let's run the numbers again from Probenet and see what they tell us." [dns] - 10https://gerrit.wikimedia.org/r/1306727 (owner: 10BCornwall) [17:15:21] (03CR) 10Ahmon Dancy: "We're at step 6 of https://wikitech.wikimedia.org/wiki/Gerrit/MOTD" [software/gerrit] (deploy/wmf/stable-3.10) - 10https://gerrit.wikimedia.org/r/1314995 (https://phabricator.wikimedia.org/T240266) (owner: 10Dzahn) [17:16:28] (03PS1) 10Krinkle: mediawiki: Disable legacy `short_urls` on vhosts where it does not work (take 2) [puppet] - 10https://gerrit.wikimedia.org/r/1321603 (https://phabricator.wikimedia.org/T107188) [17:20:53] 06SRE, 10Maps, 06Traffic: Possibility to allow Wikimedia Maps usage on all Wikibase Cloud instances - https://phabricator.wikimedia.org/T429191#12188710 (10ssingh) >>! In T429191#12186892, @Anton.Kokh wrote: > @MSantos There are currently around 800K+ pages using geo coordinates across 140 Wikibases. > 90% o... [17:21:26] (03CR) 10Dzahn: [C:03+2] "thank you! i tried to follow these steps first on deploy2002 - where they showed no unexpected diffs but then scap was disabled because of" [software/gerrit] (deploy/wmf/stable-3.10) - 10https://gerrit.wikimedia.org/r/1314995 (https://phabricator.wikimedia.org/T240266) (owner: 10Dzahn) [17:21:31] (03CR) 10Krinkle: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1321603 (https://phabricator.wikimedia.org/T107188) (owner: 10Krinkle) [17:21:38] (03CR) 10CDanis: "Thanks Balto! Can we do the private patch and the deploy together tomorrow? (Or someone who already knows how can deploy themselves, don't" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320964 (https://phabricator.wikimedia.org/T433982) (owner: 10CDanis) [17:25:44] (03PS2) 10Krinkle: mediawiki: Disable legacy `short_urls` on vhosts where it does not work (take 2) [puppet] - 10https://gerrit.wikimedia.org/r/1321603 (https://phabricator.wikimedia.org/T107188) [17:26:01] (03CR) 10Krinkle: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1321603 (https://phabricator.wikimedia.org/T107188) (owner: 10Krinkle) [17:26:29] 10ops-eqiad, 06SRE, 06Data-Persistence, 10Data-Persistence-Backup, and 3 others: decommission db1171.eqiad.wmnet - https://phabricator.wikimedia.org/T433826#12188756 (10VRiley-WMF) [17:26:53] 10ops-eqiad, 06SRE, 06Data-Persistence, 10Data-Persistence-Backup, and 3 others: decommission db1171.eqiad.wmnet - https://phabricator.wikimedia.org/T433826#12188757 (10VRiley-WMF) 05Open→03Resolved This has been completed. [17:27:31] 10ops-eqiad, 06SRE, 06Data-Persistence, 10Data-Persistence-Backup, and 3 others: decommission db1150.eqiad.wmnet - https://phabricator.wikimedia.org/T433825#12188760 (10VRiley-WMF) [17:27:43] 10ops-eqiad, 06SRE, 06Data-Persistence, 10Data-Persistence-Backup, and 3 others: decommission db1150.eqiad.wmnet - https://phabricator.wikimedia.org/T433825#12188762 (10VRiley-WMF) 05Open→03Resolved This has been completed. [17:27:47] 10ops-eqsin, 06SRE, 06DC-Ops, 06Traffic: cp5022 is unreachable - https://phabricator.wikimedia.org/T414411#12188766 (10ssingh) >>! In T414411#12183764, @RobH wrote: > pre-onsite checks: host mgmt is online and i can connect to it, but I'm having issues on its idrac interface. I want to push new ilom firmw... [17:30:20] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-07-03 - 2026-07-31): Degraded RAID on an-worker1191 - https://phabricator.wikimedia.org/T431828#12188774 (10VRiley-WMF) @brouberol Hey I wanted to check in with you on this device as I believe Ben may be out of the office for a bit. Would there be any... [17:31:51] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for BGerdemann_(WMF) - https://phabricator.wikimedia.org/T433313#12188779 (10Arnoldokoth) 05Open→03In progress [17:31:54] !log jasmine@cumin2002 END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-eqiad [17:33:25] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-07-03 - 2026-07-31): Degraded RAID on an-presto1013 - https://phabricator.wikimedia.org/T433030#12188805 (10VRiley-WMF) Currently waiting on HDD to arrive [17:34:51] (03CR) 10Ayounsi: [C:03+1] Set standard metric for new easms and drmrs transports to codfw [homer/public] - 10https://gerrit.wikimedia.org/r/1321572 (https://phabricator.wikimedia.org/T424839) (owner: 10Cathal Mooney) [17:38:53] (03CR) 10Dzahn: [C:03+1] sre.gitlab: fix rollback crash and minor pre-existing issues [cookbooks] - 10https://gerrit.wikimedia.org/r/1320978 (https://phabricator.wikimedia.org/T430655) (owner: 10Arnaudb) [17:43:09] (03CR) 10Bartosz Dziewoński: [C:03+1] wmf-config: Set JSON content model for Web2Cit configuration pages [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1314816 (https://phabricator.wikimedia.org/T305571) (owner: 10Diegodlh) [17:43:56] 06SRE, 10SRE-Access-Requests, 06Release-Engineering-Team, 13Patch-For-Review: Requesting access to Jenkins server for PWangai-WMF - https://phabricator.wikimedia.org/T433563#12188838 (10Dzahn) [17:44:10] (03CR) 10Dzahn: [C:03+2] admin: upgrade pwangai from ldap_only to contint-users shell [puppet] - 10https://gerrit.wikimedia.org/r/1321601 (https://phabricator.wikimedia.org/T433563) (owner: 10Dzahn) [17:48:27] 06SRE, 10SRE-Access-Requests: Requesting access to deployment for Chlod Alejandro - https://phabricator.wikimedia.org/T433791#12188874 (10KFrancis) Hi all, I don't have an NDA on file, so I will process one and confirm when it's complete. Thanks! [17:49:49] PROBLEM - Host ml-serve1015 is DOWN: PING CRITICAL - Packet loss = 100% [17:53:50] FIRING: KubernetesCalicoDown: ml-serve1015.eqiad.wmnet is not running calico-node Pod - https://wikitech.wikimedia.org/wiki/Calico#Operations - https://grafana.wikimedia.org/d/G8zPL7-Wz/?var-dc=eqiad%20prometheus%2Fk8s-mlserve&var-instance=ml-serve1015.eqiad.wmnet - https://alerts.wikimedia.org/?q=alertname%3DKubernetesCalicoDown [17:59:00] FIRING: NodeBGPSessionStatusNotEstablished: Kubernetes node ml-serve1015 has a BGP session which is not in the 'established' ... [17:59:00] state. [18:00:05] jnuche and jeena: Your horoscope predicts another MediaWiki train - Utc-0+Utc-7 Version (secondary timeslot) deploy. May Zuul be (nice) with you. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260805T1800). [18:02:59] (03PS1) 10Scott French: zarcillo: Allow egress to etcd in codfw [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321611 (https://phabricator.wikimedia.org/T434126) [18:04:00] FIRING: [2x] NodeBGPSessionStatusNotEstablished: Kubernetes node ml-serve1015 has a BGP session which is not in the 'established' state. [18:07:25] FIRING: SystemdUnitFailed: dump_proxy_ranges.service on puppetserver2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [18:10:43] (03CR) 10Ahmon Dancy: "Ah, yes, I ran into that too last time I did this. I have an outstanding todo list item to ask hashar about that. I'll get with him tomo" [software/gerrit] (deploy/wmf/stable-3.10) - 10https://gerrit.wikimedia.org/r/1314995 (https://phabricator.wikimedia.org/T240266) (owner: 10Dzahn) [18:13:17] !log robh@cumin2002 START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [18:13:57] 10ops-eqsin, 06SRE, 06DC-Ops, 06Traffic: cp5022 is unreachable - https://phabricator.wikimedia.org/T414411#12188950 (10RobH) >>! In T414411#12188766, @ssingh wrote: >>>! In T414411#12183764, @RobH wrote: >> pre-onsite checks: host mgmt is online and i can connect to it, but I'm having issues on its idrac i... [18:14:45] 10ops-eqiad, 06SRE, 06DC-Ops: Alert for device ps1-e2-eqiad.mgmt.eqiad.wmnet - PDU sensor over limit - https://phabricator.wikimedia.org/T434118#12188952 (10phaultfinder) [18:16:41] (03CR) 10Ssingh: zarcillo: Allow egress to etcd in codfw (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321611 (https://phabricator.wikimedia.org/T434126) (owner: 10Scott French) [18:18:02] (03CR) 10Ssingh: "Happy to merge this tomorrow at around 13:00 UTC?" [puppet] - 10https://gerrit.wikimedia.org/r/1290731 (https://phabricator.wikimedia.org/T425441) (owner: 10Arnaudb) [18:18:16] (03CR) 10CDanis: [C:03+1] zarcillo: Allow egress to etcd in codfw [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321611 (https://phabricator.wikimedia.org/T434126) (owner: 10Scott French) [18:19:33] (03CR) 10Dzahn: [C:03+2] "sounds good! thank you very much" [software/gerrit] (deploy/wmf/stable-3.10) - 10https://gerrit.wikimedia.org/r/1314995 (https://phabricator.wikimedia.org/T240266) (owner: 10Dzahn) [18:19:37] !log vriley@cumin1003 START - Cookbook sre.dns.netbox [18:22:23] !log robh@cumin2002 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [18:22:52] !log vriley@cumin1003 END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) [18:22:59] !log vriley@cumin1003 START - Cookbook sre.dns.netbox [18:23:36] 10ops-eqsin, 06SRE, 06DC-Ops, 06Traffic: cp5022 is unreachable - https://phabricator.wikimedia.org/T414411#12188983 (10RobH) a:03ssingh @ssingh: This is ready for traffic to take it back over. Currently 'planned' in netbox as its been repaired and provision cookbook successfully run. [18:25:23] (03CR) 10Scott French: "Thanks, Sukhbir!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321611 (https://phabricator.wikimedia.org/T434126) (owner: 10Scott French) [18:25:38] 10ops-eqsin, 06SRE, 06DC-Ops, 06Traffic: cp5022 is unreachable - https://phabricator.wikimedia.org/T414411#12188995 (10ssingh) a:05ssingh→03CDobbins [18:26:19] 10ops-eqsin, 06SRE, 06DC-Ops, 06Traffic: cp5022 is unreachable - https://phabricator.wikimedia.org/T414411#12188998 (10ssingh) Thanks Rob. @CDobbins will take care of it from Traffic's end. Appreciate the support here! [18:27:33] !log vriley@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" [18:27:52] (03CR) 10Ssingh: [C:03+1] zarcillo: Allow egress to etcd in codfw (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321611 (https://phabricator.wikimedia.org/T434126) (owner: 10Scott French) [18:28:00] !log vriley@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003" [18:28:00] !log vriley@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [18:28:57] (03CR) 10Scott French: [C:03+2] zarcillo: Allow egress to etcd in codfw [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321611 (https://phabricator.wikimedia.org/T434126) (owner: 10Scott French) [18:30:14] !log vriley@cumin1003 START - Cookbook sre.hosts.provision for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [18:31:15] (03Merged) 10jenkins-bot: zarcillo: Allow egress to etcd in codfw [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321611 (https://phabricator.wikimedia.org/T434126) (owner: 10Scott French) [18:34:04] !log fceratto@deploy1003 helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . [18:34:49] !log fceratto@deploy1003 helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . [18:35:11] vriley@cumin1003 provision (PID 1084972) is awaiting input [18:35:34] !log fceratto@deploy1003 helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . [18:36:21] (03CR) 10Ssingh: "So I have read the task and this CR as well. The reason we are doing this is to support V6 addresses for apt, which uses a discovery recor" [puppet] - 10https://gerrit.wikimedia.org/r/1320177 (https://phabricator.wikimedia.org/T433830) (owner: 10Majavah) [18:36:22] 06SRE, 06collaboration-services, 10Gerrit, 06Release-Engineering-Team (Seen), 07Security: gerrit.wikimedia.org:29418 uses 1024-bit RSA key only - https://phabricator.wikimedia.org/T240266#12189053 (10Dzahn) [18:52:15] !log vriley@cumin1003 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [18:52:55] !log vriley@cumin1003 START - Cookbook sre.hosts.reimage for host zuul1007.eqiad.wmnet with OS trixie [18:53:07] 10ops-eqiad, 06SRE, 06collaboration-services, 06DC-Ops, 13Patch-For-Review: Repurpose ganeti102[3456] for Zuul migration - https://phabricator.wikimedia.org/T427353#12189111 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by vriley@cumin1003 for host zuul1007.eqiad.wmnet with OS t... [18:54:30] 06SRE, 10SRE-Access-Requests, 06Release-Engineering-Team: Requesting access to Jenkins server for PWangai-WMF - https://phabricator.wikimedia.org/T433563#12189116 (10Dzahn) thanks all! @pwangai Alright, here you go! you have access now. Your user exists on these hosts: - bast1004.wikimedia.org in the eqia... [18:55:45] 06SRE, 10SRE-Access-Requests, 06Release-Engineering-Team: Requesting access to Jenkins server for PWangai-WMF - https://phabricator.wikimedia.org/T433563#12189120 (10Dzahn) 05In progress→03Resolved a:03Dzahn claiming it's resolved - always ok to reopen if there are problems [18:57:46] (03CR) 10Scott French: [C:03+2] hieradata: Prepare conf1007 for reimage to bookworm [puppet] - 10https://gerrit.wikimedia.org/r/1321107 (https://phabricator.wikimedia.org/T428495) (owner: 10Scott French) [18:58:18] (03PS1) 10Bking: Revert "cirrussearch: Enable performance governor on master-eligible hosts" [puppet] - 10https://gerrit.wikimedia.org/r/1321620 [18:58:25] (03CR) 10Bking: [V:03+2 C:03+2] Revert "cirrussearch: Enable performance governor on master-eligible hosts" [puppet] - 10https://gerrit.wikimedia.org/r/1321620 (owner: 10Bking) [19:01:50] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for BGerdemann_(WMF) - https://phabricator.wikimedia.org/T433313#12189146 (10Dzahn) @Bethany @leila Thanks as well! We were about to close this ticket out and verify that last open checkbox and then noticed it seems like: You al... [19:03:31] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for BGerdemann_(WMF) - https://phabricator.wikimedia.org/T433313#12189150 (10Dzahn) oooh! I see now in the ticket body it says "level 2" but in the first comment of @Bethany it then refers to "level 3"! is that what confused us? Is... [19:03:45] !log swfrench@cumin2002 START - Cookbook sre.hosts.reboot-single for host conf1007.eqiad.wmnet [19:07:09] !log vriley@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage [19:07:09] (03CR) 10Ahmon Dancy: "I ran "git lfs install --local --skip-repo" to resolve this and I will be working on a change to scap to make this automatic." [software/gerrit] (deploy/wmf/stable-3.10) - 10https://gerrit.wikimedia.org/r/1314995 (https://phabricator.wikimedia.org/T240266) (owner: 10Dzahn) [19:09:26] !log swfrench@cumin2002 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1007.eqiad.wmnet [19:11:01] (03CR) 10Dzahn: [C:03+2] "yay! thanks:)" [software/gerrit] (deploy/wmf/stable-3.10) - 10https://gerrit.wikimedia.org/r/1314995 (https://phabricator.wikimedia.org/T240266) (owner: 10Dzahn) [19:12:49] !log vriley@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage [19:13:48] !log swfrench@cumin2002 START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm [19:19:59] !log silenced EtcdRelicationDown 0cb709a9-f244-4f1e-971f-440ec65e7fd7 - T428495 [19:20:03] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [19:20:04] T428495: Migrate conf* hosts away from bullseye - https://phabricator.wikimedia.org/T428495 [19:23:19] (03CR) 10RLazarus: httpbb: Release v0.0.6 (031 comment) [software/httpbb] - 10https://gerrit.wikimedia.org/r/1321516 (https://phabricator.wikimedia.org/T434052) (owner: 10Blake) [19:28:01] FIRING: [2x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-1/0/2 (Transport: cr1-eqiad:et-1/1/2 (Arelion, IC-374549)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [19:29:42] !log vriley@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" [19:30:02] !log vriley@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" [19:30:03] !log vriley@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1007.eqiad.wmnet with OS trixie [19:30:20] 10ops-eqiad, 06SRE, 06collaboration-services, 06DC-Ops, 13Patch-For-Review: Repurpose ganeti102[3456] for Zuul migration - https://phabricator.wikimedia.org/T427353#12189249 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by vriley@cumin1003 for host zuul1007.eqiad.wmnet with OS trixi... [19:39:53] (03CR) 10Dzahn: [C:03+2] "ssh key verified via email" [puppet] - 10https://gerrit.wikimedia.org/r/1321601 (https://phabricator.wikimedia.org/T433563) (owner: 10Dzahn) [19:42:45] RECOVERY - OSPF status on cr1-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [19:42:51] RECOVERY - OSPF status on cr2-eqsin is OK: OSPFv2: 5/5 UP : OSPFv3: 5/5 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [19:44:05] 06SRE, 10LDAP-Access-Requests: Grant Access to ciadmin for Vaughn Walters - https://phabricator.wikimedia.org/T433615#12189313 (10Dzahn) @hashar Seeing the shell groups I am wondering if this is also needed here: ` deployment-ci-admins: gid: 721 description: users who can perform administ... [19:45:37] 06SRE, 10SRE-Access-Requests: Requesting access to deployment for Chlod Alejandro - https://phabricator.wikimedia.org/T433791#12189318 (10Dzahn) @dancy This request would also need approval from the group approver for `deployment` which lists you and Tyler. [19:46:25] (03CR) 10RLazarus: [C:03+1] hieradata: Enable the v2 API and set cluster_bootstrap false in eqiad [puppet] - 10https://gerrit.wikimedia.org/r/1321109 (https://phabricator.wikimedia.org/T428495) (owner: 10Scott French) [19:46:34] !log vriley@cumin1003 START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [19:47:26] !log bking@cumin2003 DONE (FAIL) - Cookbook sre.puppet.renew-cert (exit_code=99) for an-worker1189.eqiad.wmnet: Renew puppet certificate - bking@cumin2003 [19:49:44] 06SRE, 10SRE-Access-Requests: Grant access to analytics-privatedata-users for EAlbizzati-WMF - https://phabricator.wikimedia.org/T434034#12189330 (10Dzahn) Hi @EAlbizzati-WMF regarding the "wmf" group part of your request. This is nowadays a self-service thing. Please see the bottom part of https://wikitec... [19:51:15] vriley@cumin1003 provision (PID 1094045) is awaiting input [19:51:17] !log [bking@puppetserver1001] ~$ sudo puppetserver ca sign --certname an-worker1189.eqiad.wmnet T434142 [19:51:20] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [19:51:20] T434142: Figure out what to do with an-worker1189 - https://phabricator.wikimedia.org/T434142 [19:52:00] 06SRE, 10SRE-Access-Requests: Requesting access to deployment for Chlod Alejandro - https://phabricator.wikimedia.org/T433791#12189336 (10dancy) >>! In T433791#12189317, @Dzahn wrote: > @dancy This request would also need approval from the group approver for `deployment` which lists you and Tyler. Approved. [20:00:05] RoanKattouw, urbanecm, TheresNoTime, kindrobot, and cjming: It is that lovely time of the day again! You are hereby commanded to deploy UTC late backport window. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260805T2000). [20:00:05] James_F and cjming: A patch you scheduled for UTC late backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [20:00:23] o/ [20:01:08] Heya. [20:01:14] cjming: Did you want to go first? [20:01:17] James_F: if you're around, please go ahead with you're patches and ping me when you're done [20:01:23] Ack, I’ll do mine now then. [20:01:29] sounds good [20:01:53] (03CR) 10TrainBranchBot: [C:03+2] "Approved by jforrester@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320196 (https://phabricator.wikimedia.org/T428911) (owner: 10Jforrester) [20:01:53] (03CR) 10TrainBranchBot: [C:03+2] "Approved by jforrester@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311106 (owner: 10Jforrester) [20:01:54] (03CR) 10TrainBranchBot: [C:03+2] "Approved by jforrester@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1321585 (owner: 10Jforrester) [20:03:01] (03Merged) 10jenkins-bot: ExtensionDistributor: Drop REL1_44, EOL [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320196 (https://phabricator.wikimedia.org/T428911) (owner: 10Jforrester) [20:03:05] (03Merged) 10jenkins-bot: wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311106 (owner: 10Jforrester) [20:03:09] (03Merged) 10jenkins-bot: abstractwiki: Add dag/ml/ig/ha languages to generation script [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1321585 (owner: 10Jforrester) [20:03:28] !log jforrester@deploy1003 Started scap sync-world: Backport for [[gerrit:1320196|ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106|wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585|abstractwiki: Add dag/ml/ig/ha languages to generation script]] [20:03:32] T428911: Formally EOL MW 1.44 - https://phabricator.wikimedia.org/T428911 [20:05:19] swfrench@cumin2002 reimage (PID 117573) is awaiting input [20:07:23] !log jforrester@deploy1003 jforrester: Backport for [[gerrit:1320196|ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106|wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585|abstractwiki: Add dag/ml/ig/ha languages to generation script]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [20:08:05] !log jforrester@deploy1003 jforrester: Continuing with deployment [20:08:07] !log swfrench@cumin2002 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host conf1007.eqiad.wmnet with OS bookworm [20:12:10] !log jforrester@deploy1003 Finished scap sync-world: Backport for [[gerrit:1320196|ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106|wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585|abstractwiki: Add dag/ml/ig/ha languages to generation script]] (duration: 08m 41s) [20:12:14] cjming: Over to you. [20:12:14] T428911: Formally EOL MW 1.44 - https://phabricator.wikimedia.org/T428911 [20:12:22] James_F: thanks! [20:12:30] !log swfrench@cumin2002 START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm [20:12:48] (03CR) 10TrainBranchBot: [C:03+2] "Approved by cjming@deploy1003 using scap backport" [extensions/WikiLambda] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1321591 (https://phabricator.wikimedia.org/T433550) (owner: 10Clare Ming) [20:16:10] (03Merged) 10jenkins-bot: logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 [extensions/WikiLambda] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1321591 (https://phabricator.wikimedia.org/T433550) (owner: 10Clare Ming) [20:16:32] !log cjming@deploy1003 Started scap sync-world: Backport for [[gerrit:1321591|logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] [20:16:37] T433550: [Abstract Wikipedia] Events are not being sent from migrated wikifunctions-ui-actions instrument - https://phabricator.wikimedia.org/T433550 [20:18:01] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [20:18:30] !log cjming@deploy1003 cjming: Backport for [[gerrit:1321591|logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [20:18:56] !log cjming@deploy1003 cjming: Continuing with deployment [20:20:29] 06SRE, 06Infrastructure-Foundations, 10netops: Nokia SR-Linux: Update Homer automation to support v26 - https://phabricator.wikimedia.org/T433105#12189415 (10ayounsi) 05Open→03Resolved a:03ayounsi All done! [20:22:58] !log cjming@deploy1003 Finished scap sync-world: Backport for [[gerrit:1321591|logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] (duration: 06m 26s) [20:23:02] T433550: [Abstract Wikipedia] Events are not being sent from migrated wikifunctions-ui-actions instrument - https://phabricator.wikimedia.org/T433550 [20:24:29] !log end of UTC late backport window [20:24:32] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [20:27:42] !log swfrench@cumin2002 START - Cookbook sre.hosts.downtime for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage [20:34:17] !log swfrench@cumin2002 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage [20:38:19] 06SRE, 10SRE-Access-Requests: Requesting access to deployment for Chlod Alejandro - https://phabricator.wikimedia.org/T433791#12189451 (10Dzahn) [20:38:41] 06SRE, 10SRE-Access-Requests: Requesting access to deployment for Chlod Alejandro - https://phabricator.wikimedia.org/T433791#12189452 (10Dzahn) [20:43:16] !log T434008: changing cloudelastic:9643 from auto_expand_replicas to number_of_replicas [20:43:19] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [20:43:20] T434008: Cirrussearch: Ensure master-eligibles can restart without losing quorum - https://phabricator.wikimedia.org/T434008 [20:51:12] !log vriley@cumin1003 START - Cookbook sre.dns.netbox [20:55:53] !log swfrench@cumin2002 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1007.eqiad.wmnet with OS bookworm [20:55:54] !log vriley@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" [20:55:59] !log vriley@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003" [20:55:59] !log vriley@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [20:56:39] !log vriley@cumin1003 START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [20:59:25] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [21:00:05] Deploy window Wikifunctions Services UTC Late (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260805T2100) [21:00:05] !log vriley@cumin1003 END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [21:02:57] !log vriley@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 [21:03:57] !log vriley@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 [21:04:14] !log vriley@cumin1003 START - Cookbook sre.dns.netbox [21:07:01] !log bking@cumin2003 conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad [21:07:01] (03CR) 10Scott French: [C:03+2] hieradata: Temporarily move etcd replication to conf1007 [puppet] - 10https://gerrit.wikimedia.org/r/1321108 (https://phabricator.wikimedia.org/T428495) (owner: 10Scott French) [21:07:25] !log vriley@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [21:08:01] (03PS2) 10Scott French: hieradata: Temporarily move etcd replication to conf1007 [puppet] - 10https://gerrit.wikimedia.org/r/1321108 (https://phabricator.wikimedia.org/T428495) [21:08:01] (03PS3) 10Scott French: hieradata: Enable the v2 API and set cluster_bootstrap false in eqiad [puppet] - 10https://gerrit.wikimedia.org/r/1321109 (https://phabricator.wikimedia.org/T428495) [21:08:37] !log vriley@cumin1003 START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [21:09:12] (03CR) 10Scott French: [C:03+2] hieradata: Temporarily move etcd replication to conf1007 [puppet] - 10https://gerrit.wikimedia.org/r/1321108 (https://phabricator.wikimedia.org/T428495) (owner: 10Scott French) [21:10:49] !log bking@cumin2003 conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad [21:12:37] !log bking@cumin2003 conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad [21:12:37] !log bking@cumin2003 conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad [21:13:25] vriley@cumin1003 provision (PID 1113829) is awaiting input [21:14:42] !log vriley@cumin1003 END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [21:18:57] FIRING: [2x] JobUnavailable: Reduced availability for job etcdmirror in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [21:22:28] 06SRE, 10SRE-Access-Requests: Grant access to analytics-privatedata-users for EAlbizzati-WMF - https://phabricator.wikimedia.org/T434034#12189529 (10EAlbizzati-WMF) hey @Dzahn I already have the wmf access. I think I only need the `analytics-privatedata-users` level 1 access now. thank you [21:23:01] FIRING: [2x] JobUnavailable: Reduced availability for job etcdmirror in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [21:24:44] (03CR) 10Scott French: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1321109 (https://phabricator.wikimedia.org/T428495) (owner: 10Scott French) [21:27:03] (03CR) 10Scott French: [C:03+2] hieradata: Enable the v2 API and set cluster_bootstrap false in eqiad [puppet] - 10https://gerrit.wikimedia.org/r/1321109 (https://phabricator.wikimedia.org/T428495) (owner: 10Scott French) [21:36:02] !log swfrench@cumin2002 START - Cookbook sre.hosts.reboot-single for host conf1008.eqiad.wmnet [21:36:04] (03CR) 10Lerickson: [C:03+1] WDQSv2: Qlever index rebuild container [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320946 (https://phabricator.wikimedia.org/T432627) (owner: 10Trueg) [21:41:46] (03CR) 10Lerickson: [C:03+1] WDQSv2: Qlever index rebuild container (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320946 (https://phabricator.wikimedia.org/T432627) (owner: 10Trueg) [21:41:47] !log swfrench@cumin2002 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1008.eqiad.wmnet [21:43:06] !log vriley@cumin1003 START - Cookbook sre.dns.netbox [21:44:43] !log swfrench@cumin2002 START - Cookbook sre.hosts.reimage for host conf1008.eqiad.wmnet with OS bookworm [21:46:13] !log vriley@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [21:47:16] !log vriley@cumin1003 START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [21:48:30] !log vriley@cumin1003 END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [21:49:23] (03PS1) 10LorenMora: Turn on feature flag for custom lists for betawiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1321639 (https://phabricator.wikimedia.org/T434027) [21:54:05] FIRING: KubernetesCalicoDown: ml-serve1015.eqiad.wmnet is not running calico-node Pod - https://wikitech.wikimedia.org/wiki/Calico#Operations - https://grafana.wikimedia.org/d/G8zPL7-Wz/?var-dc=eqiad%20prometheus%2Fk8s-mlserve&var-instance=ml-serve1015.eqiad.wmnet - https://alerts.wikimedia.org/?q=alertname%3DKubernetesCalicoDown [22:00:05] Deploy window Readers deployment window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260805T2200) [22:00:55] !log swfrench@cumin2002 START - Cookbook sre.hosts.downtime for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage [22:04:00] FIRING: [2x] NodeBGPSessionStatusNotEstablished: Kubernetes node ml-serve1015 has a BGP session which is not in the 'established' state. [22:04:27] !log swfrench@cumin2002 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage [22:07:25] FIRING: SystemdUnitFailed: dump_proxy_ranges.service on puppetserver2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [22:25:09] !log swfrench@cumin2002 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1008.eqiad.wmnet with OS bookworm [22:34:14] !log swfrench@cumin2002 START - Cookbook sre.hosts.reboot-single for host conf1009.eqiad.wmnet [22:38:33] !log swfrench@cumin2002 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1009.eqiad.wmnet [22:43:29] !log swfrench@cumin2002 START - Cookbook sre.hosts.reimage for host conf1009.eqiad.wmnet with OS bookworm [22:59:18] 10SRE-swift-storage, 06Commons: file missing after move - https://phabricator.wikimedia.org/T430561#12189807 (10Pppery) 05Open→03Resolved The file seems to be back now. No idea what happened. [22:59:41] !log swfrench@cumin2002 START - Cookbook sre.hosts.downtime for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage [23:00:53] (03CR) 10Scott French: [C:03+1] "Thanks, Blake!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321561 (https://phabricator.wikimedia.org/T427668) (owner: 10Blake) [23:01:19] (03PS1) 10Dzahn: admin: add ealbizzati to analytics-privatedata-users, level 1 access [puppet] - 10https://gerrit.wikimedia.org/r/1321651 (https://phabricator.wikimedia.org/T434034) [23:02:13] (03CR) 10CI reject: [V:04-1] admin: add ealbizzati to analytics-privatedata-users, level 1 access [puppet] - 10https://gerrit.wikimedia.org/r/1321651 (https://phabricator.wikimedia.org/T434034) (owner: 10Dzahn) [23:03:39] !log swfrench@cumin2002 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage [23:03:47] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Grant access to analytics-privatedata-users for EAlbizzati-WMF - https://phabricator.wikimedia.org/T434034#12189831 (10Dzahn) @EAlbizzati-WMF gotcha! You already said that on the original ticket, my bad. Uploading the needed code change to review to move th... [23:05:51] (03PS2) 10Dzahn: admin: add ealbizzati to analytics-privatedata-users, level 1 access [puppet] - 10https://gerrit.wikimedia.org/r/1321651 (https://phabricator.wikimedia.org/T434034) [23:06:50] (03PS1) 10Ahmon Dancy: scap: Add optional logstash credentials file [puppet] - 10https://gerrit.wikimedia.org/r/1321653 (https://phabricator.wikimedia.org/T434114) [23:07:22] 10SRE-swift-storage, 10UploadWizard: Internal error: The server could not save the temporary file - https://phabricator.wikimedia.org/T353068#12189837 (10Pppery) 05Open→03Invalid This is 3 years old and lots of changes have been made to the uploading stack since then. Boldly closing as obsolete - file... [23:07:53] 10SRE-swift-storage, 06Commons, 10MediaWiki-File-management: Some or all of the undeletion failed: The file "mwstore://local-multiwrite/local-public/8/8c/Place_de_Chambésy.JPG" is in an inconsistent state within the internal storage backends - https://phabricator.wikimedia.org/T348937#12189839 (10Pppery) [23:13:12] 06SRE, 10SRE-swift-storage, 10MediaWiki-extensions-Score: Add cache key information to metadata json - https://phabricator.wikimedia.org/T257093#12189844 (10TheDJ) @MatthewVernon i mean deleting old score files, specifically their png version, as we mostly use svg now since a couple of weeks, so left unatten... [23:14:26] (03CR) 10Scott French: httpbb: Release v0.0.6 (031 comment) [software/httpbb] - 10https://gerrit.wikimedia.org/r/1321516 (https://phabricator.wikimedia.org/T434052) (owner: 10Blake) [23:24:25] !log swfrench@cumin2002 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1009.eqiad.wmnet with OS bookworm [23:28:01] FIRING: [2x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-1/0/2 (Transport: cr1-eqiad:et-1/1/2 (Arelion, IC-374549)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [23:29:41] (03PS1) 10Jforrester: wikifunctions: Lower ORCHESTRATOR_HEAP_SIZE for staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321658 (https://phabricator.wikimedia.org/T433506) [23:30:03] (03Abandoned) 10TrainBranchBot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1321156 (owner: 10TrainBranchBot) [23:33:16] (03PS1) 10Scott French: Revert "hieradata: Temporarily move etcd replication to conf1007" [puppet] - 10https://gerrit.wikimedia.org/r/1321660 (https://phabricator.wikimedia.org/T428495) [23:34:16] (03CR) 10Scott French: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1321660 (https://phabricator.wikimedia.org/T428495) (owner: 10Scott French) [23:39:25] (03CR) 10Scott French: [C:03+2] Revert "hieradata: Temporarily move etcd replication to conf1007" [puppet] - 10https://gerrit.wikimedia.org/r/1321660 (https://phabricator.wikimedia.org/T428495) (owner: 10Scott French) [23:42:33] (03PS1) 10TrainBranchBot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1321663 [23:42:33] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1321663 (owner: 10TrainBranchBot) [23:49:15] (03PS1) 10Scott French: Revert "Temporarily point all etcd client SRV records to codfw" [dns] - 10https://gerrit.wikimedia.org/r/1321667 (https://phabricator.wikimedia.org/T428495) [23:49:24] (03PS1) 10Scott French: Revert "hieradata: Temporarily point eqiad PyBals at codfw etcd" [puppet] - 10https://gerrit.wikimedia.org/r/1321668 (https://phabricator.wikimedia.org/T428495) [23:49:41] (03PS2) 10Scott French: Revert "hieradata: Temporarily point eqiad PyBals at codfw etcd" [puppet] - 10https://gerrit.wikimedia.org/r/1321668 (https://phabricator.wikimedia.org/T428495) [23:50:15] (03CR) 10Scott French: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1321668 (https://phabricator.wikimedia.org/T428495) (owner: 10Scott French) [23:52:30] (03PS1) 10Cwhite: centralserver: add utility to delete oldest files based on disk usage [puppet] - 10https://gerrit.wikimedia.org/r/1321669 [23:54:26] (03Merged) 10jenkins-bot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1321663 (owner: 10TrainBranchBot) [23:59:18] (03CR) 10Scott French: "I will follow up some time after 14:00 UTC tomorrow to move this forward." [puppet] - 10https://gerrit.wikimedia.org/r/1321668 (https://phabricator.wikimedia.org/T428495) (owner: 10Scott French) [23:59:23] (03CR) 10Scott French: "I will follow up some time after 14:00 UTC tomorrow to move this forward." [dns] - 10https://gerrit.wikimedia.org/r/1321667 (https://phabricator.wikimedia.org/T428495) (owner: 10Scott French)