[00:02:25] RESOLVED: SystemdUnitFailed: dump_proxy_ranges.service on puppetserver2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [00:59:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [01:12:49] (03PS1) 10TrainBranchBot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1321672 [01:12:49] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1321672 (owner: 10TrainBranchBot) [01:23:01] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [01:23:33] (03Merged) 10jenkins-bot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1321672 (owner: 10TrainBranchBot) [01:54:05] FIRING: KubernetesCalicoDown: ml-serve1015.eqiad.wmnet is not running calico-node Pod - https://wikitech.wikimedia.org/wiki/Calico#Operations - https://grafana.wikimedia.org/d/G8zPL7-Wz/?var-dc=eqiad%20prometheus%2Fk8s-mlserve&var-instance=ml-serve1015.eqiad.wmnet - https://alerts.wikimedia.org/?q=alertname%3DKubernetesCalicoDown [02:00:49] !log mwpresync@deploy1003 Started scap build-images: Publishing wmf/next image [02:04:15] FIRING: [2x] NodeBGPSessionStatusNotEstablished: Kubernetes node ml-serve1015 has a BGP session which is not in the 'established' state. [02:07:29] !log mwpresync@deploy1003 Finished scap build-images: Publishing wmf/next image (duration: 06m 40s) [02:08:01] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:13:57] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:18:01] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [03:28:01] FIRING: [2x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-1/0/2 (Transport: cr1-eqiad:et-1/1/2 (Arelion, IC-374549)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [04:59:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [05:54:05] FIRING: KubernetesCalicoDown: ml-serve1015.eqiad.wmnet is not running calico-node Pod - https://wikitech.wikimedia.org/wiki/Calico#Operations - https://grafana.wikimedia.org/d/G8zPL7-Wz/?var-dc=eqiad%20prometheus%2Fk8s-mlserve&var-instance=ml-serve1015.eqiad.wmnet - https://alerts.wikimedia.org/?q=alertname%3DKubernetesCalicoDown [05:56:38] (03CR) 10Trueg: WDQSv2: Qlever index rebuild container (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320946 (https://phabricator.wikimedia.org/T432627) (owner: 10Trueg) [06:00:05] Deploy window MediaWiki infrastructure (UTC early) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260806T0600) [06:00:05] marostegui, Amir1, and federico3: OwO what's this, a deployment window?? Primary database switchover. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260806T0600). nyaa~ [06:04:15] FIRING: [2x] NodeBGPSessionStatusNotEstablished: Kubernetes node ml-serve1015 has a BGP session which is not in the 'established' state. [06:18:01] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [06:25:11] (03PS4) 10Trueg: WDQSv2: Qlever index rebuild container [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320946 (https://phabricator.wikimedia.org/T432627) [06:34:25] (03PS1) 10Ayounsi: eqsin transports: active/active [homer/public] - 10https://gerrit.wikimedia.org/r/1321816 [06:36:13] (03PS1) 10Ayounsi: Make all POPs transports the same OSPF metric [homer/public] - 10https://gerrit.wikimedia.org/r/1321817 (https://phabricator.wikimedia.org/T424839) [06:41:04] (03CR) 10Cathal Mooney: "LGTM!" [homer/public] - 10https://gerrit.wikimedia.org/r/1321816 (owner: 10Ayounsi) [06:45:42] (03Abandoned) 10Ayounsi: Make all POPs transports the same OSPF metric [homer/public] - 10https://gerrit.wikimedia.org/r/1321817 (https://phabricator.wikimedia.org/T424839) (owner: 10Ayounsi) [06:46:05] (03CR) 10Ayounsi: [C:03+2] eqsin transports: active/active [homer/public] - 10https://gerrit.wikimedia.org/r/1321816 (owner: 10Ayounsi) [06:47:46] (03Merged) 10jenkins-bot: eqsin transports: active/active [homer/public] - 10https://gerrit.wikimedia.org/r/1321816 (owner: 10Ayounsi) [06:49:14] (03CR) 10Arnaudb: [C:03+1] "sounds good to me! thanks!" [puppet] - 10https://gerrit.wikimedia.org/r/1290731 (https://phabricator.wikimedia.org/T425441) (owner: 10Arnaudb) [06:53:32] PROBLEM - SSH on urldownloader1006 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/SSH/monitoring [06:54:42] FIRING: [2x] ProbeDown: Service urldownloader1006:8080 has failed probes (http_url_downloader_wikimedia_org_ip4) - https://wikitech.wikimedia.org/wiki/Url-downloader - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [07:00:05] Amir1, urbanecm, and awight: Time to snap out of that daydream and deploy UTC morning backport window. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260806T0700). [07:00:05] No Gerrit patches in the queue for this window AFAICS. [07:05:36] 10SRE-Access-Requests, 06Infrastructure-Foundations, 10LDAP-Access-Requests, 13Patch-For-Review: Superset data access request for abibendall - https://phabricator.wikimedia.org/T430938#12190296 (10jcrespo) 05In progress→03Stalled This is still waiting on user feedback. [07:06:36] !log marostegui@cumin1003 START - Cookbook sre.mysql.depool depool db1153: Cloning [07:06:37] !log marostegui@cumin1003 START - Cookbook sre.mysql.parsercache [07:07:36] !log marostegui@cumin1003 END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) [07:07:37] !log marostegui@cumin1003 END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Cloning [07:08:05] !log marostegui@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2252.codfw.wmnet,db1153.eqiad.wmnet with reason: cloning [07:08:48] (03PS1) 10Marostegui: mariadb: Productionize db1268 [puppet] - 10https://gerrit.wikimedia.org/r/1321819 (https://phabricator.wikimedia.org/T407942) [07:09:31] (03PS2) 10JMeybohm: coredns: Rename coredns to coredns1.11 to support multiple versions [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1321573 (https://phabricator.wikimedia.org/T433590) [07:09:31] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for iberker - https://phabricator.wikimedia.org/T433258#12190303 (10jcrespo) > I don't really need production server access No worries. No SSH access is scheduled, but this is a privileged LDAP access request, and it requires speci... [07:09:33] (03PS2) 10JMeybohm: Add coredns 1.12 [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1321574 (https://phabricator.wikimedia.org/T433590) [07:10:09] (03CR) 10Marostegui: [C:03+2] mariadb: Productionize db1268 [puppet] - 10https://gerrit.wikimedia.org/r/1321819 (https://phabricator.wikimedia.org/T407942) (owner: 10Marostegui) [07:15:15] (03CR) 10JMeybohm: [C:03+1] admin_ng: pin kube-state-metrics chart to current version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320945 (https://phabricator.wikimedia.org/T427405) (owner: 10Blake) [07:16:21] 06SRE, 10SRE-Access-Requests, 10LDAP-Access-Requests: Grant Access to wmf for Chandler Diggs - https://phabricator.wikimedia.org/T433302#12190310 (10jcrespo) 05In progress→03Resolved It seems the access has been deployed already. @CDiggs-WMF didn't reported of any issue after last message. Clinic dut... [07:16:59] 06SRE, 10SRE-Access-Requests, 10LDAP-Access-Requests: Grant Access to wmf for Chandler Diggs - https://phabricator.wikimedia.org/T433302#12190312 (10jcrespo) a:05CDiggs-WMF→03Dzahn [07:22:30] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for BGerdemann_(WMF) - https://phabricator.wikimedia.org/T433313#12190319 (10jcrespo) a:03Bethany @Bethany could you provide feedback to @Dzahn 's latest comment? If level 3 is required, could you update the original request accor... [07:28:01] FIRING: [2x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-1/0/2 (Transport: cr1-eqiad:et-1/1/2 (Arelion, IC-374549)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [07:32:35] (03PS1) 10Arthur taylor: wikidata-query-gui: bump to latest version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321822 (https://phabricator.wikimedia.org/T433277) [07:34:53] (03PS1) 10Clare Ming: Test Kitchen UI: Deploy v1.5.2 release to staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321824 (https://phabricator.wikimedia.org/T432629) [07:37:41] !log updated istio to 1.29.4 on wikikube codfw - T427401 [07:37:45] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [07:37:46] T427401: Update istio to 1.29 - https://phabricator.wikimedia.org/T427401 [07:41:04] (03PS1) 10Clare Ming: Test Kitchen UI: Deploy v1.5.2 release to production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321827 (https://phabricator.wikimedia.org/T432629) [07:46:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [07:47:09] 06SRE, 10SRE-Access-Requests: Requesting access to deployment for Chlod Alejandro - https://phabricator.wikimedia.org/T433791#12190367 (10jcrespo) @Chlod while NDA approval is ongoing, do you mind generating a new gerrit PR (I believe it won't be hard for you) on the operations/puppet repo, path production/mod... [07:48:36] !log bwojtowicz@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . [07:51:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.62% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [07:51:58] !log bwojtowicz@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . [07:53:04] !log bwojtowicz@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . [07:54:51] !log bwojtowicz@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . [07:56:49] !log bwojtowicz@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . [07:57:50] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Grant access to analytics-privatedata-users for EAlbizzati-WMF - https://phabricator.wikimedia.org/T434034#12190399 (10jcrespo) @EAlbizzati-WMF The patch @Dzahn created is ready, but because this has an extra server privileged account, this is now only pendi... [07:58:20] !log bwojtowicz@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . [07:59:29] (03CR) 10Jcrespo: [C:04-1] "Virtual +1 for the patch itself, but voting -1 pending on Seddon's approval on ticket." [puppet] - 10https://gerrit.wikimedia.org/r/1321651 (https://phabricator.wikimedia.org/T434034) (owner: 10Dzahn) [08:00:05] jnuche and jeena: #bothumor My software never has bugs. It just develops random features. Rise for MediaWiki train - Utc-0+Utc-7 Version. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260806T0800). [08:00:10] !log bwojtowicz@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . [08:00:17] hi there, train will roll out in a few minutes [08:00:57] !log bwojtowicz@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . [08:06:39] (03PS1) 10TrainBranchBot: group2 to 1.47.0-wmf.14 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1321882 (https://phabricator.wikimedia.org/T430833) [08:06:41] (03CR) 10TrainBranchBot: [C:03+2] "Initiated by jnuche@deploy1003" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1321882 (https://phabricator.wikimedia.org/T430833) (owner: 10TrainBranchBot) [08:07:23] (03PS3) 10JMeybohm: Update to v3.30.7 [debs/calico] (v3.30) - 10https://gerrit.wikimedia.org/r/1305139 (https://phabricator.wikimedia.org/T427400) (owner: 10Jelto) [08:07:57] (03Merged) 10jenkins-bot: group2 to 1.47.0-wmf.14 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1321882 (https://phabricator.wikimedia.org/T430833) (owner: 10TrainBranchBot) [08:09:40] (03CR) 10JMeybohm: "I only now realized this change was initially based off the wrong branch (main instead of v2.29). For some reason I had created the v3.30 " [debs/calico] (v3.30) - 10https://gerrit.wikimedia.org/r/1305139 (https://phabricator.wikimedia.org/T427400) (owner: 10Jelto) [08:12:01] 06SRE, 10LDAP-Access-Requests: Grant Access to NDA for jaleman-vdr-wmf - https://phabricator.wikimedia.org/T433417#12190438 (10jcrespo) Waiting for @KFrancis to ok, to proceed. When she give the ok, we will onboard @JAATPH on the wmf group- that will have to be self-requested through bitu, not ticket (I will... [08:14:07] !log jnuche@deploy1003 rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.14 refs T430833 [08:14:11] T430833: 1.47.0-wmf.14 deployment blockers - https://phabricator.wikimedia.org/T430833 [08:20:39] (03PS1) 10Bartosz Wójtowicz: ml-services: Remove revertrisk-multilingual from experimental staging. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321883 (https://phabricator.wikimedia.org/T431089) [08:29:37] (03CR) 10Klausman: [C:03+2] namespaces: webrequest-pageview [puppet] - 10https://gerrit.wikimedia.org/r/1320823 (https://phabricator.wikimedia.org/T433962) (owner: 10JavierMonton) [08:29:40] !log push pfw policy - T434115 [08:29:42] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [08:30:06] (03CR) 10Klausman: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1320823 (https://phabricator.wikimedia.org/T433962) (owner: 10JavierMonton) [08:35:52] (03PS1) 10Klausman: dse-k8s: add webrequest-pageview{,-next} namespaces [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321884 (https://phabricator.wikimedia.org/T433962) [08:36:36] (03CR) 10Jcrespo: [C:04-1] "I will be testing fabfur's proposed approach and will soon send an amend. We have a month to implement it, so it may take a while, bit it " [puppet] - 10https://gerrit.wikimedia.org/r/1321123 (https://phabricator.wikimedia.org/T430023) (owner: 10Jcrespo) [08:39:10] !log jynus@cumin1003 START - Cookbook sre.hosts.decommission for hosts backup1003.eqiad.wmnet [08:40:48] (03CR) 10JavierMonton: [C:03+1] dse-k8s: add webrequest-pageview{,-next} namespaces [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321884 (https://phabricator.wikimedia.org/T433962) (owner: 10Klausman) [08:43:33] (03CR) 10Klausman: [C:03+2] dse-k8s: add webrequest-pageview{,-next} namespaces [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321884 (https://phabricator.wikimedia.org/T433962) (owner: 10Klausman) [08:43:42] (03CR) 10Klausman: [V:03+2 C:03+2] dse-k8s: add webrequest-pageview{,-next} namespaces [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321884 (https://phabricator.wikimedia.org/T433962) (owner: 10Klausman) [08:46:37] !log jynus@cumin1003 START - Cookbook sre.dns.netbox [08:51:05] 06SRE, 06MediaWiki-Core-Platform-Team, 05Front-end Modernization: Introduce a Front-end Build Step for MediaWiki Skins and Extensions - https://phabricator.wikimedia.org/T279108#12190486 (10jcrespo) Trying to triage this from the SRE side, my guess is they will need either feedback or resources (servers) for... [08:53:04] (03Merged) 10jenkins-bot: dse-k8s: add webrequest-pageview{,-next} namespaces [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321884 (https://phabricator.wikimedia.org/T433962) (owner: 10Klausman) [08:53:28] !log jynus@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" [08:54:23] !log marostegui@cumin1003 dbctl commit (dc=all): 'Pool back ms3', diff saved to https://phabricator.wikimedia.org/P95925 and previous config saved to /var/cache/conftool/dbconfig/20260806-085422-marostegui.json [08:55:24] !log jynus@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" [08:55:24] !log jynus@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [08:55:24] (03PS23) 10Slyngshede: varnish: Configure text for thumbnails [puppet] - 10https://gerrit.wikimedia.org/r/1306896 (https://phabricator.wikimedia.org/T427465) [08:55:25] !log jynus@cumin1003 END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup1003.eqiad.wmnet [08:55:54] (03PS5) 10Jcrespo: bacula: Remove last references to backup1003 & backup2003 [puppet] - 10https://gerrit.wikimedia.org/r/1320928 (https://phabricator.wikimedia.org/T420506) [08:57:07] !log jynus@cumin1003 START - Cookbook sre.hosts.decommission for hosts backup2003.codfw.wmnet [08:57:37] 06SRE, 06Infrastructure-Foundations: Migrate diffscan VM to Trixie - https://phabricator.wikimedia.org/T415347#12190513 (10ayounsi) I think there is something wrong with the new diffscan version, it does seem to run properly but it haven't sent any email since the upgrade, which is very suspicious. [08:58:23] RECOVERY - SSH on urldownloader1006 is OK: SSH OK - OpenSSH_10.0p2 Debian-7+deb13u4 (protocol 2.0) https://wikitech.wikimedia.org/wiki/SSH/monitoring [08:59:01] (03CR) 10Kevin Bazira: [C:03+1] ml-services: Remove revertrisk-multilingual from experimental staging. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321883 (https://phabricator.wikimedia.org/T431089) (owner: 10Bartosz Wójtowicz) [08:59:04] (03CR) 10Jcrespo: [C:03+2] bacula: Remove last references to backup1003 & backup2003 [puppet] - 10https://gerrit.wikimedia.org/r/1320928 (https://phabricator.wikimedia.org/T420506) (owner: 10Jcrespo) [08:59:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:01:06] 10ops-eqiad, 10bacula, 06Data-Persistence, 10Data-Persistence-Backup, and 2 others: decommission backup1003.eqiad.wmnet - https://phabricator.wikimedia.org/T433970#12190518 (10jcrespo) a:05jcrespo→03None [09:01:51] 10ops-eqiad, 10bacula, 06Data-Persistence, 10Data-Persistence-Backup, and 2 others: decommission backup1003.eqiad.wmnet - https://phabricator.wikimedia.org/T433970#12190536 (10jcrespo) This is ready for dcops to unrack or recycling. Thank you! [09:02:00] !log klausman@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. [09:02:26] !log jynus@cumin1003 START - Cookbook sre.dns.netbox [09:02:31] (03PS24) 10Slyngshede: varnish: Configure text for thumbnails [puppet] - 10https://gerrit.wikimedia.org/r/1306896 (https://phabricator.wikimedia.org/T427465) [09:03:37] !log klausman@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. [09:04:10] (03PS1) 10JMeybohm: calico: Add support for v3.30, drop v3.23 [puppet] - 10https://gerrit.wikimedia.org/r/1321888 (https://phabricator.wikimedia.org/T427400) [09:04:12] (03PS1) 10JMeybohm: wikikube: Update staging codfw to calico 3.30 [puppet] - 10https://gerrit.wikimedia.org/r/1321889 (https://phabricator.wikimedia.org/T427400) [09:04:42] RESOLVED: [2x] ProbeDown: Service urldownloader1006:8080 has failed probes (http_url_downloader_wikimedia_org_ip4) - https://wikitech.wikimedia.org/wiki/Url-downloader - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [09:05:01] (03CR) 10Bartosz Wójtowicz: [C:03+2] ml-services: Remove revertrisk-multilingual from experimental staging. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321883 (https://phabricator.wikimedia.org/T431089) (owner: 10Bartosz Wójtowicz) [09:06:58] !log jynus@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" [09:07:13] !log jynus@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003" [09:07:14] !log jynus@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [09:07:15] !log jynus@cumin1003 END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup2003.codfw.wmnet [09:07:27] 10ops-codfw, 10bacula, 06Data-Persistence, 10Data-Persistence-Backup, and 2 others: decommission backup2003.codfw.wmnet - https://phabricator.wikimedia.org/T433971#12190561 (10jcrespo) a:05jcrespo→03None [09:07:42] (03Merged) 10jenkins-bot: ml-services: Remove revertrisk-multilingual from experimental staging. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321883 (https://phabricator.wikimedia.org/T431089) (owner: 10Bartosz Wójtowicz) [09:08:10] 10ops-codfw, 10bacula, 06Data-Persistence, 10Data-Persistence-Backup, and 2 others: decommission backup2003.codfw.wmnet - https://phabricator.wikimedia.org/T433971#12190568 (10jcrespo) This is ready for onsite ops to unrack or recycle. [09:08:18] (03PS3) 10Blake: admin_ng: pin kube-state-metrics chart to current version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320945 (https://phabricator.wikimedia.org/T427405) [09:09:48] !log bwojtowicz@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . [09:10:25] !log bwojtowicz@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . [09:17:23] (03Abandoned) 10Trueg: WIP: wdqs: pin qlever memory request == limit [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311052 (https://phabricator.wikimedia.org/T431395) (owner: 10Gmodena) [09:18:50] (03PS1) 10Marostegui: mariadb: Productionize db1266 [puppet] - 10https://gerrit.wikimedia.org/r/1321890 (https://phabricator.wikimedia.org/T407942) [09:19:10] !log marostegui@cumin1003 START - Cookbook sre.mysql.depool depool db1151: Cloning [09:19:10] !log marostegui@cumin1003 START - Cookbook sre.mysql.parsercache [09:20:09] !log marostegui@cumin1003 END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) [09:20:09] !log marostegui@cumin1003 END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Cloning [09:20:36] !log marostegui@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2253.codfw.wmnet,db1151.eqiad.wmnet with reason: cloning [09:22:48] (03CR) 10Blake: [C:03+2] admin_ng: pin kube-state-metrics chart to current version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320945 (https://phabricator.wikimedia.org/T427405) (owner: 10Blake) [09:27:54] (03PS25) 10Slyngshede: varnish: Configure text for thumbnails [puppet] - 10https://gerrit.wikimedia.org/r/1306896 (https://phabricator.wikimedia.org/T427465) [09:30:34] (03Merged) 10jenkins-bot: admin_ng: pin kube-state-metrics chart to current version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320945 (https://phabricator.wikimedia.org/T427405) (owner: 10Blake) [09:30:55] !log bounce cr3-eqsin<->cr2-eqiad bgp session to disable no-prepend command [09:30:58] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [09:31:12] !log blake@deploy1003 helmfile [codfw] START helmfile.d/admin 'apply'. [09:33:17] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/admin 'apply'. [09:34:12] (03PS1) 10Marostegui: instances.yaml: Remove db1178 from dbctl [puppet] - 10https://gerrit.wikimedia.org/r/1321895 (https://phabricator.wikimedia.org/T433471) [09:34:58] (03CR) 10Marostegui: [C:03+2] instances.yaml: Remove db1178 from dbctl [puppet] - 10https://gerrit.wikimedia.org/r/1321895 (https://phabricator.wikimedia.org/T433471) (owner: 10Marostegui) [09:36:33] !log marostegui@cumin1003 dbctl commit (dc=all): 'Remove db1178 from dbctl T433471', diff saved to https://phabricator.wikimedia.org/P95928 and previous config saved to /var/cache/conftool/dbconfig/20260806-093632-marostegui.json [09:36:37] T433471: decommission db1178.eqiad.wmnet - https://phabricator.wikimedia.org/T433471 [09:39:09] !log marostegui@cumin1003 dbctl commit (dc=all): 'Pool back ms2', diff saved to https://phabricator.wikimedia.org/P95929 and previous config saved to /var/cache/conftool/dbconfig/20260806-093908-marostegui.json [09:39:16] (03PS1) 10AikoChou: ml-services: scale revertrisk-wikidata to 1 replica in eqiad [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321896 (https://phabricator.wikimedia.org/T420883) [09:40:05] (03PS2) 10AikoChou: ml-services: scale revertrisk-wikidata to 1 replica [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321896 (https://phabricator.wikimedia.org/T420883) [09:40:42] (03PS1) 10Marostegui: db1178: Disable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1321897 (https://phabricator.wikimedia.org/T433471) [09:41:33] (03CR) 10Marostegui: [C:03+2] db1178: Disable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1321897 (https://phabricator.wikimedia.org/T433471) (owner: 10Marostegui) [09:50:41] (03PS26) 10Slyngshede: varnish: Configure text for thumbnails [puppet] - 10https://gerrit.wikimedia.org/r/1306896 (https://phabricator.wikimedia.org/T427465) [09:54:05] FIRING: KubernetesCalicoDown: ml-serve1015.eqiad.wmnet is not running calico-node Pod - https://wikitech.wikimedia.org/wiki/Calico#Operations - https://grafana.wikimedia.org/d/G8zPL7-Wz/?var-dc=eqiad%20prometheus%2Fk8s-mlserve&var-instance=ml-serve1015.eqiad.wmnet - https://alerts.wikimedia.org/?q=alertname%3DKubernetesCalicoDown [09:55:30] (03PS1) 10AikoChou: ml-services: scale revertrisk-wikidata to 3 replicas [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321899 (https://phabricator.wikimedia.org/T420883) [09:57:31] (03PS27) 10Slyngshede: varnish: Configure text for thumbnails [puppet] - 10https://gerrit.wikimedia.org/r/1306896 (https://phabricator.wikimedia.org/T427465) [10:00:04] Deploy window MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260806T1000) [10:01:42] (03CR) 10Joal: "Nits and questions" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320828 (https://phabricator.wikimedia.org/T433962) (owner: 10JavierMonton) [10:03:56] (03PS1) 10Clément Goubert: alertmanager: 24h repeat interval for TSP slack [puppet] - 10https://gerrit.wikimedia.org/r/1321900 [10:04:15] FIRING: [2x] NodeBGPSessionStatusNotEstablished: Kubernetes node ml-serve1015 has a BGP session which is not in the 'established' state. [10:07:43] !log cwilliams@cumin1003 START - Cookbook sre.mysql.depool depool db2187: Security update [10:09:08] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2187: Security update [10:09:18] (03PS28) 10Slyngshede: varnish: Configure text for thumbnails [puppet] - 10https://gerrit.wikimedia.org/r/1306896 (https://phabricator.wikimedia.org/T427465) [10:10:09] (03CR) 10JavierMonton: stream: webrequest-pageview (032 comments) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320828 (https://phabricator.wikimedia.org/T433962) (owner: 10JavierMonton) [10:12:12] (03CR) 10Joal: [C:03+1] stream: webrequest-pageview [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320828 (https://phabricator.wikimedia.org/T433962) (owner: 10JavierMonton) [10:17:46] (03PS1) 10Mvolz: Revert "citoid: update version" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321904 [10:17:54] (03CR) 10Mvolz: [C:03+2] Revert "citoid: update version" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321904 (owner: 10Mvolz) [10:18:01] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [10:20:19] (03Merged) 10jenkins-bot: Revert "citoid: update version" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321904 (owner: 10Mvolz) [10:44:00] (03CR) 10Phuedx: [C:03+1] Test Kitchen UI: Deploy v1.5.2 release to staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321824 (https://phabricator.wikimedia.org/T432629) (owner: 10Clare Ming) [10:44:14] (03CR) 10Phuedx: [C:03+1] Test Kitchen UI: Deploy v1.5.2 release to production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321827 (https://phabricator.wikimedia.org/T432629) (owner: 10Clare Ming) [10:44:27] (03CR) 10Audrey Penven: [C:03+1] wikidata-query-gui: bump to latest version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321822 (https://phabricator.wikimedia.org/T433277) (owner: 10Arthur taylor) [10:45:00] 06SRE, 10SRE-swift-storage, 10MediaWiki-extensions-Score, 06Reader Experience Team: Add cache key information to metadata json - https://phabricator.wikimedia.org/T257093#12190953 (10MatthewVernon) OK, right, yes, I think that is likely worth doing. Looking at the container now: ` root@ms-fe2009:/home/mver... [10:48:37] (03CR) 10Arthur taylor: [C:03+2] wikidata-query-gui: bump to latest version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321822 (https://phabricator.wikimedia.org/T433277) (owner: 10Arthur taylor) [10:49:17] jouncebot: nowandnext [10:49:17] For the next 0 hour(s) and 10 minute(s): MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260806T1000) [10:49:17] In 1 hour(s) and 10 minute(s): Mobileapps/RESTBase/Wikifeeds (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260806T1200) [10:50:47] 06SRE, 10SRE-swift-storage, 10MediaWiki-extensions-Score, 06Reader Experience Team: Add cache key information to metadata json - https://phabricator.wikimedia.org/T257093#12190977 (10MatthewVernon) My uninformed view is that "we want to keep them except when we upgrade lilypond and then want to regenerate... [10:50:54] PROBLEM - Docker registry HTTPS interface on registry1004 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/Docker [10:51:07] (03Merged) 10jenkins-bot: wikidata-query-gui: bump to latest version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321822 (https://phabricator.wikimedia.org/T433277) (owner: 10Arthur taylor) [10:51:44] RECOVERY - Docker registry HTTPS interface on registry1004 is OK: HTTP OK: HTTP/1.1 200 OK - 3745 bytes in 0.187 second response time https://wikitech.wikimedia.org/wiki/Docker [10:53:37] !log arthurtaylor@deploy1003 helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply [10:54:02] !log arthurtaylor@deploy1003 helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply [10:56:08] !log arthurtaylor@deploy1003 helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply [10:56:28] !log arthurtaylor@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply [10:56:35] !log arthurtaylor@deploy1003 helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply [10:56:52] !log arthurtaylor@deploy1003 helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply [10:56:57] (03PS1) 10Ayounsi: WIP: Add TransportLinksInUsageNoRedundancy alert [alerts] - 10https://gerrit.wikimedia.org/r/1321906 [10:57:53] !log cwilliams@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2187.codfw.wmnet with reason: Maintenance [10:58:53] (03PS1) 10Marostegui: mariadb: Add test-s4 [puppet] - 10https://gerrit.wikimedia.org/r/1321907 (https://phabricator.wikimedia.org/T427059) [10:59:28] (03CR) 10CI reject: [V:04-1] mariadb: Add test-s4 [puppet] - 10https://gerrit.wikimedia.org/r/1321907 (https://phabricator.wikimedia.org/T427059) (owner: 10Marostegui) [11:00:10] (03CR) 10CI reject: [V:04-1] WIP: Add TransportLinksInUsageNoRedundancy alert [alerts] - 10https://gerrit.wikimedia.org/r/1321906 (owner: 10Ayounsi) [11:03:49] (03PS2) 10Marostegui: mariadb: Add test-s4 [puppet] - 10https://gerrit.wikimedia.org/r/1321907 (https://phabricator.wikimedia.org/T427059) [11:07:52] (03CR) 10JavierMonton: [C:03+2] stream: webrequest-pageview [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320828 (https://phabricator.wikimedia.org/T433962) (owner: 10JavierMonton) [11:10:25] !log cwilliams@cumin1003 START - Cookbook sre.mysql.pool pool db2187: Security update [11:14:37] (03PS29) 10Slyngshede: varnish: Configure text for thumbnails [puppet] - 10https://gerrit.wikimedia.org/r/1306896 (https://phabricator.wikimedia.org/T427465) [11:15:07] !log brouberol@deploy1003 helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. [11:16:03] !log brouberol@deploy1003 helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. [11:17:20] (03Merged) 10jenkins-bot: stream: webrequest-pageview [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320828 (https://phabricator.wikimedia.org/T433962) (owner: 10JavierMonton) [11:18:03] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Grant access to analytics-privatedata-users for EAlbizzati-WMF - https://phabricator.wikimedia.org/T434034#12191087 (10Seddon) Approved. [11:21:00] (03PS30) 10Slyngshede: varnish: Configure text for thumbnails [puppet] - 10https://gerrit.wikimedia.org/r/1306896 (https://phabricator.wikimedia.org/T427465) [11:21:03] (03PS2) 10Blake: httpbb: Release v0.0.6 [software/httpbb] - 10https://gerrit.wikimedia.org/r/1321516 (https://phabricator.wikimedia.org/T434052) [11:21:51] (03CR) 10Blake: httpbb: Release v0.0.6 (031 comment) [software/httpbb] - 10https://gerrit.wikimedia.org/r/1321516 (https://phabricator.wikimedia.org/T434052) (owner: 10Blake) [11:23:39] (03PS2) 10Ayounsi: WIP: Add TransportLinksInUsageNoRedundancy alert [alerts] - 10https://gerrit.wikimedia.org/r/1321906 [11:24:18] !log javiermonton@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply [11:26:50] !log javiermonton@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply [11:26:54] (03CR) 10CI reject: [V:04-1] WIP: Add TransportLinksInUsageNoRedundancy alert [alerts] - 10https://gerrit.wikimedia.org/r/1321906 (owner: 10Ayounsi) [11:28:01] FIRING: [2x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-1/0/2 (Transport: cr1-eqiad:et-1/1/2 (Arelion, IC-374549)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [11:31:31] !log javiermonton@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply [11:31:41] !log javiermonton@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply [11:36:59] (03CR) 10Gkyziridis: [C:03+1] "Thnx for working on this!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321896 (https://phabricator.wikimedia.org/T420883) (owner: 10AikoChou) [11:37:33] (03CR) 10Gkyziridis: [C:03+1] "LGTM!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321899 (https://phabricator.wikimedia.org/T420883) (owner: 10AikoChou) [11:40:24] James_F: I +2'ed the patch but if you want me to wait, I'm okay with it [11:40:25] https://gerrit.wikimedia.org/r/c/mediawiki/tools/codesniffer/+/1321910 [11:47:58] (03PS1) 10Bartosz Wójtowicz: ml-services: Declare ephemeral-storage on prod services. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321915 (https://phabricator.wikimedia.org/T431089) [11:53:01] FIRING: [2x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-1/0/2 (Transport: cr1-eqiad:et-1/1/2 (Arelion, IC-374549)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [11:53:57] RESOLVED: [2x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-1/0/2 (Transport: cr1-eqiad:et-1/1/2 (Arelion, IC-374549)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [11:55:57] 06SRE, 06Infrastructure-Foundations, 10vm-requests: eqiad/codfw: 3 trixie VMs for rdb-lock - https://phabricator.wikimedia.org/T434188 (10Clement_Goubert) 03NEW [11:56:30] 06SRE, 06Infrastructure-Foundations, 06ServiceOps, 10ServiceOps-Datastores, and 2 others: eqiad/codfw: 3 trixie VMs for rdb-lock - https://phabricator.wikimedia.org/T434188#12191178 (10Clement_Goubert) p:05Triage→03Medium [11:58:29] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2187: Security update [11:59:02] (03PS1) 10PipelineBot: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321917 [12:00:05] Deploy window Mobileapps/RESTBase/Wikifeeds (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260806T1200) [12:08:09] (03CR) 10Kevin Bazira: [C:03+1] ml-services: Declare ephemeral-storage on prod services. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321915 (https://phabricator.wikimedia.org/T431089) (owner: 10Bartosz Wójtowicz) [12:08:53] !log brouberol@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. [12:09:58] !log brouberol@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. [12:10:49] (03PS1) 10Sadiya.mohammed13: Use QLever example queries in query-next [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321920 (https://phabricator.wikimedia.org/T432638) [12:12:20] !log javiermonton@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply [12:12:25] !log javiermonton@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply [12:21:36] (03CR) 10CI reject: [V:04-1] Localisation updates from https://translatewiki.net. [phabricator/translations] (wmf/stable) - 10https://gerrit.wikimedia.org/r/1321925 (owner: 10L10n-bot) [12:25:02] (03CR) 10Audrey Penven: "looks good, but we should also add this to `helmfile.d / services / wikidata-query-gui / values-wikidata-query-scholarly-next-gui.yaml`" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321920 (https://phabricator.wikimedia.org/T432638) (owner: 10Sadiya.mohammed13) [12:25:53] (03PS1) 10Brouberol: an-test-client1002: stop trying to install airflow [puppet] - 10https://gerrit.wikimedia.org/r/1321932 (https://phabricator.wikimedia.org/T406746) [12:26:49] (03CR) 10Audrey Penven: "(accidentally added as "resolved" comment)" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321920 (https://phabricator.wikimedia.org/T432638) (owner: 10Sadiya.mohammed13) [12:28:53] (03CR) 10CDanis: [C:03+1] "thx!" [puppet] - 10https://gerrit.wikimedia.org/r/1319483 (https://phabricator.wikimedia.org/T431683) (owner: 10Ayounsi) [12:32:30] (03CR) 10Jcrespo: [C:03+1] "This is now ready, do you want me to merge it?" [puppet] - 10https://gerrit.wikimedia.org/r/1321651 (https://phabricator.wikimedia.org/T434034) (owner: 10Dzahn) [12:32:31] (03CR) 10Joal: [C:03+1] "LGTM!" [puppet] - 10https://gerrit.wikimedia.org/r/1321932 (https://phabricator.wikimedia.org/T406746) (owner: 10Brouberol) [12:32:36] (03CR) 10Brouberol: [C:03+2] an-test-client1002: stop trying to install airflow [puppet] - 10https://gerrit.wikimedia.org/r/1321932 (https://phabricator.wikimedia.org/T406746) (owner: 10Brouberol) [12:32:45] (03CR) 10Bartosz Wójtowicz: [C:03+2] ml-services: Declare ephemeral-storage on prod services. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321915 (https://phabricator.wikimedia.org/T431089) (owner: 10Bartosz Wójtowicz) [12:33:31] (03CR) 10Ayounsi: [C:03+2] Fastnetmon bump threshold_pps to 2Mpps and add threshold_udp/icmp [puppet] - 10https://gerrit.wikimedia.org/r/1319483 (https://phabricator.wikimedia.org/T431683) (owner: 10Ayounsi) [12:36:44] (03PS2) 10Sadiya.mohammed13: Use QLever example queries in query-next [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321920 (https://phabricator.wikimedia.org/T432638) [12:38:03] (03CR) 10Audrey Penven: "looks good to me!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321920 (https://phabricator.wikimedia.org/T432638) (owner: 10Sadiya.mohammed13) [12:38:03] (03Merged) 10jenkins-bot: ml-services: Declare ephemeral-storage on prod services. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321915 (https://phabricator.wikimedia.org/T431089) (owner: 10Bartosz Wójtowicz) [12:38:04] !log brouberol@cumin1003 START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm [12:38:23] (03CR) 10Audrey Penven: [C:03+1] Use QLever example queries in query-next [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321920 (https://phabricator.wikimedia.org/T432638) (owner: 10Sadiya.mohammed13) [12:39:35] !log bwojtowicz@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . [12:41:04] !log bwojtowicz@deploy1003 helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . [12:42:00] (03CR) 10AikoChou: [C:03+2] ml-services: scale revertrisk-wikidata to 1 replica [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321896 (https://phabricator.wikimedia.org/T420883) (owner: 10AikoChou) [12:43:55] (03CR) 10Cathal Mooney: [C:03+2] Set standard metric for new easms and drmrs transports to codfw [homer/public] - 10https://gerrit.wikimedia.org/r/1321572 (https://phabricator.wikimedia.org/T424839) (owner: 10Cathal Mooney) [12:44:32] (03Merged) 10jenkins-bot: ml-services: scale revertrisk-wikidata to 1 replica [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321896 (https://phabricator.wikimedia.org/T420883) (owner: 10AikoChou) [12:45:07] !log bwojtowicz@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-models' for release 'main' . [12:45:37] (03Merged) 10jenkins-bot: Set standard metric for new easms and drmrs transports to codfw [homer/public] - 10https://gerrit.wikimedia.org/r/1321572 (https://phabricator.wikimedia.org/T424839) (owner: 10Cathal Mooney) [12:46:12] !log bwojtowicz@deploy1003 helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . [12:48:21] !log bwojtowicz@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . [12:50:27] !log bwojtowicz@deploy1003 helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . [12:51:50] !log bwojtowicz@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' . [12:53:19] !log bwojtowicz@deploy1003 helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' . [12:53:24] !log brouberol@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage [12:54:46] !log bwojtowicz@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . [12:55:29] !log bwojtowicz@deploy1003 helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . [12:57:02] !log aikochou@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . [12:57:03] !log bwojtowicz@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'logo-detection' for release 'main' . [12:57:51] !log bwojtowicz@deploy1003 helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . [12:58:22] !log brouberol@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage [12:59:01] !log bwojtowicz@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'readability' for release 'main' . [12:59:08] (03CR) 10AikoChou: [C:03+2] ml-services: scale revertrisk-wikidata to 3 replicas [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321899 (https://phabricator.wikimedia.org/T420883) (owner: 10AikoChou) [12:59:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:00:05] Lucas_WMDE, urbanecm, and TheresNoTime: OwO what's this, a deployment window?? UTC afternoon backport window. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260806T1300). nyaa~ [13:00:05] No Gerrit patches in the queue for this window AFAICS. [13:00:37] !log bwojtowicz@deploy1003 helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . [13:00:51] (03Merged) 10jenkins-bot: ml-services: scale revertrisk-wikidata to 3 replicas [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321899 (https://phabricator.wikimedia.org/T420883) (owner: 10AikoChou) [13:02:00] !log aikochou@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' . [13:03:22] !log bwojtowicz@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . [13:03:59] !log bwojtowicz@deploy1003 helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . [13:04:25] RESOLVED: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:04:32] !log brouberol@cumin1003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-client1002.eqiad.wmnet with OS bookworm [13:05:01] !log bwojtowicz@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revision-models' for release 'main' . [13:05:42] jouncebot: refresh [13:05:43] I refreshed my knowledge about deployments. [13:05:46] !log aikochou@deploy1003 helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . [13:05:48] jouncebot: nowandnext [13:05:48] For the next 0 hour(s) and 54 minute(s): UTC afternoon backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260806T1300) [13:05:48] In 1 hour(s) and 24 minute(s): Test Kitchen Experiment Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260806T1430) [13:05:50] !log bwojtowicz@deploy1003 helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . [13:06:35] !log brouberol@cumin1003 START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm [13:06:51] !log bwojtowicz@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . [13:08:15] (03CR) 10TrainBranchBot: [C:03+2] "Approved by hashar@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1314816 (https://phabricator.wikimedia.org/T305571) (owner: 10Diegodlh) [13:08:32] that change only affects beta cluster (it only touches `-labs.php` files) [13:09:19] !log bwojtowicz@deploy1003 helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . [13:09:32] (03Merged) 10jenkins-bot: wmf-config: Set JSON content model for Web2Cit configuration pages [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1314816 (https://phabricator.wikimedia.org/T305571) (owner: 10Diegodlh) [13:11:44] > 13:10:12 Skipping sync since all commits were beta/labs-only changes. Operation completed. [13:11:45] all good [13:11:55] !log bwojtowicz@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . [13:13:58] !log bwojtowicz@deploy1003 helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' . [13:14:09] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Thursday, August 06 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-i" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1315128 (https://phabricator.wikimedia.org/T432412) (owner: 10Tsevener) [13:16:16] !log bwojtowicz@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . [13:16:46] (03PS1) 10Trueg: WDQS: Enable prometheus metrics on Qlever [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321948 (https://phabricator.wikimedia.org/T433040) [13:17:16] !log bwojtowicz@deploy1003 helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' . [13:18:33] !log bwojtowicz@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . [13:18:33] !log brouberol@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage [13:18:55] (03PS1) 10JMeybohm: Update to 1.29.4 [debs/istio] - 10https://gerrit.wikimedia.org/r/1321949 (https://phabricator.wikimedia.org/T427401) [13:19:48] !log bwojtowicz@deploy1003 helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' . [13:20:12] (03PS31) 10Slyngshede: varnish: Configure text for thumbnails [puppet] - 10https://gerrit.wikimedia.org/r/1306896 (https://phabricator.wikimedia.org/T427465) [13:22:22] !log bwojtowicz@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . [13:23:19] !log brouberol@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage [13:24:19] (03PS2) 10JMeybohm: Update to 1.29.4 [debs/istio] - 10https://gerrit.wikimedia.org/r/1321949 (https://phabricator.wikimedia.org/T434198) [13:26:08] (03PS3) 10Atsuko: airflow3 support [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320131 (https://phabricator.wikimedia.org/T433388) [13:26:10] !log bwojtowicz@deploy1003 helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . [13:28:07] (03PS4) 10Atsuko: airflow3 support [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320131 (https://phabricator.wikimedia.org/T433388) [13:29:11] !log bwojtowicz@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . [13:31:25] (03CR) 10Brouberol: [C:03+1] airflow3 support [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320131 (https://phabricator.wikimedia.org/T433388) (owner: 10Atsuko) [13:31:47] !log bwojtowicz@deploy1003 helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . [13:32:30] FIRING: Traffic bill over quota: Alert for device cr2-eqsin.wikimedia.org - Traffic bill over quota - https://alerts.wikimedia.org/?q=alertname%3DTraffic+bill+over+quota [13:41:40] (03PS5) 10Atsuko: airflow3 support [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320131 (https://phabricator.wikimedia.org/T433388) [13:42:32] (03CR) 10Scott French: [C:03+1] "Thanks, Blake!" [software/httpbb] - 10https://gerrit.wikimedia.org/r/1321516 (https://phabricator.wikimedia.org/T434052) (owner: 10Blake) [13:43:17] (03CR) 10Bking: [C:03+1] opensearch-semantic-search: bump to opensearch 3.7.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319845 (https://phabricator.wikimedia.org/T433697) (owner: 10DCausse) [13:44:01] !log aikochou@deploy1003 helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . [13:45:27] (03CR) 10Atsuko: [C:03+2] airflow3 support [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320131 (https://phabricator.wikimedia.org/T433388) (owner: 10Atsuko) [13:48:57] (03Merged) 10jenkins-bot: airflow3 support [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320131 (https://phabricator.wikimedia.org/T433388) (owner: 10Atsuko) [13:49:13] !log aikochou@deploy1003 helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . [13:51:18] (03PS2) 10Ssingh: wmf-config/ProductionServices: set URL for urldownloader to discovery record [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1313985 (https://phabricator.wikimedia.org/T429175) [13:52:11] (03PS1) 10Bartosz Wójtowicz: ml-services: Lower ephemeral-storage requests for revscoring services. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321967 (https://phabricator.wikimedia.org/T431089) [13:52:30] RESOLVED: Traffic bill over quota: Alert for device cr2-eqsin.wikimedia.org - Traffic bill over quota - https://alerts.wikimedia.org/?q=alertname%3DTraffic+bill+over+quota [13:54:04] (03CR) 10CDobbins: [C:03+2] Revert "hieradata: Temporarily point eqiad PyBals at codfw etcd" [puppet] - 10https://gerrit.wikimedia.org/r/1321668 (https://phabricator.wikimedia.org/T428495) (owner: 10Scott French) [13:54:05] FIRING: KubernetesCalicoDown: ml-serve1015.eqiad.wmnet is not running calico-node Pod - https://wikitech.wikimedia.org/wiki/Calico#Operations - https://grafana.wikimedia.org/d/G8zPL7-Wz/?var-dc=eqiad%20prometheus%2Fk8s-mlserve&var-instance=ml-serve1015.eqiad.wmnet - https://alerts.wikimedia.org/?q=alertname%3DKubernetesCalicoDown [13:55:52] (03CR) 10Scott French: [C:03+2] Revert "Temporarily point all etcd client SRV records to codfw" [dns] - 10https://gerrit.wikimedia.org/r/1321667 (https://phabricator.wikimedia.org/T428495) (owner: 10Scott French) [13:56:08] !log swfrench@dns1004 START - running authdns-update [13:58:07] !log swfrench@dns1004 END - running authdns-update [13:58:13] !log authdns update to direct eqiad-associated etcd clients back to eqiad - T428495 [13:58:17] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [13:58:17] T428495: Migrate conf* hosts away from bullseye - https://phabricator.wikimedia.org/T428495 [13:59:19] (03PS3) 10Ssingh: wmf-config/ProductionServices: set URL for urldownloader to service record [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1313985 (https://phabricator.wikimedia.org/T429175) [14:02:22] !log brouberol@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-client1002.eqiad.wmnet with OS bookworm [14:03:43] (03CR) 10CDanis: [C:03+1] wmf-config/ProductionServices: set URL for urldownloader to service record [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1313985 (https://phabricator.wikimedia.org/T429175) (owner: 10Ssingh) [14:04:15] FIRING: [2x] NodeBGPSessionStatusNotEstablished: Kubernetes node ml-serve1015 has a BGP session which is not in the 'established' state. [14:04:22] !log brouberol@cumin1003 START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm [14:04:30] !log brouberol@cumin1003 START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm [14:04:43] !log begin rolling restart of confd in drmrs, eqiad, esams, magru - T428495 [14:04:46] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:04:47] T428495: Migrate conf* hosts away from bullseye - https://phabricator.wikimedia.org/T428495 [14:09:17] (03CR) 10Ssingh: [C:03+1] "Thanks for the explanation in the commit message." [puppet] - 10https://gerrit.wikimedia.org/r/1290731 (https://phabricator.wikimedia.org/T425441) (owner: 10Arnaudb) [14:09:25] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [14:11:48] !log restarted navtiming on webperf1003 - T428495 [14:11:51] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:11:52] T428495: Migrate conf* hosts away from bullseye - https://phabricator.wikimedia.org/T428495 [14:12:29] !log sudo cumin 'A:cp-text' "disable-puppet 'merging CR 1290731'": T425441 [14:12:33] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:12:33] T425441: gitlab behind CDN - https://phabricator.wikimedia.org/T425441 [14:12:40] PROBLEM - PyBal connections to etcd on lvs1017 is CRITICAL: CRITICAL: 0 connections established with conf1007.eqiad.wmnet:4001 (min=12) https://wikitech.wikimedia.org/wiki/PyBal [14:14:55] !log brouberol@cumin1003 START - Cookbook sre.hosts.reimage for host an-test-ui1001.eqiad.wmnet with OS bookworm [14:16:06] (03CR) 10Ssingh: [C:03+2] trafficserver: add a map for gitlab instances as a backend [puppet] - 10https://gerrit.wikimedia.org/r/1290731 (https://phabricator.wikimedia.org/T425441) (owner: 10Arnaudb) [14:16:13] !log brouberol@cumin1003 START - Cookbook sre.hosts.reimage for host an-test-presto1001.eqiad.wmnet with OS bookworm [14:18:01] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [14:18:41] (03PS1) 10JavierMonton: namespaces: webrequest-page-view [puppet] - 10https://gerrit.wikimedia.org/r/1321974 (https://phabricator.wikimedia.org/T433962) [14:18:55] !log begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - T428495 [14:18:58] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:18:59] T428495: Migrate conf* hosts away from bullseye - https://phabricator.wikimedia.org/T428495 [14:21:05] (03PS1) 10JavierMonton: topic: webrequest-page-view [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321976 [14:21:56] (03PS2) 10JavierMonton: topic: webrequest-page-view [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321976 (https://phabricator.wikimedia.org/T433962) [14:22:41] RECOVERY - PyBal connections to etcd on lvs1017 is OK: OK: 12 connections established with conf1007.eqiad.wmnet:4001 (min=12) https://wikitech.wikimedia.org/wiki/PyBal [14:23:01] !log sudo cumin -b2 'A:cp-text' "run-puppet-agent --enable 'merging CR 1290731'": T425441 [14:23:05] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:23:06] T425441: gitlab behind CDN - https://phabricator.wikimedia.org/T425441 [14:27:44] !log brouberol@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage [14:28:03] !log brouberol@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage [14:30:05] Deploy window Test Kitchen Experiment Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260806T1430) [14:34:48] !log brouberol@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage [14:38:25] !log brouberol@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage [14:40:30] !log brouberol@cumin1003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm [14:40:57] !log cdobbins@cumin1003 START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo (T428495) [14:41:01] T428495: Migrate conf* hosts away from bullseye - https://phabricator.wikimedia.org/T428495 [14:42:36] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo (T428495) [14:42:42] !log brouberol@cumin1003 START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm [14:42:46] !log brouberol@cumin1003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-worker1002.eqiad.wmnet with OS bookworm [14:43:13] !log brouberol@cumin1003 START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm [14:44:20] !log cdobbins@cumin1003 START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams (T428495) [14:46:14] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams (T428495) [14:46:18] T428495: Migrate conf* hosts away from bullseye - https://phabricator.wikimedia.org/T428495 [14:47:38] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for iberker - https://phabricator.wikimedia.org/T433258#12191908 (10SCherukuwada) Manager approves. [14:47:55] (03PS1) 10Ssingh: P:kubernetes::deployment_server::global_config: update IPs for urldownloader [puppet] - 10https://gerrit.wikimedia.org/r/1321987 (https://phabricator.wikimedia.org/T429175) [14:48:42] !log brouberol@cumin1003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm [14:49:28] !log cdobbins@cumin1003 START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru (T428495) [14:51:51] (03CR) 10Ahmon Dancy: "wmf-beta-update-all has been failing since this merged:" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1314816 (https://phabricator.wikimedia.org/T305571) (owner: 10Diegodlh) [14:51:54] (03CR) 10Ssingh: "I suspect the PCC failure is unrelated" [puppet] - 10https://gerrit.wikimedia.org/r/1321987 (https://phabricator.wikimedia.org/T429175) (owner: 10Ssingh) [14:52:20] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru (T428495) [14:52:25] T428495: Migrate conf* hosts away from bullseye - https://phabricator.wikimedia.org/T428495 [14:53:32] (03PS2) 10Ssingh: P:kubernetes::deployment_server::global_config: update IPs for urldownloader [puppet] - 10https://gerrit.wikimedia.org/r/1321987 (https://phabricator.wikimedia.org/T429175) [14:54:41] !log brouberol@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-presto1001.eqiad.wmnet with OS bookworm [14:55:15] !log brouberol@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-ui1001.eqiad.wmnet with OS bookworm [14:55:45] !log cdobbins@cumin1003 START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs (T428495) [14:57:42] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs (T428495) [14:57:47] T428495: Migrate conf* hosts away from bullseye - https://phabricator.wikimedia.org/T428495 [14:58:33] !log brouberol@cumin1003 START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm [14:59:56] (03CR) 10Klausman: [C:03+2] namespaces: webrequest-page-view [puppet] - 10https://gerrit.wikimedia.org/r/1321974 (https://phabricator.wikimedia.org/T433962) (owner: 10JavierMonton) [15:00:05] jnuche and jeena: Your horoscope predicts another Train log triage deploy. May Zuul be (nice) with you. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260806T1500). [15:00:06] (03CR) 10Klausman: [C:03+1] topic: webrequest-page-view [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321976 (https://phabricator.wikimedia.org/T433962) (owner: 10JavierMonton) [15:02:59] (03CR) 10Brennen Bearnes: [C:03+1] scap: Add optional logstash credentials file [puppet] - 10https://gerrit.wikimedia.org/r/1321653 (https://phabricator.wikimedia.org/T434114) (owner: 10Ahmon Dancy) [15:07:50] (03PS3) 10Ssingh: P:kubernetes::deployment_server::global_config: update IPs for urldownloader [puppet] - 10https://gerrit.wikimedia.org/r/1321987 (https://phabricator.wikimedia.org/T429175) [15:10:57] (03PS1) 10Kamila Součková: deployment_server: fix mariadb_master_ips handling of undef [puppet] - 10https://gerrit.wikimedia.org/r/1321999 [15:11:02] (03CR) 10Dzahn: [C:03+1] "thanks. sure, feel free to join it. or I can in a little while." [puppet] - 10https://gerrit.wikimedia.org/r/1321651 (https://phabricator.wikimedia.org/T434034) (owner: 10Dzahn) [15:13:54] (03CR) 10Ssingh: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1321999 (owner: 10Kamila Součková) [15:16:46] (03CR) 10Ssingh: [C:03+1] "Looks good to me and in PCC." [puppet] - 10https://gerrit.wikimedia.org/r/1321999 (owner: 10Kamila Součková) [15:18:11] (03CR) 10Scott French: [C:03+1] "Thanks, Raine!" [puppet] - 10https://gerrit.wikimedia.org/r/1321999 (owner: 10Kamila Součková) [15:21:50] (03CR) 10Kamila Součková: [C:03+2] deployment_server: fix mariadb_master_ips handling of undef [puppet] - 10https://gerrit.wikimedia.org/r/1321999 (owner: 10Kamila Součková) [15:22:52] (03CR) 10Ssingh: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1321987 (https://phabricator.wikimedia.org/T429175) (owner: 10Ssingh) [15:25:55] (03PS1) 10Brouberol: preseed: add grub-installer stanzas to an-test-worker hosts [puppet] - 10https://gerrit.wikimedia.org/r/1322002 (https://phabricator.wikimedia.org/T406746) [15:29:02] !log brouberol@cumin1003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm [15:29:09] !log brouberol@cumin1003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm [15:38:37] (03CR) 10CDanis: [C:03+1] preseed: add grub-installer stanzas to an-test-worker hosts [puppet] - 10https://gerrit.wikimedia.org/r/1322002 (https://phabricator.wikimedia.org/T406746) (owner: 10Brouberol) [15:38:54] (03CR) 10Brouberol: [C:03+2] preseed: add grub-installer stanzas to an-test-worker hosts [puppet] - 10https://gerrit.wikimedia.org/r/1322002 (https://phabricator.wikimedia.org/T406746) (owner: 10Brouberol) [15:43:20] !log brouberol@cumin1003 START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm [15:45:51] Hi! I had a change to CommonSettings-labs.php merged recently with the help of hashar and matmarex. But as dancy commented on the patch, the change is causing wmf-beta-update-all to fail. I commented about this in https://phabricator.wikimedia.org/T305571#12191466. Shall I submit a separate patch fixing the problem and reference the old change id [15:45:51] in the commit message? [15:46:09] Sorry, changeset is https://gerrit.wikimedia.org/r/c/operations/mediawiki-config/+/1314816 [15:46:11] sure! that helps reviewers I guess [15:46:33] I am happy to merge it [15:46:47] dancy: sorry I did not check whether the beta update job actually worked :\ [15:47:31] I will submit another patch right away [15:47:40] Thanks and sorry for the inconveniences [15:50:51] (03CR) 10Ssingh: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1321987 (https://phabricator.wikimedia.org/T429175) (owner: 10Ssingh) [15:52:43] (03PS1) 10CDanis: Blocklist SCTP [puppet] - 10https://gerrit.wikimedia.org/r/1322009 [15:53:20] (03CR) 10CI reject: [V:04-1] Blocklist SCTP [puppet] - 10https://gerrit.wikimedia.org/r/1322009 (owner: 10CDanis) [15:53:47] (03PS1) 10JHathaway: base::kernel: blacklist sctp module [puppet] - 10https://gerrit.wikimedia.org/r/1322010 [15:54:00] (03CR) 10CDanis: [C:03+1] base::kernel: blacklist sctp module [puppet] - 10https://gerrit.wikimedia.org/r/1322010 (owner: 10JHathaway) [15:54:08] (03CR) 10Kevin Bazira: [C:03+1] ml-services: Lower ephemeral-storage requests for revscoring services. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321967 (https://phabricator.wikimedia.org/T431089) (owner: 10Bartosz Wójtowicz) [15:54:18] (03Abandoned) 10CDanis: Blocklist SCTP [puppet] - 10https://gerrit.wikimedia.org/r/1322009 (owner: 10CDanis) [15:56:42] (03CR) 10JHathaway: [C:03+2] base::kernel: blacklist sctp module [puppet] - 10https://gerrit.wikimedia.org/r/1322010 (owner: 10JHathaway) [15:56:54] (03PS4) 10Ssingh: P:kubernetes::deployment_server::global_config: update IPs for urldownloader [puppet] - 10https://gerrit.wikimedia.org/r/1321987 (https://phabricator.wikimedia.org/T429175) [15:57:53] !log brouberol@cumin1003 START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm [15:59:53] (03CR) 10Ssingh: [V:03+1] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9185/co" [puppet] - 10https://gerrit.wikimedia.org/r/1321987 (https://phabricator.wikimedia.org/T429175) (owner: 10Ssingh) [16:00:01] !log brouberol@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage [16:00:05] jhathaway and rzl: #bothumor I � Unicode. All rise for Puppet request window deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260806T1600). [16:00:05] No Gerrit patches in the queue for this window AFAICS. [16:00:14] (03PS1) 10Diegodlh: wmf-config: Import Title class in CommonSettings-labs.php [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1322012 (https://phabricator.wikimedia.org/T305571) [16:05:44] !log brouberol@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage [16:07:36] (03PS2) 10Krinkle: Remove Mathoid SVG option from wgMathValidModes [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311062 (https://phabricator.wikimedia.org/T271001) [16:08:01] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:16:02] !log brouberol@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage [16:18:01] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:19:49] !log brouberol@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage [16:23:03] (03CR) 10CDobbins: "36-set-cookie-logging-false-positives.vtc passes when I run it against cp2047" [puppet] - 10https://gerrit.wikimedia.org/r/1306896 (https://phabricator.wikimedia.org/T427465) (owner: 10Slyngshede) [16:24:29] (03CR) 10Ssingh: "Confirming the same, all tests pass on cp1100 (cache_text)" [puppet] - 10https://gerrit.wikimedia.org/r/1306896 (https://phabricator.wikimedia.org/T427465) (owner: 10Slyngshede) [16:35:27] !log brouberol@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1001.eqiad.wmnet with OS bookworm [16:37:01] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Thursday, August 06 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-i" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319939 (https://phabricator.wikimedia.org/T430409) (owner: 10Chlod Alejandro) [16:49:40] 10ops-eqiad, 06SRE, 10bacula, 06Data-Persistence, and 3 others: decommission backup1003.eqiad.wmnet - https://phabricator.wikimedia.org/T433970#12192316 (10VRiley-WMF) a:03VRiley-WMF [16:50:09] !log brouberol@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1002.eqiad.wmnet with OS bookworm [16:51:56] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Thursday, August 06 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-i" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320230 (https://phabricator.wikimedia.org/T204089) (owner: 10Ebernhardson) [17:00:04] bd808: Your horoscope predicts another Cloud Services/Technical Documentation weekly deploy (Toolhub, Developer portal, Striker) deploy. May Zuul be (nice) with you. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260806T1700). [17:00:05] Deploy window MediaWiki infrastructure (UTC late) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260806T1700) [17:07:08] (03CR) 10Ahmon Dancy: [C:03+2] wmf-config: Import Title class in CommonSettings-labs.php [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1322012 (https://phabricator.wikimedia.org/T305571) (owner: 10Diegodlh) [17:07:12] (03CR) 10Ahmon Dancy: [C:03+2] "Thank you!" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1322012 (https://phabricator.wikimedia.org/T305571) (owner: 10Diegodlh) [17:08:10] (03CR) 10Dzahn: [C:03+2] admin: add ealbizzati to analytics-privatedata-users, level 1 access [puppet] - 10https://gerrit.wikimedia.org/r/1321651 (https://phabricator.wikimedia.org/T434034) (owner: 10Dzahn) [17:08:19] (03Merged) 10jenkins-bot: wmf-config: Import Title class in CommonSettings-labs.php [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1322012 (https://phabricator.wikimedia.org/T305571) (owner: 10Diegodlh) [17:10:31] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Grant access to analytics-privatedata-users for EAlbizzati-WMF - https://phabricator.wikimedia.org/T434034#12192412 (10Dzahn) 05Open→03Resolved a:03Dzahn thanks all! done! @EAlbizzati-WMF you are in the group now. It can take a maximum of about 3... [17:13:25] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for iberker - https://phabricator.wikimedia.org/T433258#12192433 (10Dzahn) >>! In T433258#12178332, @IBerker-WMF wrote: > I signed the L3 document, but I don't really need production server access, I need analytics-privatedata-users... [17:13:42] 10ops-eqiad, 06SRE, 10bacula, 06Data-Persistence, and 3 others: decommission backup1003.eqiad.wmnet - https://phabricator.wikimedia.org/T433970#12192436 (10VRiley-WMF) [17:13:53] PROBLEM - SSH on urldownloader2005 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/SSH/monitoring [17:13:56] 10ops-eqiad, 06SRE, 10bacula, 06Data-Persistence, and 3 others: decommission backup1003.eqiad.wmnet - https://phabricator.wikimedia.org/T433970#12192437 (10VRiley-WMF) 05Open→03Resolved This is completed. [17:14:42] FIRING: [2x] ProbeDown: Service urldownloader2005:8080 has failed probes (http_url_downloader_wikimedia_org_ip4) - https://wikitech.wikimedia.org/wiki/Url-downloader - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [17:15:55] 10SRE-SLO, 06Abstract Wikipedia team (27Q1 (Jul–Sep)), 07OKR-Work: new SLI (1 of 2): server-side metrics on Abstract Wikipedia preview - https://phabricator.wikimedia.org/T434231 (10ecarg) 03NEW [17:17:21] 10SRE-SLO, 06Abstract Wikipedia team (27Q1 (Jul–Sep)), 07OKR-Work: new SLI (1 of 2): server-side metrics on Abstract Wikipedia preview - https://phabricator.wikimedia.org/T434231#12192482 (10ecarg) p:05Triage→03High [17:19:35] I won't be using my deploy window this week. The potential changes to deploy for developer-portal are 4 Luxembourgish translation units which seems not enough to ship at the moment. [17:19:48] 10SRE-SLO, 06Abstract Wikipedia team (27Q1 (Jul–Sep)), 07OKR-Work: new SLI (1 of 2): server-side metrics on Abstract Wikipedia preview - https://phabricator.wikimedia.org/T434231#12192512 (10ecarg) [17:20:04] 10SRE-SLO, 06Abstract Wikipedia team (27Q1 (Jul–Sep)), 07OKR-Work: new SLI (1 of 2): server-side metrics on Abstract Wikipedia preview - https://phabricator.wikimedia.org/T434231#12192515 (10ecarg) [17:20:54] 10SRE-SLO, 06Abstract Wikipedia team (27Q1 (Jul–Sep)), 07OKR-Work: new SLI (1 of 2): server-side metrics on Abstract Wikipedia preview - https://phabricator.wikimedia.org/T434231#12192521 (10ecarg) Welcoming feedback on any of the above (objective %, SLI definition), thank you in advance! [17:21:50] 06SRE: Evaluation Virtualization management platforms for a private VPS cluster - https://phabricator.wikimedia.org/T87251#12192523 (10Dzahn) for the record because I was asked. The title of this doc is `Virtualization Evaluation - Private VPS Service Goal`. [17:34:24] (03CR) 10Dzahn: [C:03+1] sre.gitlab.failover: check discovery CNAMEs and gitlab-ssh records [cookbooks] - 10https://gerrit.wikimedia.org/r/1320977 (https://phabricator.wikimedia.org/T430655) (owner: 10Arnaudb) [17:38:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.38% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [17:41:03] FIRING: MediaWikiEditFailures: Elevated MediaWiki edit failures (session_loss) for cluster - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://grafana.wikimedia.org/d/000000208/edit-count?orgId=1&viewPanel=13 - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiEditFailures [17:43:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.41% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [17:46:02] (03PS1) 10Ebernhardson: cirrus sup: Enable first class redirect handling [deployment-charts] - 10https://gerrit.wikimedia.org/r/1322039 (https://phabricator.wikimedia.org/T204089) [17:46:03] RESOLVED: MediaWikiEditFailures: Elevated MediaWiki edit failures (session_loss) for cluster - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://grafana.wikimedia.org/d/000000208/edit-count?orgId=1&viewPanel=13 - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiEditFailures [17:46:47] RECOVERY - SSH on urldownloader2005 is OK: SSH OK - OpenSSH_10.0p2 Debian-7+deb13u4 (protocol 2.0) https://wikitech.wikimedia.org/wiki/SSH/monitoring [17:47:14] uh [17:47:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 18.94% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [17:51:54] cdobbins@cumin1003 reimage (PID 1298858) is awaiting input [17:52:11] 06SRE, 10SRE-Access-Requests: Grant access to analytics-privatedata-users for EAlbizzati-WMF - https://phabricator.wikimedia.org/T434034#12192669 (10jcrespo) Additionally, if you have any doubt or problem logging it, you can reopen the ticket or contact the SRE on clinic duty directly. Thank you. [17:52:53] PROBLEM - SSH on urldownloader2005 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/SSH/monitoring [17:53:15] (03PS1) 10Andrew Bogott: ceph.conf: pass slow ops threshold config to template [puppet] - 10https://gerrit.wikimedia.org/r/1322044 (https://phabricator.wikimedia.org/T429387) [17:54:05] FIRING: KubernetesCalicoDown: ml-serve1015.eqiad.wmnet is not running calico-node Pod - https://wikitech.wikimedia.org/wiki/Calico#Operations - https://grafana.wikimedia.org/d/G8zPL7-Wz/?var-dc=eqiad%20prometheus%2Fk8s-mlserve&var-instance=ml-serve1015.eqiad.wmnet - https://alerts.wikimedia.org/?q=alertname%3DKubernetesCalicoDown [17:54:22] (03CR) 10Andrew Bogott: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1322044 (https://phabricator.wikimedia.org/T429387) (owner: 10Andrew Bogott) [17:55:13] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for iberker - https://phabricator.wikimedia.org/T433258#12192705 (10jcrespo) [17:55:58] !log cdobbins@cumin1003 START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie [17:56:09] 10ops-eqsin, 06SRE, 06DC-Ops, 06Traffic: cp5022 is unreachable - https://phabricator.wikimedia.org/T414411#12192715 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by cdobbins@cumin1003 for host cp5022.eqsin.wmnet with OS trixie [17:57:18] (03CR) 10Andrew Bogott: [C:03+2] ceph.conf: pass slow ops threshold config to template [puppet] - 10https://gerrit.wikimedia.org/r/1322044 (https://phabricator.wikimedia.org/T429387) (owner: 10Andrew Bogott) [17:57:54] jouncebot: nowandnext [17:57:55] For the next 0 hour(s) and 2 minute(s): Cloud Services/Technical Documentation weekly deploy (Toolhub, Developer portal, Striker) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260806T1700) [17:57:55] For the next 0 hour(s) and 2 minute(s): MediaWiki infrastructure (UTC late) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260806T1700) [17:57:55] In 0 hour(s) and 2 minute(s): MediaWiki train - Utc-0+Utc-7 Version (secondary timeslot) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260806T1800) [17:58:30] 06SRE, 10SRE-swift-storage, 10MediaWiki-extensions-Score, 06Reader Experience Team: Add cache key information to metadata json - https://phabricator.wikimedia.org/T257093#12192725 (10HFan-WMF) thanks so much for the tag-in here :D -- while REx is listed as the maintainer, I have to admit that we have littl... [17:58:43] the train was done on the EU time, so I steal it to deploy a change [17:59:05] (03PS3) 10Krinkle: mediawiki: Disable legacy `short_urls` on vhosts where it does not work (take 2) [puppet] - 10https://gerrit.wikimedia.org/r/1321603 (https://phabricator.wikimedia.org/T107188) [17:59:08] (03CR) 10Ladsgroup: [C:03+2] mediawiki: Disable legacy `short_urls` on vhosts where it does not work (take 2) [puppet] - 10https://gerrit.wikimedia.org/r/1321603 (https://phabricator.wikimedia.org/T107188) (owner: 10Krinkle) [17:59:11] (03CR) 10Ladsgroup: [V:03+2 C:03+2] mediawiki: Disable legacy `short_urls` on vhosts where it does not work (take 2) [puppet] - 10https://gerrit.wikimedia.org/r/1321603 (https://phabricator.wikimedia.org/T107188) (owner: 10Krinkle) [18:00:05] jnuche and jeena: How many deployers does it take to do MediaWiki train - Utc-0+Utc-7 Version (secondary timeslot) deploy? (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260806T1800). [18:02:48] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for iberker - https://phabricator.wikimedia.org/T433258#12192769 (10jcrespo) a:05SCherukuwada→03jcrespo Thank you @SCherukuwada . I will deploy Web/Ldap only/ssh-less access now (the lowest privileged level) -which is implicitl... [18:04:15] FIRING: [2x] NodeBGPSessionStatusNotEstablished: Kubernetes node ml-serve1015 has a BGP session which is not in the 'established' state. [18:09:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [18:12:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.69% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [18:17:04] !log ladsgroup@deploy1003 Started scap sync-world: Deploying gerrit:1321603 (T107188) [18:17:09] T107188: Sunset ShortUrl extension in favour of UrlShortener extension - https://phabricator.wikimedia.org/T107188 [18:17:12] (03CR) 10Clare Ming: [C:03+2] Test Kitchen UI: Deploy v1.5.2 release to staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321824 (https://phabricator.wikimedia.org/T432629) (owner: 10Clare Ming) [18:17:34] (03CR) 10Clare Ming: [C:03+2] Test Kitchen UI: Deploy v1.5.2 release to production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321827 (https://phabricator.wikimedia.org/T432629) (owner: 10Clare Ming) [18:18:14] !log ladsgroup@deploy1003 Stopping before sync operations [18:19:34] (03Merged) 10jenkins-bot: Test Kitchen UI: Deploy v1.5.2 release to staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321824 (https://phabricator.wikimedia.org/T432629) (owner: 10Clare Ming) [18:19:46] !log ladsgroup@deploy1003 Started scap sync-world: Deploying gerrit:1321603 (T107188) [18:20:17] (03Merged) 10jenkins-bot: Test Kitchen UI: Deploy v1.5.2 release to production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321827 (https://phabricator.wikimedia.org/T432629) (owner: 10Clare Ming) [18:23:05] 06SRE, 10SRE-Access-Requests, 10LDAP-Access-Requests: Grant Access to wmf for Chandler Diggs - https://phabricator.wikimedia.org/T433302#12192871 (10CDiggs-WMF) @Dzahn Just figured out the entire process and now I'm in Matomo today, thanks for your help on this! [18:23:53] RECOVERY - SSH on urldownloader2005 is OK: SSH OK - OpenSSH_10.0p2 Debian-7+deb13u4 (protocol 2.0) https://wikitech.wikimedia.org/wiki/SSH/monitoring [18:25:27] !log ladsgroup@deploy1003 Finished scap sync-world: Deploying gerrit:1321603 (T107188) (duration: 06m 08s) [18:25:32] T107188: Sunset ShortUrl extension in favour of UrlShortener extension - https://phabricator.wikimedia.org/T107188 [18:26:34] 10ops-eqiad, 06SRE, 06DC-Ops: Alert for device ps1-e2-eqiad.mgmt.eqiad.wmnet - PDU sensor over limit - https://phabricator.wikimedia.org/T434118#12192894 (10VRiley-WMF) 05Open→03Resolved a:03VRiley-WMF Will monitor this. [18:26:53] PROBLEM - SSH on urldownloader2005 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/SSH/monitoring [18:33:03] (03PS2) 10Aude: Turn on feature flag for custom lists for betawiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1321639 (https://phabricator.wikimedia.org/T434027) (owner: 10LorenMora) [18:33:03] (03PS1) 10Aude: Enable ReadingLists for all logged-in users on beta cluster [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1322060 (https://phabricator.wikimedia.org/T434213) [18:34:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.17% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [18:34:21] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Thursday, August 06 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-i" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1322060 (https://phabricator.wikimedia.org/T434213) (owner: 10Aude) [18:34:38] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Thursday, August 06 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-i" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1321639 (https://phabricator.wikimedia.org/T434027) (owner: 10LorenMora) [18:39:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.17% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [18:40:39] 10ops-eqiad, 06SRE, 10Data-Persistence-Backup, 06DC-Ops: db1245 crashed - https://phabricator.wikimedia.org/T431115#12192955 (10VRiley-WMF) Hey @jcrespo I updated the firmware and reseated some of the fans. Would you be able to try to test putting a load on it? I was looking for thermal paste, but can't se... [18:50:20] 06SRE, 10SRE-Access-Requests, 10LDAP-Access-Requests: Grant Access to wmf for Chandler Diggs - https://phabricator.wikimedia.org/T433302#12192982 (10Dzahn) @CDiggs-WMF cool! thanks for confirming. it's always the best to end tickets with the user actually verifying. cheers. [18:51:52] (03PS2) 10Anzx: tcywiki: update logos for 10years anniversary [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1322054 (https://phabricator.wikimedia.org/T434176) [18:52:12] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Thursday, August 06 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-i" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1322054 (https://phabricator.wikimedia.org/T434176) (owner: 10Anzx) [18:52:51] (03PS3) 10Anzx: outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1322038 (https://phabricator.wikimedia.org/T431959) [18:53:00] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Thursday, August 06 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-i" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1322038 (https://phabricator.wikimedia.org/T431959) (owner: 10Anzx) [18:59:32] cdobbins@cumin1003 reimage (PID 1298858) is awaiting input [19:00:44] !log cdobbins@cumin1003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp5022.eqsin.wmnet with OS trixie [19:00:53] 10ops-eqsin, 06SRE, 06DC-Ops, 06Traffic: cp5022 is unreachable - https://phabricator.wikimedia.org/T414411#12192999 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by cdobbins@cumin1003 for host cp5022.eqsin.wmnet with OS trixie executed with errors: - cp5022 (**FAIL**) - Removed from... [19:20:47] RECOVERY - SSH on urldownloader2005 is OK: SSH OK - OpenSSH_10.0p2 Debian-7+deb13u4 (protocol 2.0) https://wikitech.wikimedia.org/wiki/SSH/monitoring [19:21:39] (03PS1) 10Chlod Alejandro: admin: add shell and key for chlod [puppet] - 10https://gerrit.wikimedia.org/r/1322077 (https://phabricator.wikimedia.org/T433791) [19:23:43] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to deployment for Chlod Alejandro - https://phabricator.wikimedia.org/T433791#12193056 (10Chlod) ^ Done! I've also added an SSH key for my backup Yubikey. If additional permission is needed for that or if that's something that should only be... [19:24:42] RESOLVED: [2x] ProbeDown: Service urldownloader2005:8080 has failed probes (http_url_downloader_wikimedia_org_ip4) - https://wikitech.wikimedia.org/wiki/Url-downloader - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [19:28:35] (03PS1) 10Cwhite: beta-logs: enable and configure logs-api [puppet] - 10https://gerrit.wikimedia.org/r/1322080 (https://phabricator.wikimedia.org/T434114) [19:28:39] (03PS1) 10Cwhite: opensearch httpd: add manage_local_users toggle [puppet] - 10https://gerrit.wikimedia.org/r/1322081 (https://phabricator.wikimedia.org/T434114) [19:31:27] !log cjming@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply [19:31:54] !log cjming@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply [19:33:22] (03CR) 10CDanis: [C:03+1] "Looks good, thanks!" [puppet] - 10https://gerrit.wikimedia.org/r/1321987 (https://phabricator.wikimedia.org/T429175) (owner: 10Ssingh) [19:40:59] !log cjming@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply [19:41:28] !log cjming@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply [20:00:05] RoanKattouw, urbanecm, TheresNoTime, kindrobot, and cjming: #bothumor I � Unicode. All rise for UTC late backport window deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260806T2000). [20:00:05] toni_, chlod, ebernhardson, aude, and anzx: A patch you scheduled for UTC late backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [20:00:07] o/ [20:00:10] \o [20:00:10] hi [20:00:13] \o/ [20:00:19] mine are patches for the beta cluster only [20:01:38] o/ [20:02:39] i can deploy for those who need a deployer but i have to cut out before the backport window ends in about 45 minutes [20:03:04] my config patch is safe to bundle with others [20:03:26] think they are all config patches so should be quicker [20:04:21] toni_: you're first in the queue - do you want me to take care of it? [20:04:45] i also don't mind if my config paych is bundled with other [20:05:02] likewise here; my patch is simple enough and shouldn't cause any trouble [20:05:17] sure, just let me know what I need to do. not 100% sure how to test, it's a new docroot that nothing references yet. [20:05:26] (03PS2) 10Cwhite: opensearch httpd: add manage_local_users toggle [puppet] - 10https://gerrit.wikimedia.org/r/1322081 (https://phabricator.wikimedia.org/T434114) [20:05:52] cool - i'll deploy ebernhardson's, aude's, chlod's, and anzx's patches out together then after doing toni_'s patch [20:05:59] thx [20:05:59] thank you! [20:06:05] ok [20:06:10] gotcha, thanks! [20:06:21] (03PS4) 10Tsevener: Add new Apple app site association file for Test Wiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1315128 (https://phabricator.wikimedia.org/T432412) [20:07:52] (03CR) 10TrainBranchBot: [C:03+2] "Approved by cjming@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1315128 (https://phabricator.wikimedia.org/T432412) (owner: 10Tsevener) [20:08:00] (03PS3) 10Aaron Schulz: Rename api-gateway to common-application-gateway [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319560 (https://phabricator.wikimedia.org/T434112) [20:09:14] (03Merged) 10jenkins-bot: Add new Apple app site association file for Test Wiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1315128 (https://phabricator.wikimedia.org/T432412) (owner: 10Tsevener) [20:09:24] !log cjming@deploy1003 Started scap sync-world: Backport for [[gerrit:1315128|Add new Apple app site association file for Test Wiki (T432412)]] [20:09:29] T432412: [Eng] iOS - Update app site association file - https://phabricator.wikimedia.org/T432412 [20:11:14] !log cjming@deploy1003 cjming, tsev: Backport for [[gerrit:1315128|Add new Apple app site association file for Test Wiki (T432412)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [20:11:47] toni_: gtg? lmk when to sync - i'm not sure how to test either tbh [20:11:53] 06SRE, 10SRE-swift-storage, 10MediaWiki-extensions-Score, 06Reader Experience Team: Add cache key information to metadata json - https://phabricator.wikimedia.org/T257093#12193179 (10TheDJ) There is no rush, it’s just that we have 315GB of storage of which a significant part isn’t used anymore because we s... [20:14:17] yeah, really I'm just testing the https://test.wikipedia.org/.well-known/apple-app-site-association files, they look unchanged against mw debug extension, which is what I expect since I don't have the puppet change pointing to it yet (that'll be a followup patch). So I think we're good? No regressions [20:14:35] cool - then syncing [20:14:40] !log cjming@deploy1003 cjming, tsev: Continuing with deployment [20:18:01] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [20:18:46] !log cjming@deploy1003 Finished scap sync-world: Backport for [[gerrit:1315128|Add new Apple app site association file for Test Wiki (T432412)]] (duration: 09m 22s) [20:18:52] T432412: [Eng] iOS - Update app site association file - https://phabricator.wikimedia.org/T432412 [20:20:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.41% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [20:20:18] ok - idk that i feel comfortable shipping everyone else's patches out together -- some of them probably should be staggered [20:20:25] here goes [20:20:29] thanks! [20:22:55] (03CR) 10TrainBranchBot: [C:03+2] "Approved by cjming@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319939 (https://phabricator.wikimedia.org/T430409) (owner: 10Chlod Alejandro) [20:22:56] (03CR) 10TrainBranchBot: [C:03+2] "Approved by cjming@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320230 (https://phabricator.wikimedia.org/T204089) (owner: 10Ebernhardson) [20:23:57] FIRING: ProbeDown: Service chart-renderer:30443 has failed probes (http_chart-renderer_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#chart-renderer:30443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [20:24:06] (03Merged) 10jenkins-bot: Revert "frwiki: change to Wikipedia 25 logo" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319939 (https://phabricator.wikimedia.org/T430409) (owner: 10Chlod Alejandro) [20:24:10] (03Merged) 10jenkins-bot: cirrus: Enable building redirect documents [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320230 (https://phabricator.wikimedia.org/T204089) (owner: 10Ebernhardson) [20:24:25] !log cjming@deploy1003 Started scap sync-world: Backport for [[gerrit:1319939|Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230|cirrus: Enable building redirect documents (T204089)]] [20:24:32] T430409: Requesting temporary logo change for fr.wikipedia.org - https://phabricator.wikimedia.org/T430409 [20:24:32] T204089: CirrusSearch: Add filter for exclusion of redirects or finding only them - https://phabricator.wikimedia.org/T204089 [20:24:35] (03CR) 10Ebernhardson: [C:03+2] cirrus sup: Enable first class redirect handling [deployment-charts] - 10https://gerrit.wikimedia.org/r/1322039 (https://phabricator.wikimedia.org/T204089) (owner: 10Ebernhardson) [20:25:15] RESOLVED: [2x] PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-int releases routed via main at eqiad: 24.26% idle - https://bit.ly/wmf-fpmsat - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [20:25:30] FIRING: [2x] PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-int releases routed via main at eqiad: 25% idle - https://bit.ly/wmf-fpmsat - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [20:26:15] !log cjming@deploy1003 cjming, ebernhardson, chlod: Backport for [[gerrit:1319939|Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230|cirrus: Enable building redirect documents (T204089)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [20:26:19] aude: you can deploy right? [20:26:36] ebernhardson: and chlod: ok to sync? [20:26:45] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-int releases routed via main at eqiad: 24.26% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-int&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [20:26:51] all looks good to me :) [20:26:59] cjming: yup [20:27:02] (03Merged) 10jenkins-bot: cirrus sup: Enable first class redirect handling [deployment-charts] - 10https://gerrit.wikimedia.org/r/1322039 (https://phabricator.wikimedia.org/T204089) (owner: 10Ebernhardson) [20:27:04] !log cjming@deploy1003 cjming, ebernhardson, chlod: Continuing with deployment [20:27:05] I can deploy mine though they are changes to the beta cluster only [20:27:41] (03PS1) 10Cwhite: profile: set default accounts and groups for api::httpd_proxy [puppet] - 10https://gerrit.wikimedia.org/r/1322115 (https://phabricator.wikimedia.org/T434114) [20:28:14] anzx: do you need a deployer? [20:28:57] RESOLVED: ProbeDown: Service chart-renderer:30443 has failed probes (http_chart-renderer_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#chart-renderer:30443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [20:29:08] cjming: yes [20:29:26] aude: are you comfortable deploying anzx's changes? i have to run here soon [20:30:17] yes i can deploy them [20:30:45] awesome thanks -- then after the current deployment finishes, i'll hand over the reigns to you if that's ok [20:31:06] !log cjming@deploy1003 Finished scap sync-world: Backport for [[gerrit:1319939|Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230|cirrus: Enable building redirect documents (T204089)]] (duration: 06m 41s) [20:31:13] T430409: Requesting temporary logo change for fr.wikipedia.org - https://phabricator.wikimedia.org/T430409 [20:31:14] T204089: CirrusSearch: Add filter for exclusion of redirects or finding only them - https://phabricator.wikimedia.org/T204089 [20:31:22] ebernhardson: chlod: live! [20:31:26] ok [20:31:29] thank you cjming! :D [20:31:29] aude: all yours - ty!! [20:31:52] (03CR) 10Cwhite: [C:03+2] profile: set default accounts and groups for api::httpd_proxy [puppet] - 10https://gerrit.wikimedia.org/r/1322115 (https://phabricator.wikimedia.org/T434114) (owner: 10Cwhite) [20:32:04] !log ebernhardson@deploy1003 helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply [20:32:08] !log ebernhardson@deploy1003 helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply [20:32:12] cjming: thanks! [20:32:17] np! [20:34:20] i am taking a quick look at the patches [20:37:00] !log ebernhardson@deploy1003 helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply [20:37:07] !log ebernhardson@deploy1003 helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply [20:38:03] (03CR) 10TrainBranchBot: [C:03+2] "Approved by aude@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1322038 (https://phabricator.wikimedia.org/T431959) (owner: 10Anzx) [20:38:04] (03CR) 10TrainBranchBot: [C:03+2] "Approved by aude@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1321639 (https://phabricator.wikimedia.org/T434027) (owner: 10LorenMora) [20:38:04] (03CR) 10TrainBranchBot: [C:03+2] "Approved by aude@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1322054 (https://phabricator.wikimedia.org/T434176) (owner: 10Anzx) [20:39:37] (03Merged) 10jenkins-bot: outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1322038 (https://phabricator.wikimedia.org/T431959) (owner: 10Anzx) [20:39:41] (03Merged) 10jenkins-bot: Turn on feature flag for custom lists for betawiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1321639 (https://phabricator.wikimedia.org/T434027) (owner: 10LorenMora) [20:39:45] (03Merged) 10jenkins-bot: tcywiki: update logos for 10years anniversary [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1322054 (https://phabricator.wikimedia.org/T434176) (owner: 10Anzx) [20:39:57] !log aude@deploy1003 Started scap sync-world: Backport for [[gerrit:1322038|outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639|Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054|tcywiki: update logos for 10years anniversary (T434176)]] [20:40:04] T431959: Modify user group configuration settings for Outreach Wiki - https://phabricator.wikimedia.org/T431959 [20:40:05] T434027: Turn on feature flag for custom lists for betawiki - https://phabricator.wikimedia.org/T434027 [20:40:05] T434176: Requesting temporary logo change for tcy.wikipedia.org (WP10) - https://phabricator.wikimedia.org/T434176 [20:41:34] !log ebernhardson@deploy1003 helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply [20:41:38] !log ebernhardson@deploy1003 helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply [20:41:46] !log aude@deploy1003 lmora, aude, anzx: Backport for [[gerrit:1322038|outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639|Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054|tcywiki: update logos for 10years anniversary (T434176)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be [20:41:46] verified there. [20:41:57] aude: checking [20:42:14] thanks [20:43:27] aude: both changes looks good, ok to continue [20:44:01] ok [20:44:06] !log aude@deploy1003 lmora, aude, anzx: Continuing with deployment [20:48:08] !log aude@deploy1003 Finished scap sync-world: Backport for [[gerrit:1322038|outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639|Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054|tcywiki: update logos for 10years anniversary (T434176)]] (duration: 08m 12s) [20:48:16] T431959: Modify user group configuration settings for Outreach Wiki - https://phabricator.wikimedia.org/T431959 [20:48:17] T434027: Turn on feature flag for custom lists for betawiki - https://phabricator.wikimedia.org/T434027 [20:48:21] T434176: Requesting temporary logo change for tcy.wikipedia.org (WP10) - https://phabricator.wikimedia.org/T434176 [20:48:31] aude: thanks for deploying [20:49:08] you're welcome! [20:49:17] I have one more patch for the beta cluster [20:49:52] (03CR) 10TrainBranchBot: [C:03+2] "Approved by aude@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1322060 (https://phabricator.wikimedia.org/T434213) (owner: 10Aude) [20:53:18] (03CR) 10CI reject: [V:04-1] Enable ReadingLists for all logged-in users on beta cluster [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1322060 (https://phabricator.wikimedia.org/T434213) (owner: 10Aude) [20:53:51] (03PS2) 10Cwhite: beta-logs: enable and configure logs-api [puppet] - 10https://gerrit.wikimedia.org/r/1322080 (https://phabricator.wikimedia.org/T434114) [20:55:39] (03PS2) 10Aude: Enable ReadingLists for all logged-in users on beta cluster [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1322060 (https://phabricator.wikimedia.org/T434213) [21:02:20] FIRING: [2x] CirrusSearchSaneitizerFixRateTooHigh: MediaWiki CirrusSearch Saneitizer is fixing an abnormally high number of documents in cloudelastic - https://wikitech.wikimedia.org/wiki/Search/CirrusStreamingUpdater#San(e)itizing - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchSaneitizerFixRateTooHigh [21:02:36] aude are you still deploying? [21:02:47] I would like to deploy two security patches [21:02:49] yes [21:03:23] (03CR) 10TrainBranchBot: "Approved by aude@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1322060 (https://phabricator.wikimedia.org/T434213) (owner: 10Aude) [21:03:35] if the patch merges this time, then it shouldn't take long [21:04:16] aude fingers crossed [21:04:48] i had to rebase the second patch, since it was a chain of patches (and is beta cluster only) [21:06:52] (03Merged) 10jenkins-bot: Enable ReadingLists for all logged-in users on beta cluster [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1322060 (https://phabricator.wikimedia.org/T434213) (owner: 10Aude) [21:07:20] FIRING: [3x] CirrusSearchSaneitizerFixRateTooHigh: MediaWiki CirrusSearch Saneitizer is fixing an abnormally high number of documents in cloudelastic - https://wikitech.wikimedia.org/wiki/Search/CirrusStreamingUpdater#San(e)itizing - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchSaneitizerFixRateTooHigh [21:07:37] I am done [21:07:51] 21:07:02 Skipping sync since all commits were beta/labs-only changes. Operation completed. [21:09:10] thank you aude!! [21:09:52] getting started with the first security patch [21:20:41] (03PS1) 10Cwhite: profile: add parameter to fix duplicate declaration in beta [puppet] - 10https://gerrit.wikimedia.org/r/1322144 (https://phabricator.wikimedia.org/T434114) [21:23:39] (03PS1) 10Aude: Fix beta config for wgReadingListBetaFeature [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1322145 (https://phabricator.wikimedia.org/T434213) [21:23:41] (03CR) 10Cwhite: [C:03+2] profile: add parameter to fix duplicate declaration in beta [puppet] - 10https://gerrit.wikimedia.org/r/1322144 (https://phabricator.wikimedia.org/T434114) (owner: 10Cwhite) [21:24:15] maryum when you are done, I have one more patch for the beta cluster since my previous patch is causing an issue with scap on the beta cluster [21:24:22] no hurry though [21:24:38] aude okay I'll let you know, still on the first patch and have one more [21:24:43] thanks [21:26:28] (03PS1) 10Cwhite: profile: fix typo in opensearch::api::httpd_proxy [puppet] - 10https://gerrit.wikimedia.org/r/1322149 (https://phabricator.wikimedia.org/T434114) [21:26:57] (03CR) 10Cwhite: [V:03+2 C:03+2] profile: fix typo in opensearch::api::httpd_proxy [puppet] - 10https://gerrit.wikimedia.org/r/1322149 (https://phabricator.wikimedia.org/T434114) (owner: 10Cwhite) [21:29:12] !log Deploy security patch for T434189 [21:29:14] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [21:29:21] running scap for the second patch [21:30:02] (03CR) 10Cwhite: [C:03+2] alertmanager: 24h repeat interval for TSP slack [puppet] - 10https://gerrit.wikimedia.org/r/1321900 (owner: 10Clément Goubert) [21:39:05] !log Deploy security patch for T433070 [21:39:08] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [21:39:13] aude finished with scap [21:39:31] thanks [21:39:46] (03PS1) 10Cwhite: beta-logs: disable ssl requirement for logs-api [puppet] - 10https://gerrit.wikimedia.org/r/1322159 (https://phabricator.wikimedia.org/T434114) [21:39:46] (03CR) 10TrainBranchBot: [C:03+2] "Approved by aude@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1322145 (https://phabricator.wikimedia.org/T434213) (owner: 10Aude) [21:41:45] (03Merged) 10jenkins-bot: Fix beta config for wgReadingListBetaFeature [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1322145 (https://phabricator.wikimedia.org/T434213) (owner: 10Aude) [21:42:26] done but monitoring the beta cluster updates to make sure it's good now [21:46:52] (03CR) 10Cwhite: [C:03+2] beta-logs: disable ssl requirement for logs-api [puppet] - 10https://gerrit.wikimedia.org/r/1322159 (https://phabricator.wikimedia.org/T434114) (owner: 10Cwhite) [21:48:27] (03CR) 10Cwhite: [C:03+1] "LGTM!" [puppet] - 10https://gerrit.wikimedia.org/r/1321653 (https://phabricator.wikimedia.org/T434114) (owner: 10Ahmon Dancy) [21:50:06] (03CR) 10Cwhite: [C:03+1] scap: Add optional logstash credentials file (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1321653 (https://phabricator.wikimedia.org/T434114) (owner: 10Ahmon Dancy) [21:54:43] FIRING: KubernetesCalicoDown: ml-serve1015.eqiad.wmnet is not running calico-node Pod - https://wikitech.wikimedia.org/wiki/Calico#Operations - https://grafana.wikimedia.org/d/G8zPL7-Wz/?var-dc=eqiad%20prometheus%2Fk8s-mlserve&var-instance=ml-serve1015.eqiad.wmnet - https://alerts.wikimedia.org/?q=alertname%3DKubernetesCalicoDown [22:04:43] FIRING: [2x] NodeBGPSessionStatusNotEstablished: Kubernetes node ml-serve1015 has a BGP session which is not in the 'established' state. [22:06:24] (03PS1) 10DLynch: CodeMirror: turn on the new 2017 editor integration [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1322175 (https://phabricator.wikimedia.org/T432558) [22:09:43] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [22:32:33] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to deployment for Chlod Alejandro - https://phabricator.wikimedia.org/T433791#12193596 (10KFrancis) Hello all, the NDA is out for signatures. I'll confirm when it's complete. Thanks! [22:41:20] 06SRE, 10SRE-swift-storage, 10MediaWiki-extensions-Score, 06Reader Experience Team: Add cache key information to metadata json - https://phabricator.wikimedia.org/T257093#12193663 (10HFan-WMF) Ah! thanks for the info :) Looks like Timo is reviewing the patch (thank you Timo!), so I'll move this onto our te... [23:07:33] (03PS1) 10PipelineBot: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1322199 [23:08:40] (03PS1) 10PipelineBot: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1322201 [23:10:28] (03PS1) 10PipelineBot: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1322203 [23:41:51] PROBLEM - Host pki1002 is DOWN: PING CRITICAL - Packet loss = 100% [23:42:27] (03PS1) 10TrainBranchBot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1322217 [23:42:27] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1322217 (owner: 10TrainBranchBot) [23:43:24] FIRING: [15x] ProbeDown: Service pki1002:443 has failed probes (http_PKI_cassandra_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#pki1002:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [23:47:53] FIRING: [2x] JobUnavailable: Reduced availability for job cfssl in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [23:48:24] FIRING: [44x] ProbeDown: Service pki1002:443 has failed probes (http_PKI_aux_front_proxy_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#pki1002:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [23:55:28] (03Merged) 10jenkins-bot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1322217 (owner: 10TrainBranchBot)