[00:14:45] FIRING: WidespreadPuppetFailure: Puppet has failed in esams - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [00:23:00] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [00:39:45] RESOLVED: WidespreadPuppetFailure: Puppet has failed in esams - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [01:07:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [01:12:24] (03PS1) 10TrainBranchBot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1319959 [01:12:24] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1319959 (owner: 10TrainBranchBot) [01:25:18] (03Merged) 10jenkins-bot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1319959 (owner: 10TrainBranchBot) [01:42:40] FIRING: SystemdUnitFailed: generate_vrts_aliases.service on mx-in2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [01:58:00] FIRING: HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-rollback - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs-next - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [02:00:16] !log mwpresync@deploy1003 Started scap build-images: Publishing wmf/next image [02:07:03] !log mwpresync@deploy1003 Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) [02:08:00] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:13:00] FIRING: [3x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-3/1/2 (Transport: Hurricane Electric (dc4841.dal5)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [02:13:56] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:18:00] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [03:12:38] FIRING: CertAlmostExpired: gNMI TLS certificate for lsw1-d6-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [03:53:00] FIRING: [2x] GanetiBGPNoInPrefixes: ganeti2033 is not sending any prefix to lsw1-b7-codfw - group Ganeti4 - https://wikitech.wikimedia.org/wiki/Ganeti#GanetiBGPNoInPrefixes - https://alerts.wikimedia.org/?q=alertname%3DGanetiBGPNoInPrefixes [03:58:49] 06SRE, 06Data-Engineering: WE 5.4 FY 25/26: Improve automata detection at the edge and pass it to the refinery pipeline - https://phabricator.wikimedia.org/T396562#12177505 (10nshahquinn-wmf) [05:07:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [05:13:05] PROBLEM - OSPF status on cr1-eqiad is CRITICAL: OSPFv2: 8/9 UP : OSPFv3: 9/9 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [05:13:07] PROBLEM - OSPF status on cr2-esams is CRITICAL: OSPFv2: 4/5 UP : OSPFv3: 5/5 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [05:14:05] RECOVERY - OSPF status on cr1-eqiad is OK: OSPFv2: 9/9 UP : OSPFv3: 9/9 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [05:14:07] RECOVERY - OSPF status on cr2-esams is OK: OSPFv2: 5/5 UP : OSPFv3: 5/5 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [05:42:40] FIRING: SystemdUnitFailed: generate_vrts_aliases.service on mx-in2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [05:49:26] (03PS1) 10KartikMistry: Update cxserver to 2026-07-16-140518-production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319962 (https://phabricator.wikimedia.org/T97231) [05:58:00] FIRING: HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-rollback - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs-next - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [06:05:33] 10ops-codfw, 06SRE, 06DC-Ops, 06ServiceOps: Rerack mc-gp2006 and mc2046 - https://phabricator.wikimedia.org/T432716#12177582 (10ayounsi) Networking is no magic :) The server is still connected to the `B7` rack though, what is its new switch port ? For the re-image, the "hack" is to set the host's switch... [06:18:00] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [06:33:20] Deploying cxserver, staging first. [06:36:25] (03CR) 10KartikMistry: [C:03+2] Update cxserver to 2026-07-16-140518-production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319962 (https://phabricator.wikimedia.org/T97231) (owner: 10KartikMistry) [06:39:27] (03Merged) 10jenkins-bot: Update cxserver to 2026-07-16-140518-production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319962 (https://phabricator.wikimedia.org/T97231) (owner: 10KartikMistry) [06:51:22] (03CR) 10Dpogorzelski: [C:03+2] ml-serve: enlarge kubelet LV on GPU nodes [puppet] - 10https://gerrit.wikimedia.org/r/1307368 (https://phabricator.wikimedia.org/T431017) (owner: 10Dpogorzelski) [06:54:43] !log kartik@deploy1003 helmfile [staging] START helmfile.d/services/cxserver: apply [06:55:17] !log kartik@deploy1003 helmfile [staging] DONE helmfile.d/services/cxserver: apply [06:56:13] (03CR) 10Arnaudb: "same here, feel free to add me again when needed!" [puppet] - 10https://gerrit.wikimedia.org/r/1178874 (https://phabricator.wikimedia.org/T378028) (owner: 10AOkoth) [06:57:05] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, August 03 UTC morning backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1315895 (owner: 10Jdlrobson) [06:57:12] 06SRE, 06MediaWiki-Core-Platform-Team, 05Front-end Modernization: Introduce a Front-end Build Step for MediaWiki Skins and Extensions - https://phabricator.wikimedia.org/T279108#12177636 (10LSobanski) @MGoncalves-WMF could I ask you to clarify what is the ask for SRE here? [07:00:05] Amir1, urbanecm, and awight: gettimeofday() says it's time for UTC morning backport window. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260803T0700) [07:00:05] jdlrobson: A patch you scheduled for UTC morning backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [07:01:03] o/ [07:01:09] ill go ahead and spiderpigthis one patch. [07:03:11] (03CR) 10TrainBranchBot: [C:03+2] "Approved by jdlrobson@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1315895 (owner: 10Jdlrobson) [07:04:16] (03Merged) 10jenkins-bot: Enable page images on recipe namespace [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1315895 (owner: 10Jdlrobson) [07:04:58] !log jdlrobson@deploy1003 Started scap sync-world: Backport for [[gerrit:1315895|Enable page images on recipe namespace]] [07:06:17] (03CR) 10Ayounsi: "recheck" [software/homer] - 10https://gerrit.wikimedia.org/r/1319796 (owner: 10Ayounsi) [07:09:47] !log Drop renamed tables T425074 [07:09:50] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [07:09:51] T425074: Drop CheckUser tables from closed wikis - https://phabricator.wikimedia.org/T425074 [07:12:38] FIRING: CertAlmostExpired: gNMI TLS certificate for lsw1-d6-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [07:20:34] (03PS1) 10Marostegui: Revert "db1218: Disable notifications" [puppet] - 10https://gerrit.wikimedia.org/r/1319964 [07:21:30] !log jdlrobson@deploy1003 jdlrobson: Backport for [[gerrit:1315895|Enable page images on recipe namespace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [07:23:00] (03CR) 10Marostegui: [C:03+2] Revert "db1218: Disable notifications" [puppet] - 10https://gerrit.wikimedia.org/r/1319964 (owner: 10Marostegui) [07:23:27] 10ops-eqiad, 06SRE, 06DBA, 06DC-Ops: db1218 crashed - https://phabricator.wikimedia.org/T433565#12177673 (10Marostegui) I am going to start repooling this host now that both BIOS and idrac have been updated. [07:23:54] !log marostegui@cumin1003 START - Cookbook sre.mysql.pool pool db1218: Repool after a crash [07:24:07] 10ops-eqiad, 06SRE, 06DBA, 06DC-Ops: db1218 crashed - https://phabricator.wikimedia.org/T433565#12177675 (10ops-monitoring-bot) Starting pool of db1218 by marostegui@cumin1003: Repool after a crash [07:24:12] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-be1085.eqiad.wmnet with OS trixie [07:24:23] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12177676 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-be1085.eqiad.wmnet with OS trixie [07:24:53] 10ops-eqiad, 06SRE, 06DBA, 06DC-Ops: db1218 crashed - https://phabricator.wikimedia.org/T433565#12177677 (10Marostegui) 05Open→03Resolved a:03Marostegui Thanks for the help @VRiley-WMF and @CWilliams-WMF We will see if it crashes again. [07:24:58] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host ms-be2083.codfw.wmnet with OS trixie [07:25:06] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12177680 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host ms-be2083.codfw.wmnet with OS trixie [07:25:35] !log jdlrobson@deploy1003 jdlrobson: Continuing with deployment [07:26:38] (03CR) 10CI reject: [V:04-1] Add Python 3.14 support [software/homer] - 10https://gerrit.wikimedia.org/r/1319796 (owner: 10Ayounsi) [07:33:20] !log kartik@deploy1003 helmfile [codfw] START helmfile.d/services/cxserver: apply [07:33:50] !log kartik@deploy1003 helmfile [codfw] DONE helmfile.d/services/cxserver: apply [07:36:28] (03PS1) 10Dpogorzelski: ml-serve: lower kubelet LV target to 495G [puppet] - 10https://gerrit.wikimedia.org/r/1319968 (https://phabricator.wikimedia.org/T431017) [07:37:38] !log jdlrobson@deploy1003 Finished scap sync-world: Backport for [[gerrit:1315895|Enable page images on recipe namespace]] (duration: 32m 40s) [07:37:56] (03CR) 10Dpogorzelski: [C:03+2] ml-serve: lower kubelet LV target to 495G [puppet] - 10https://gerrit.wikimedia.org/r/1319968 (https://phabricator.wikimedia.org/T431017) (owner: 10Dpogorzelski) [07:38:03] !log mvernon@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage [07:38:50] !log kartik@deploy1003 helmfile [eqiad] START helmfile.d/services/cxserver: apply [07:39:09] (done) [07:39:18] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage [07:39:27] !log kartik@deploy1003 helmfile [eqiad] DONE helmfile.d/services/cxserver: apply [07:40:39] !log Updated cxsever to 2026-07-16-140518-production (T97231) [07:40:43] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [07:40:44] T97231: Put references before the dots for some languages - https://phabricator.wikimedia.org/T97231 [07:42:27] (03PS1) 10Marostegui: production-m1.sql.erb: Add new grants for etherpad [puppet] - 10https://gerrit.wikimedia.org/r/1319969 (https://phabricator.wikimedia.org/T433814) [07:42:56] (03CR) 10Marostegui: "This is a noop, and it is just for tracking. grants needs to be added to the DB." [puppet] - 10https://gerrit.wikimedia.org/r/1319969 (https://phabricator.wikimedia.org/T433814) (owner: 10Marostegui) [07:44:02] (03PS1) 10Slyngshede: data.yaml Offboarding jstep [puppet] - 10https://gerrit.wikimedia.org/r/1319970 [07:44:28] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage [07:46:55] (03PS1) 10Ozge: ml-services: update jina-embeddings image and tune staging inference config [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319971 (https://phabricator.wikimedia.org/T432717) [07:47:41] (03CR) 10Ozge: "final patch to make staging and prod identical and use the latest clean up image." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319971 (https://phabricator.wikimedia.org/T432717) (owner: 10Ozge) [07:48:25] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage [07:52:45] (03CR) 10Marostegui: sre.mysql.depool: regression of T427381 (031 comment) [cookbooks] - 10https://gerrit.wikimedia.org/r/1319509 (https://phabricator.wikimedia.org/T433564) (owner: 10CWilliams) [07:53:37] (03CR) 10Marostegui: [C:03+2] production-m1.sql.erb: Add new grants for etherpad [puppet] - 10https://gerrit.wikimedia.org/r/1319969 (https://phabricator.wikimedia.org/T433814) (owner: 10Marostegui) [07:58:51] (03PS1) 10Marostegui: production-m1.sql.erb: Fix username [puppet] - 10https://gerrit.wikimedia.org/r/1319972 (https://phabricator.wikimedia.org/T433814) [07:59:30] (03CR) 10Slyngshede: [C:03+2] data.yaml Offboarding jstep [puppet] - 10https://gerrit.wikimedia.org/r/1319970 (owner: 10Slyngshede) [08:06:55] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1085.eqiad.wmnet with OS trixie [08:07:10] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12177771 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-be1085.eqiad.wmnet with OS trixie completed: - ms-be1085 (**WARN*... [08:07:15] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2083.codfw.wmnet with OS trixie [08:07:23] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12177772 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-be2083.codfw.wmnet with OS trixie completed: - ms-be2083 (**PASS*... [08:08:29] (03PS1) 10Slyngshede: data.yaml Offboarding slong-wmf [puppet] - 10https://gerrit.wikimedia.org/r/1319973 [08:08:49] (03PS2) 10Marostegui: production-m1.sql.erb: Fix username [puppet] - 10https://gerrit.wikimedia.org/r/1319972 (https://phabricator.wikimedia.org/T433814) [08:08:49] (03PS1) 10Marostegui: production-m1.sql.erb: Add codfw etherpad new grants [puppet] - 10https://gerrit.wikimedia.org/r/1319974 (https://phabricator.wikimedia.org/T433814) [08:09:23] !log marostegui@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Repool after a crash [08:09:31] 10ops-eqiad, 06SRE, 06DBA, 06DC-Ops: db1218 crashed - https://phabricator.wikimedia.org/T433565#12177776 (10ops-monitoring-bot) Completed pooling of db1218 by marostegui@cumin1003: Repool after a crash [08:13:18] (03CR) 10Slyngshede: [C:03+2] data.yaml Offboarding slong-wmf [puppet] - 10https://gerrit.wikimedia.org/r/1319973 (owner: 10Slyngshede) [08:15:09] (03CR) 10Marostegui: [C:03+2] production-m1.sql.erb: Fix username [puppet] - 10https://gerrit.wikimedia.org/r/1319972 (https://phabricator.wikimedia.org/T433814) (owner: 10Marostegui) [08:15:23] (03CR) 10Marostegui: [C:03+2] production-m1.sql.erb: Add codfw etherpad new grants [puppet] - 10https://gerrit.wikimedia.org/r/1319974 (https://phabricator.wikimedia.org/T433814) (owner: 10Marostegui) [08:21:41] (03CR) 10Kevin Bazira: [C:03+1] ml-services: update jina-embeddings image and tune staging inference config [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319971 (https://phabricator.wikimedia.org/T432717) (owner: 10Ozge) [08:22:04] (03CR) 10Ozge: [C:03+2] ml-services: update jina-embeddings image and tune staging inference config [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319971 (https://phabricator.wikimedia.org/T432717) (owner: 10Ozge) [08:22:25] (03PS4) 10CWilliams: sre.mysql.depool: regression of T427381 [cookbooks] - 10https://gerrit.wikimedia.org/r/1319509 (https://phabricator.wikimedia.org/T433564) [08:22:40] (03PS1) 10Slyngshede: data.yaml Offboarding Sarmbruster [puppet] - 10https://gerrit.wikimedia.org/r/1320016 [08:23:53] (03CR) 10Slyngshede: [C:03+2] data.yaml Offboarding Sarmbruster [puppet] - 10https://gerrit.wikimedia.org/r/1320016 (owner: 10Slyngshede) [08:24:09] (03CR) 10CWilliams: sre.mysql.depool: regression of T427381 (031 comment) [cookbooks] - 10https://gerrit.wikimedia.org/r/1319509 (https://phabricator.wikimedia.org/T433564) (owner: 10CWilliams) [08:24:59] (03Merged) 10jenkins-bot: ml-services: update jina-embeddings image and tune staging inference config [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319971 (https://phabricator.wikimedia.org/T432717) (owner: 10Ozge) [08:25:15] (03CR) 10Arnaudb: "agreeing with @dzahn@wikimedia.org here. I think this is worth merging as a redundant safeguard. I'm not sure about what you'd need to ana" [puppet] - 10https://gerrit.wikimedia.org/r/1308081 (https://phabricator.wikimedia.org/T431398) (owner: 10Arnaudb) [08:25:25] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-be1086.eqiad.wmnet with OS trixie [08:25:35] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12177818 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-be1086.eqiad.wmnet with OS trixie [08:26:00] jouncebot: nowandnext [08:26:00] No deployments scheduled for the next 1 hour(s) and 33 minute(s) [08:26:01] In 1 hour(s) and 33 minute(s): MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260803T1000) [08:27:07] !log ozge@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . [08:29:45] (03CR) 10Arthur taylor: [C:03+2] wikidata-query-gui: bump to latest version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319451 (https://phabricator.wikimedia.org/T433233) (owner: 10Arthur taylor) [08:33:02] (03Merged) 10jenkins-bot: wikidata-query-gui: bump to latest version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319451 (https://phabricator.wikimedia.org/T433233) (owner: 10Arthur taylor) [08:33:46] (03CR) 10Slyngshede: [C:03+2] Release version 0.1.18 [software/bitu] - 10https://gerrit.wikimedia.org/r/1319812 (owner: 10Slyngshede) [08:34:24] !log arthurtaylor@deploy1003 helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply [08:34:54] !log arthurtaylor@deploy1003 helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply [08:35:47] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host ms-be2084.codfw.wmnet with OS trixie [08:35:57] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12177851 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host ms-be2084.codfw.wmnet with OS trixie [08:37:17] (03Merged) 10jenkins-bot: Release version 0.1.18 [software/bitu] - 10https://gerrit.wikimedia.org/r/1319812 (owner: 10Slyngshede) [08:37:31] !log arthurtaylor@deploy1003 helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply [08:37:44] (03PS1) 10KartikMistry: cxserver: Add missing referencePunctuation [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320102 (https://phabricator.wikimedia.org/T97231) [08:37:56] !log arthurtaylor@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply [08:38:05] !log arthurtaylor@deploy1003 helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply [08:38:27] !log arthurtaylor@deploy1003 helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply [08:39:43] !log mvernon@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage [08:41:30] (03CR) 10Arnaudb: "I'm wondering if we would benefit from using the discovery records here, avoiding going through the CDN layers" [puppet] - 10https://gerrit.wikimedia.org/r/1314095 (owner: 10Dzahn) [08:41:31] (03CR) 10KartikMistry: [C:03+2] cxserver: Add missing referencePunctuation [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320102 (https://phabricator.wikimedia.org/T97231) (owner: 10KartikMistry) [08:44:00] (03Merged) 10jenkins-bot: cxserver: Add missing referencePunctuation [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320102 (https://phabricator.wikimedia.org/T97231) (owner: 10KartikMistry) [08:44:30] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage [08:46:02] (03CR) 10Arnaudb: [C:03+1] "the context is from 1306459 where we prepared for gitlab to move behind CDN. it was to be able to revert quicker if anything was going wro" [dns] - 10https://gerrit.wikimedia.org/r/1310135 (owner: 10Dzahn) [08:50:35] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage [08:52:28] 06SRE, 06Data-Platform-SRE (2026-07-03 - 2026-07-31): archiva1002 has stale jobs in /var/cache/archiva that uses all the disk space - https://phabricator.wikimedia.org/T425083#12177906 (10jcrespo) I got the following backup error, I wonder if it is related to these ongoing issues: ` 02-Aug 04:05 backup1014.eq... [08:53:04] (03PS1) 10Arthur taylor: wikidata-query-gui: bump to latest version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320103 (https://phabricator.wikimedia.org/T433234) [08:53:05] (03PS1) 10Effie Mouzeli: profile::pki::root_ca: add new intermediates for liftwing [puppet] - 10https://gerrit.wikimedia.org/r/1320104 (https://phabricator.wikimedia.org/T353511) [08:53:44] (03CR) 10CI reject: [V:04-1] profile::pki::root_ca: add new intermediates for liftwing [puppet] - 10https://gerrit.wikimedia.org/r/1320104 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [08:53:55] (03CR) 10Marostegui: [C:03+1] "Thanks!" [cookbooks] - 10https://gerrit.wikimedia.org/r/1319509 (https://phabricator.wikimedia.org/T433564) (owner: 10CWilliams) [08:54:15] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage [08:54:55] (03PS2) 10Effie Mouzeli: profile::pki::root_ca: add new intermediates for liftwing [puppet] - 10https://gerrit.wikimedia.org/r/1320104 (https://phabricator.wikimedia.org/T353511) [08:56:38] (03CR) 10Audrey Penven: [C:03+1] wikidata-query-gui: bump to latest version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320103 (https://phabricator.wikimedia.org/T433234) (owner: 10Arthur taylor) [08:58:06] (03PS2) 10Arthur taylor: wikidata-query-gui: bump to latest version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320103 (https://phabricator.wikimedia.org/T433234) [08:58:32] (03PS5) 10CWilliams: sre.mysql.depool: regression of T427381 [cookbooks] - 10https://gerrit.wikimedia.org/r/1319509 (https://phabricator.wikimedia.org/T433564) [09:01:50] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1086.eqiad.wmnet with OS trixie [09:02:00] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12177940 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-be1086.eqiad.wmnet with OS trixie completed: - ms-be1086 (**PASS*... [09:07:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:10:01] 06SRE, 06Data-Platform-SRE (2026-07-03 - 2026-07-31): archiva1002 has stale jobs in /var/cache/archiva that uses all the disk space - https://phabricator.wikimedia.org/T425083#12177946 (10jcrespo) 05Stalled→03Open The host seems at the moment hung, either extremely slow or completely unresponsive through s... [09:10:18] 06SRE, 06Data-Platform-SRE (2026-07-03 - 2026-07-31): archiva1002 has stale jobs in /var/cache/archiva that uses all the disk space - https://phabricator.wikimedia.org/T425083#12177949 (10jcrespo) p:05Low→03High [09:11:14] (03CR) 10CWilliams: [C:03+2] sre.mysql.depool: regression of T427381 [cookbooks] - 10https://gerrit.wikimedia.org/r/1319509 (https://phabricator.wikimedia.org/T433564) (owner: 10CWilliams) [09:12:33] (03CR) 10Ladsgroup: Allow setting a separate thumbUrl in production [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319920 (https://phabricator.wikimedia.org/T427465) (owner: 10Ladsgroup) [09:13:12] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2084.codfw.wmnet with OS trixie [09:13:26] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12177957 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-be2084.codfw.wmnet with OS trixie completed: - ms-be2084 (**PASS*... [09:15:15] (03Merged) 10jenkins-bot: sre.mysql.depool: regression of T427381 [cookbooks] - 10https://gerrit.wikimedia.org/r/1319509 (https://phabricator.wikimedia.org/T433564) (owner: 10CWilliams) [09:22:37] (03PS3) 10Jcrespo: mariadb: Reenable notifications for db1265 & db1285 [puppet] - 10https://gerrit.wikimedia.org/r/1319809 (https://phabricator.wikimedia.org/T407942) [09:22:37] (03PS1) 10Jcrespo: dbbackups: Add backups for new m1 database: etherpad [puppet] - 10https://gerrit.wikimedia.org/r/1320105 (https://phabricator.wikimedia.org/T433814) [09:23:35] (03CR) 10Jcrespo: "Blocked on +1 and actual db deployment, otherwise backups will generate errors for missing db." [puppet] - 10https://gerrit.wikimedia.org/r/1320105 (https://phabricator.wikimedia.org/T433814) (owner: 10Jcrespo) [09:23:58] (03PS1) 10KartikMistry: Revert "cxserver: Add missing referencePunctuation" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320107 [09:24:41] (03CR) 10Jcrespo: [C:03+2] "Thanks for the review." [puppet] - 10https://gerrit.wikimedia.org/r/1319809 (https://phabricator.wikimedia.org/T407942) (owner: 10Jcrespo) [09:26:27] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host ms-be2085.codfw.wmnet with OS trixie [09:26:34] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-be1087.eqiad.wmnet with OS trixie [09:26:35] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12178005 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host ms-be2085.codfw.wmnet with OS trixie [09:26:42] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12178006 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-be1087.eqiad.wmnet with OS trixie [09:31:50] (03CR) 10KartikMistry: [C:03+2] Revert "cxserver: Add missing referencePunctuation" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320107 (owner: 10KartikMistry) [09:32:20] (03CR) 10Trueg: [C:03+2] WDQSv2: Scheduling based on resource requests [deployment-charts] - 10https://gerrit.wikimedia.org/r/1318733 (https://phabricator.wikimedia.org/T431395) (owner: 10Trueg) [09:34:27] (03Merged) 10jenkins-bot: Revert "cxserver: Add missing referencePunctuation" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320107 (owner: 10KartikMistry) [09:34:57] (03Merged) 10jenkins-bot: WDQSv2: Scheduling based on resource requests [deployment-charts] - 10https://gerrit.wikimedia.org/r/1318733 (https://phabricator.wikimedia.org/T431395) (owner: 10Trueg) [09:40:14] !log mvernon@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage [09:40:21] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage [09:42:25] RESOLVED: SystemdUnitFailed: generate_vrts_aliases.service on mx-in2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:43:56] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage [09:45:25] FIRING: SystemdUnitFailed: generate_vrts_aliases.service on mx-in2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:46:56] (03CR) 10Marostegui: [C:03+1] "Db has been created" [puppet] - 10https://gerrit.wikimedia.org/r/1320105 (https://phabricator.wikimedia.org/T433814) (owner: 10Jcrespo) [09:48:00] FIRING: [3x] HelmReleaseBadStatus: Helm release wdqs/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [09:48:13] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage [09:49:34] (03CR) 10Jcrespo: [C:03+2] dbbackups: Add backups for new m1 database: etherpad [puppet] - 10https://gerrit.wikimedia.org/r/1320105 (https://phabricator.wikimedia.org/T433814) (owner: 10Jcrespo) [09:51:17] (03PS1) 10KartikMistry: cxserver: Add referencePunctuation config [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320110 (https://phabricator.wikimedia.org/T97231) [09:55:35] (03CR) 10KartikMistry: [C:03+2] cxserver: Add referencePunctuation config [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320110 (https://phabricator.wikimedia.org/T97231) (owner: 10KartikMistry) [09:56:09] (03PS1) 10Marostegui: wmnet: Remove x2-master CNAME [dns] - 10https://gerrit.wikimedia.org/r/1320111 (https://phabricator.wikimedia.org/T433685) [09:57:13] (03CR) 10CI reject: [V:04-1] wmnet: Remove x2-master CNAME [dns] - 10https://gerrit.wikimedia.org/r/1320111 (https://phabricator.wikimedia.org/T433685) (owner: 10Marostegui) [09:58:15] (03Merged) 10jenkins-bot: cxserver: Add referencePunctuation config [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320110 (https://phabricator.wikimedia.org/T97231) (owner: 10KartikMistry) [10:00:04] Deploy window MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260803T1000) [10:01:58] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1087.eqiad.wmnet with OS trixie [10:02:06] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12178149 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-be1087.eqiad.wmnet with OS trixie completed: - ms-be1087 (**PASS*... [10:02:30] (03PS1) 10Jelto: cert-manager: import CRDs for v1.19.6 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320114 (https://phabricator.wikimedia.org/T427402) [10:03:25] (03CR) 10Marostegui: "recheck" [dns] - 10https://gerrit.wikimedia.org/r/1320111 (https://phabricator.wikimedia.org/T433685) (owner: 10Marostegui) [10:03:37] (03PS1) 10Ladsgroup: Remove more $wmg = $wg hacks [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320115 (https://phabricator.wikimedia.org/T119117) [10:04:19] (03CR) 10CI reject: [V:04-1] cert-manager: import CRDs for v1.19.6 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320114 (https://phabricator.wikimedia.org/T427402) (owner: 10Jelto) [10:06:40] (03PS1) 10KartikMistry: cxserver: Add referencePunctuation config [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320116 (https://phabricator.wikimedia.org/T97231) [10:12:48] (03CR) 10KartikMistry: [C:03+2] cxserver: Add referencePunctuation config [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320116 (https://phabricator.wikimedia.org/T97231) (owner: 10KartikMistry) [10:15:31] (03CR) 10Clément Goubert: [C:03+2] Update subpage parameters in restbase sandbox redirect [deployment-charts] - 10https://gerrit.wikimedia.org/r/1312695 (owner: 10Aaron Schulz) [10:15:46] (03Merged) 10jenkins-bot: cxserver: Add referencePunctuation config [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320116 (https://phabricator.wikimedia.org/T97231) (owner: 10KartikMistry) [10:18:00] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [10:18:14] (03Merged) 10jenkins-bot: Update subpage parameters in restbase sandbox redirect [deployment-charts] - 10https://gerrit.wikimedia.org/r/1312695 (owner: 10Aaron Schulz) [10:18:34] !log kartik@deploy1003 helmfile [staging] START helmfile.d/services/cxserver: apply [10:18:56] !log kartik@deploy1003 helmfile [staging] DONE helmfile.d/services/cxserver: apply [10:19:36] (03PS2) 10Mpostoronca: WikimediaAntiAbuse: Document required load order after Echo [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319086 (https://phabricator.wikimedia.org/T432452) [10:20:14] !log kartik@deploy1003 helmfile [codfw] START helmfile.d/services/cxserver: apply [10:20:35] 10ops-codfw, 06SRE, 06DBA, 06DC-Ops: db2209 firmware upgrade after a crash - https://phabricator.wikimedia.org/T431952#12178185 (10Marostegui) :-/ Thank you, let me know if I can help further. [10:20:49] !log kartik@deploy1003 helmfile [codfw] DONE helmfile.d/services/cxserver: apply [10:21:17] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2085.codfw.wmnet with OS trixie [10:21:23] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12178189 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-be2085.codfw.wmnet with OS trixie completed: - ms-be2085 (**PASS*... [10:21:40] !log kartik@deploy1003 helmfile [eqiad] START helmfile.d/services/cxserver: apply [10:22:01] !log kartik@deploy1003 helmfile [eqiad] DONE helmfile.d/services/cxserver: apply [10:22:31] !log cgoubert@deploy1003 helmfile [staging] START helmfile.d/services/rest-gateway: apply [10:23:03] !log cgoubert@deploy1003 helmfile [staging] DONE helmfile.d/services/rest-gateway: apply [10:23:12] !log cgoubert@deploy1003 helmfile [staging] START helmfile.d/services/rest-gateway: apply [10:23:22] !log cgoubert@deploy1003 helmfile [staging] DONE helmfile.d/services/rest-gateway: apply [10:24:36] !log cgoubert@deploy1003 helmfile [codfw] START helmfile.d/services/rest-gateway: apply [10:24:51] !log cxserver: Add referencePunctuation config (T97231) [10:24:54] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [10:24:55] T97231: Put references before the dots for some languages - https://phabricator.wikimedia.org/T97231 [10:25:37] (03PS1) 10Marostegui: production-core.sql.erb: Add ms1,ms2 and ms3 [puppet] - 10https://gerrit.wikimedia.org/r/1320117 (https://phabricator.wikimedia.org/T433685) [10:26:25] !log cgoubert@deploy1003 helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply [10:27:00] !log cgoubert@deploy1003 helmfile [eqiad] START helmfile.d/services/rest-gateway: apply [10:27:26] !log cgoubert@deploy1003 helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply [10:31:21] (03CR) 10Marostegui: "This is really a NOOP, this is just for tracking." [puppet] - 10https://gerrit.wikimedia.org/r/1320117 (https://phabricator.wikimedia.org/T433685) (owner: 10Marostegui) [10:31:23] (03CR) 10Marostegui: [C:03+2] production-core.sql.erb: Add ms1,ms2 and ms3 [puppet] - 10https://gerrit.wikimedia.org/r/1320117 (https://phabricator.wikimedia.org/T433685) (owner: 10Marostegui) [10:35:36] !log marostegui@cumin1003 dbctl commit (dc=all): 'Depool db2248 from s4 T433610', diff saved to https://phabricator.wikimedia.org/P95854 and previous config saved to /var/cache/conftool/dbconfig/20260803-103535-marostegui.json [10:35:40] T433610: FIRING: [2x] SystemdUnitFailed: ifup@eno12399np0.service on db2248:9100 - https://phabricator.wikimedia.org/T433610 [10:36:02] (03PS1) 10Clément Goubert: deployment_server: Install make for test suite [puppet] - 10https://gerrit.wikimedia.org/r/1320119 [10:36:16] jouncebot: nowandnext [10:36:16] For the next 0 hour(s) and 23 minute(s): MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260803T1000) [10:36:16] In 2 hour(s) and 23 minute(s): UTC afternoon backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260803T1300) [10:36:53] !log marostegui@cumin1003 dbctl commit (dc=all): 'Depool db2245, db2246 and db2247 T433610', diff saved to https://phabricator.wikimedia.org/P95855 and previous config saved to /var/cache/conftool/dbconfig/20260803-103652-marostegui.json [10:37:55] (03PS2) 10Jelto: cert-manager: import CRDs for v1.19.6 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320114 (https://phabricator.wikimedia.org/T427402) [10:37:55] (03PS1) 10Jelto: cert-manager: update chart for v1.19.6 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320120 (https://phabricator.wikimedia.org/T427402) [10:38:12] 10ops-codfw, 06SRE, 06DC-Ops, 13Patch-For-Review: Q4:rack/setup/install sretest2010 Config J 1P test host - https://phabricator.wikimedia.org/T394357#12178227 (10MatthewVernon) Hi folks, thanks for coming up with some possible solutions :) If @Jhancock.wm has a suitable disk for option 4, that's great (an... [10:39:04] (03CR) 10TrainBranchBot: [C:03+2] "Approved by ladsgroup@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320115 (https://phabricator.wikimedia.org/T119117) (owner: 10Ladsgroup) [10:39:39] (03PS1) 10Cathal Mooney: Reverse PTR Includes: add missing ones for eqsin switch<->cr links [dns] - 10https://gerrit.wikimedia.org/r/1320121 (https://phabricator.wikimedia.org/T418439) [10:40:00] (03CR) 10CI reject: [V:04-1] cert-manager: update chart for v1.19.6 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320120 (https://phabricator.wikimedia.org/T427402) (owner: 10Jelto) [10:40:02] (03CR) 10CI reject: [V:04-1] cert-manager: import CRDs for v1.19.6 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320114 (https://phabricator.wikimedia.org/T427402) (owner: 10Jelto) [10:40:37] (03CR) 10CI reject: [V:04-1] Reverse PTR Includes: add missing ones for eqsin switch<->cr links [dns] - 10https://gerrit.wikimedia.org/r/1320121 (https://phabricator.wikimedia.org/T418439) (owner: 10Cathal Mooney) [10:41:42] !log cmooney@cumin1003 START - Cookbook sre.dns.netbox [10:42:30] (03Merged) 10jenkins-bot: Remove more $wmg = $wg hacks [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320115 (https://phabricator.wikimedia.org/T119117) (owner: 10Ladsgroup) [10:42:46] !log ladsgroup@deploy1003 Started scap sync-world: Backport for [[gerrit:1320115|Remove more $wmg = $wg hacks (T119117)]] [10:42:49] T119117: Get rid of $wg = $wmg hack - https://phabricator.wikimedia.org/T119117 [10:44:59] (03PS1) 10Effie Mouzeli: profile::memcached::instance: migrate to PKI [puppet] - 10https://gerrit.wikimedia.org/r/1320123 (https://phabricator.wikimedia.org/T353511) [10:45:22] (03PS2) 10Effie Mouzeli: profile::memcached::instance: migrate to PKI [puppet] - 10https://gerrit.wikimedia.org/r/1320123 (https://phabricator.wikimedia.org/T353511) [10:45:37] (03CR) 10Effie Mouzeli: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1320123 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [10:45:58] (03PS3) 10Effie Mouzeli: profile::pki::root_ca: add new intermediates for memcached [puppet] - 10https://gerrit.wikimedia.org/r/1320104 (https://phabricator.wikimedia.org/T353511) [10:46:08] (03CR) 10CI reject: [V:04-1] profile::memcached::instance: migrate to PKI [puppet] - 10https://gerrit.wikimedia.org/r/1320123 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [10:46:09] !log cmooney@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" [10:46:20] (03CR) 10Effie Mouzeli: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1320104 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [10:46:24] !log cmooney@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003" [10:46:24] !log cmooney@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [10:46:41] !log ladsgroup@deploy1003 ladsgroup: Backport for [[gerrit:1320115|Remove more $wmg = $wg hacks (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [10:46:44] (03PS2) 10Cathal Mooney: Reverse PTR Includes: add missing ones for eqsin switch<->cr links [dns] - 10https://gerrit.wikimedia.org/r/1320121 (https://phabricator.wikimedia.org/T418439) [10:46:51] (03PS3) 10Effie Mouzeli: profile::memcached::instance: migrate to PKI [puppet] - 10https://gerrit.wikimedia.org/r/1320123 (https://phabricator.wikimedia.org/T353511) [10:47:09] (03PS4) 10Effie Mouzeli: profile::memcached::instance: migrate to PKI [puppet] - 10https://gerrit.wikimedia.org/r/1320123 (https://phabricator.wikimedia.org/T353511) [10:47:10] !log ladsgroup@deploy1003 ladsgroup: Continuing with deployment [10:47:28] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host ms-be2086.codfw.wmnet with OS trixie [10:47:39] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12178264 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host ms-be2086.codfw.wmnet with OS trixie [10:47:42] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-be1088.eqiad.wmnet with OS trixie [10:47:49] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12178265 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-be1088.eqiad.wmnet with OS trixie [10:47:52] (03CR) 10CI reject: [V:04-1] profile::memcached::instance: migrate to PKI [puppet] - 10https://gerrit.wikimedia.org/r/1320123 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [10:48:31] (03CR) 10Ladsgroup: [C:03+2] Use ImageMagick's -background=none flag for Alpha WebP images [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1319567 (https://phabricator.wikimedia.org/T283646) (owner: 10Samwilson) [10:49:09] (03CR) 10Cathal Mooney: [C:03+2] Reverse PTR Includes: add missing ones for eqsin switch<->cr links [dns] - 10https://gerrit.wikimedia.org/r/1320121 (https://phabricator.wikimedia.org/T418439) (owner: 10Cathal Mooney) [10:49:38] !log cmooney@dns3003 START - running authdns-update [10:50:17] I'm gonna deploy some changeprop chart changes, should be minor but we never know [10:51:50] (03CR) 10Clément Goubert: [C:03+2] Remove redundant fetchGoogleCloudVisionAnnotations entry from _jobqueue.yaml [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310654 (owner: 10Aaron Schulz) [10:51:54] !log cmooney@dns3003 END - running authdns-update [10:52:00] (technically noop actually) [10:53:43] !log ladsgroup@deploy1003 Finished scap sync-world: Backport for [[gerrit:1320115|Remove more $wmg = $wg hacks (T119117)]] (duration: 10m 57s) [10:53:48] T119117: Get rid of $wg = $wmg hack - https://phabricator.wikimedia.org/T119117 [10:54:09] (03PS1) 10Slyngshede: IDM: Update IDM to version 0.1.18 [dns] - 10https://gerrit.wikimedia.org/r/1320124 [10:54:33] (03Merged) 10jenkins-bot: Remove redundant fetchGoogleCloudVisionAnnotations entry from _jobqueue.yaml [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310654 (owner: 10Aaron Schulz) [10:55:55] (03CR) 10Marostegui: "recheck" [dns] - 10https://gerrit.wikimedia.org/r/1320111 (https://phabricator.wikimedia.org/T433685) (owner: 10Marostegui) [10:56:19] (03PS3) 10Jelto: cert-manager: import CRDs for v1.19.6 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320114 (https://phabricator.wikimedia.org/T427402) [10:56:29] (03CR) 10Slyngshede: [C:03+1] Use os.environ.get for checking DJANGO_DEBUGGING [software/bitu] - 10https://gerrit.wikimedia.org/r/1319150 (owner: 10Perryprog) [10:57:28] (03CR) 10Slyngshede: [C:03+1] Use os.environ.get for checking DJANGO_DEBUGGING (031 comment) [software/bitu] - 10https://gerrit.wikimedia.org/r/1319150 (owner: 10Perryprog) [10:57:29] (03CR) 10Marostegui: [C:03+2] wmnet: Remove x2-master CNAME [dns] - 10https://gerrit.wikimedia.org/r/1320111 (https://phabricator.wikimedia.org/T433685) (owner: 10Marostegui) [10:57:32] (03CR) 10Slyngshede: [C:03+2] Use os.environ.get for checking DJANGO_DEBUGGING [software/bitu] - 10https://gerrit.wikimedia.org/r/1319150 (owner: 10Perryprog) [10:58:10] (03CR) 10Effie Mouzeli: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1320123 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [10:58:10] (03PS4) 10Jelto: cert-manager: import CRDs for v1.19.6 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320114 (https://phabricator.wikimedia.org/T427402) [10:58:35] (03PS2) 10Jelto: cert-manager: update chart for v1.19.6 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320120 (https://phabricator.wikimedia.org/T427402) [10:58:38] (03CR) 10Effie Mouzeli: "recheck" [puppet] - 10https://gerrit.wikimedia.org/r/1320123 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [10:58:40] (03Merged) 10jenkins-bot: Use ImageMagick's -background=none flag for Alpha WebP images [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1319567 (https://phabricator.wikimedia.org/T283646) (owner: 10Samwilson) [11:00:18] !log marostegui@dns1004 START - running authdns-update [11:00:21] (03CR) 10CI reject: [V:04-1] cert-manager: import CRDs for v1.19.6 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320114 (https://phabricator.wikimedia.org/T427402) (owner: 10Jelto) [11:00:21] (03Merged) 10jenkins-bot: Use os.environ.get for checking DJANGO_DEBUGGING [software/bitu] - 10https://gerrit.wikimedia.org/r/1319150 (owner: 10Perryprog) [11:01:26] !log mvernon@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage [11:02:00] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage [11:02:59] !log marostegui@dns1004 END - running authdns-update [11:04:22] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for iberker - https://phabricator.wikimedia.org/T433258#12178332 (10IBerker-WMF) I signed the L3 document, but I don't really need production server access, I need analytics-privatedata-users access. Does that require server access? [11:04:25] !log cgoubert@deploy1003 helmfile [staging] START helmfile.d/services/changeprop: apply [11:04:32] !log cgoubert@deploy1003 helmfile [staging] DONE helmfile.d/services/changeprop: apply [11:04:38] !log cgoubert@deploy1003 helmfile [codfw] START helmfile.d/services/changeprop: apply [11:05:06] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for iberker - https://phabricator.wikimedia.org/T433258#12178337 (10IBerker-WMF) @SCherukuwada can you approve? Thanks. [11:05:15] !log cgoubert@deploy1003 helmfile [codfw] DONE helmfile.d/services/changeprop: apply [11:05:21] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage [11:06:29] !log cgoubert@deploy1003 helmfile [eqiad] START helmfile.d/services/changeprop: apply [11:06:52] !log cgoubert@deploy1003 helmfile [eqiad] DONE helmfile.d/services/changeprop: apply [11:07:21] !log cgoubert@deploy1003 helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply [11:07:37] !log cgoubert@deploy1003 helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply [11:07:54] !log cgoubert@deploy1003 helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply [11:08:01] !log cgoubert@deploy1003 helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply [11:08:11] !log cgoubert@deploy1003 helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply [11:09:15] !log cgoubert@deploy1003 helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply [11:09:32] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage [11:10:25] !log cgoubert@deploy1003 helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply [11:10:59] !log cgoubert@deploy1003 helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply [11:12:38] FIRING: CertAlmostExpired: gNMI TLS certificate for lsw1-d6-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [11:13:47] (03PS1) 10Marostegui: wmnet: Add ms1 and ms2 [dns] - 10https://gerrit.wikimedia.org/r/1320128 (https://phabricator.wikimedia.org/T433685) [11:16:08] (03CR) 10Arthur taylor: [C:03+2] wikidata-query-gui: bump to latest version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320103 (https://phabricator.wikimedia.org/T433234) (owner: 10Arthur taylor) [11:16:51] (03PS1) 10Slyngshede: Enable CAS Delegated Authentication [software/cas-overlay-template] - 10https://gerrit.wikimedia.org/r/1320129 (https://phabricator.wikimedia.org/T402006) [11:17:36] (03CR) 10Marostegui: [C:03+2] wmnet: Add ms1 and ms2 [dns] - 10https://gerrit.wikimedia.org/r/1320128 (https://phabricator.wikimedia.org/T433685) (owner: 10Marostegui) [11:17:41] !log marostegui@dns1004 START - running authdns-update [11:18:38] (03Merged) 10jenkins-bot: wikidata-query-gui: bump to latest version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320103 (https://phabricator.wikimedia.org/T433234) (owner: 10Arthur taylor) [11:18:47] (03PS5) 10Effie Mouzeli: profile::memcached::instance: migrate to PKI [puppet] - 10https://gerrit.wikimedia.org/r/1320123 (https://phabricator.wikimedia.org/T353511) [11:19:49] (03PS2) 10Slyngshede: Enable CAS Delegated Authentication [software/cas-overlay-template] - 10https://gerrit.wikimedia.org/r/1320129 (https://phabricator.wikimedia.org/T402006) [11:19:53] !log marostegui@dns1004 END - running authdns-update [11:20:41] !log arthurtaylor@deploy1003 helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply [11:21:01] (03CR) 10CI reject: [V:04-1] profile::memcached::instance: migrate to PKI [puppet] - 10https://gerrit.wikimedia.org/r/1320123 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [11:21:03] !log arthurtaylor@deploy1003 helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply [11:22:51] (03PS6) 10Effie Mouzeli: profile::memcached::instance: migrate to PKI [puppet] - 10https://gerrit.wikimedia.org/r/1320123 (https://phabricator.wikimedia.org/T353511) [11:22:52] !log cwilliams@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db[2245-2248].codfw.wmnet with reason: Checking network [11:24:35] (03CR) 10Slyngshede: [V:03+2 C:03+2] Enable CAS Delegated Authentication [software/cas-overlay-template] - 10https://gerrit.wikimedia.org/r/1320129 (https://phabricator.wikimedia.org/T402006) (owner: 10Slyngshede) [11:24:47] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1088.eqiad.wmnet with OS trixie [11:24:53] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12178369 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-be1088.eqiad.wmnet with OS trixie completed: - ms-be1088 (**PASS*... [11:25:36] !log arthurtaylor@deploy1003 helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply [11:25:57] !log arthurtaylor@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply [11:26:04] !log arthurtaylor@deploy1003 helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply [11:26:24] !log arthurtaylor@deploy1003 helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply [11:26:41] (03CR) 10Effie Mouzeli: [C:03+1] "lgtm one nit" [puppet] - 10https://gerrit.wikimedia.org/r/1320119 (owner: 10Clément Goubert) [11:26:52] (03CR) 10Effie Mouzeli: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1320123 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [11:27:18] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2086.codfw.wmnet with OS trixie [11:27:27] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12178372 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-be2086.codfw.wmnet with OS trixie completed: - ms-be2086 (**PASS*... [11:29:03] (03PS4) 10Effie Mouzeli: profile::pki::root_ca: add new intermediates for memcached [puppet] - 10https://gerrit.wikimedia.org/r/1320104 (https://phabricator.wikimedia.org/T353511) [11:30:28] (03PS5) 10Jelto: cert-manager: import CRDs for v1.19.6 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320114 (https://phabricator.wikimedia.org/T427402) [11:30:32] (03CR) 10Effie Mouzeli: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1320104 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [11:30:50] (03PS7) 10Effie Mouzeli: profile::memcached::instance: migrate to PKI [puppet] - 10https://gerrit.wikimedia.org/r/1320123 (https://phabricator.wikimedia.org/T353511) [11:35:43] !log cwilliams@cumin1003 START - Cookbook sre.hosts.remove-downtime for db[2245-2248].codfw.wmnet [11:35:45] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2245-2248].codfw.wmnet [11:38:00] FIRING: [3x] HelmReleaseBadStatus: Helm release wdqs/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [11:39:32] (03PS3) 10Kamila Součková: php: rebuild to include Timo's APCU metrics fix [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1319453 (https://phabricator.wikimedia.org/T433312) [11:41:09] (03CR) 10Kamila Součková: [V:03+1] "checked build locally" [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1319453 (https://phabricator.wikimedia.org/T433312) (owner: 10Kamila Součková) [11:41:38] (03CR) 10Kamila Součková: [V:03+1 C:03+2] "pre-rebase version had a +1, rebase was trivial" [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1319453 (https://phabricator.wikimedia.org/T433312) (owner: 10Kamila Součková) [11:41:42] (03CR) 10Kamila Součková: [V:03+2 C:03+2] php: rebuild to include Timo's APCU metrics fix [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1319453 (https://phabricator.wikimedia.org/T433312) (owner: 10Kamila Součková) [11:49:46] (03PS1) 10Effie Mouzeli: hieradata: switch mc1059 to PKI [puppet] - 10https://gerrit.wikimedia.org/r/1320132 (https://phabricator.wikimedia.org/T353511) [11:49:58] (03CR) 10Effie Mouzeli: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1320132 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [12:00:39] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host ms-be2087.codfw.wmnet with OS trixie [12:00:40] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-be1089.eqiad.wmnet with OS trixie [12:00:48] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12178469 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host ms-be2087.codfw.wmnet with OS trixie [12:00:50] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12178470 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-be1089.eqiad.wmnet with OS trixie [12:00:57] I just missed the infra window, but if nobody else is deploying MW, I'm going to do a deploy to pick up new base images [12:03:38] !log kamila@deploy1003 Started scap sync-world: rebuild after base image update [12:04:03] (03PS2) 10Effie Mouzeli: hieradata: switch mc1059 to PKI [puppet] - 10https://gerrit.wikimedia.org/r/1320132 (https://phabricator.wikimedia.org/T353511) [12:13:35] (03PS1) 10Kevin Bazira: ml-services: update tts-section-generator image to fix misread symbols, abbreviations, and roman numerals [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320135 (https://phabricator.wikimedia.org/T433820) [12:14:12] !log mvernon@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage [12:14:47] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage [12:16:32] (03CR) 10Kevin Bazira: [C:03+2] ml-services: update tts-section-generator image to fix misread symbols, abbreviations, and roman numerals [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320135 (https://phabricator.wikimedia.org/T433820) (owner: 10Kevin Bazira) [12:18:46] jouncebot: nowandnext [12:18:46] No deployments scheduled for the next 0 hour(s) and 41 minute(s) [12:18:46] In 0 hour(s) and 41 minute(s): UTC afternoon backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260803T1300) [12:19:05] I will deploy thumbor stuff. not mediawiki related [12:19:10] (03Merged) 10jenkins-bot: ml-services: update tts-section-generator image to fix misread symbols, abbreviations, and roman numerals [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320135 (https://phabricator.wikimedia.org/T433820) (owner: 10Kevin Bazira) [12:19:25] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage [12:20:45] FIRING: WidespreadPuppetFailure: Puppet has failed in esams - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [12:22:04] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage [12:22:37] (03CR) 10Effie Mouzeli: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1320132 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [12:22:49] !log kevinbazira@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . [12:26:39] (03CR) 10Clément Goubert: deployment_server: Install make for test suite (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1320119 (owner: 10Clément Goubert) [12:27:37] (03CR) 10CI reject: [V:04-1] Localisation updates from https://translatewiki.net. [phabricator/translations] (wmf/stable) - 10https://gerrit.wikimedia.org/r/1320140 (owner: 10L10n-bot) [12:28:16] (03PS2) 10Clément Goubert: deployment_server: Install make for test suite [puppet] - 10https://gerrit.wikimedia.org/r/1320119 [12:28:21] (03CR) 10Clément Goubert: deployment_server: Install make for test suite (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1320119 (owner: 10Clément Goubert) [12:28:26] (03CR) 10Clément Goubert: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1320119 (owner: 10Clément Goubert) [12:28:51] !log kevinbazira@deploy1003 helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . [12:32:56] !log kamila@deploy1003 Finished scap sync-world: rebuild after base image update (duration: 30m 26s) [12:33:24] (03CR) 10Clément Goubert: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1320119 (owner: 10Clément Goubert) [12:34:02] !log cwilliams@cumin1003 START - Cookbook sre.mysql.multiinstance_reboot for db[2245-2247].codfw.wmnet [12:34:29] !log cwilliams@cumin1003 START - Cookbook sre.mysql.depool depool db2245: Rebooting db2245.codfw.wmnet [12:34:38] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2245: Rebooting db2245.codfw.wmnet [12:37:14] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1089.eqiad.wmnet with OS trixie [12:37:21] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12178539 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-be1089.eqiad.wmnet with OS trixie completed: - ms-be1089 (**PASS*... [12:38:45] (03PS1) 10Majavah: hieradata: Add missing v6 addresses for apt.wikimedia.org [puppet] - 10https://gerrit.wikimedia.org/r/1320148 (https://phabricator.wikimedia.org/T433830) [12:39:59] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2087.codfw.wmnet with OS trixie [12:40:07] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12178574 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-be2087.codfw.wmnet with OS trixie completed: - ms-be2087 (**PASS*... [12:42:08] !log cwilliams@cumin1003 START - Cookbook sre.mysql.depool depool db2246: Rebooting db2246.codfw.wmnet [12:42:18] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2246: Rebooting db2246.codfw.wmnet [12:42:42] (03PS1) 10Aude: Enable Special:ChartWizard on Wikimedia Commons [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320149 (https://phabricator.wikimedia.org/T433831) [12:43:59] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, August 03 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployca" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320149 (https://phabricator.wikimedia.org/T433831) (owner: 10Aude) [12:44:40] 10SRE-tools, 06Infrastructure-Foundations, 10Spicerack: Clarify when spicerack/cumin is going to retry - https://phabricator.wikimedia.org/T433698#12178581 (10Mahveotm) Hi @fgiunchedi, I’d like to work on this. I'm propose retaining the delay while also adding a timezone-aware UTC retry timestamp, for exampl... [12:46:03] (03PS1) 10Mahveotm: decorators: include UTC time in retry warnings [software/pywmflib] - 10https://gerrit.wikimedia.org/r/1320151 (https://phabricator.wikimedia.org/T433698) [12:49:12] !log cwilliams@cumin1003 START - Cookbook sre.mysql.depool depool db2247: Rebooting db2247.codfw.wmnet [12:49:22] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2247: Rebooting db2247.codfw.wmnet [12:50:45] RESOLVED: WidespreadPuppetFailure: Puppet has failed in esams - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [12:56:05] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db[2245-2247].codfw.wmnet [12:58:42] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host ms-be2088.codfw.wmnet with OS trixie [12:58:57] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12178641 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host ms-be2088.codfw.wmnet with OS trixie [13:00:02] (03PS1) 10DCausse: opensearch-semantic-search-test: bump to opensearch 3.7.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320153 (https://phabricator.wikimedia.org/T433697) [13:00:04] Lucas_WMDE, urbanecm, and TheresNoTime: May I have your attention please! UTC afternoon backport window. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260803T1300) [13:00:04] HouseOfM and aude: A patch you scheduled for UTC afternoon backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [13:00:20] (03PS3) 10Clément Goubert: deployment_server: Install make for test suite [puppet] - 10https://gerrit.wikimedia.org/r/1320119 [13:00:28] (03CR) 10Clément Goubert: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1320119 (owner: 10Clément Goubert) [13:00:50] hi, i can deploy the config changes [13:01:23] idk if HouseOfM is here? [13:02:25] RESOLVED: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:02:44] o/ [13:03:09] hi, i can deploy yours [13:03:12] (03PS2) 10Mahveotm: decorators: Add UTC timestamp to retry warnings [software/pywmflib] - 10https://gerrit.wikimedia.org/r/1320151 (https://phabricator.wikimedia.org/T433698) [13:03:58] Thank you [13:04:36] (03CR) 10TrainBranchBot: [C:03+2] "Approved by aude@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320149 (https://phabricator.wikimedia.org/T433831) (owner: 10Aude) [13:04:37] (03CR) 10TrainBranchBot: [C:03+2] "Approved by aude@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319854 (https://phabricator.wikimedia.org/T429507) (owner: 10Mhorsey) [13:04:54] (03PS2) 10DCausse: opensearch-semantic-search: bump to opensearch 3.7.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319845 (https://phabricator.wikimedia.org/T433697) [13:05:57] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-be1090.eqiad.wmnet with OS trixie [13:06:05] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12178674 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-be1090.eqiad.wmnet with OS trixie [13:06:33] (03CR) 10Majavah: [V:03+1] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9096/co" [puppet] - 10https://gerrit.wikimedia.org/r/1320148 (https://phabricator.wikimedia.org/T433830) (owner: 10Majavah) [13:07:53] (03Merged) 10jenkins-bot: Enable Special:ChartWizard on Wikimedia Commons [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320149 (https://phabricator.wikimedia.org/T433831) (owner: 10Aude) [13:07:56] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, August 03 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-ite" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319903 (https://phabricator.wikimedia.org/T414338) (owner: 10Krinkle) [13:07:57] (03Merged) 10jenkins-bot: enable CampaignEvents worklists [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319854 (https://phabricator.wikimedia.org/T429507) (owner: 10Mhorsey) [13:08:09] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, August 03 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-ite" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311041 (https://phabricator.wikimedia.org/T271001) (owner: 10Krinkle) [13:08:13] !log aude@deploy1003 Started scap sync-world: Backport for [[gerrit:1320149|Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854|enable CampaignEvents worklists (T429507 T429508)]] [13:08:21] T433831: Enable Special:ChartWizard on Wikimedia Commons - https://phabricator.wikimedia.org/T433831 [13:08:22] T429507: Enable worklists in beta - https://phabricator.wikimedia.org/T429507 [13:08:22] T429508: Enable worklists in production - https://phabricator.wikimedia.org/T429508 [13:09:15] 10SRE-tools, 06Infrastructure-Foundations, 10Spicerack, 13Patch-For-Review: Clarify when spicerack/cumin is going to retry - https://phabricator.wikimedia.org/T433698#12178692 (10Mahveotm) 05Open→03In progress [13:10:57] (03CR) 10Jforrester: "Shouldn't we wait for https://gerrit.wikimedia.org/r/c/mediawiki/extensions/CategoryTree/+/293318 to ship to production first?" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319945 (https://phabricator.wikimedia.org/T152294) (owner: 10TheDJ) [13:12:00] (03CR) 10Effie Mouzeli: [C:03+1] deployment_server: Install make for test suite [puppet] - 10https://gerrit.wikimedia.org/r/1320119 (owner: 10Clément Goubert) [13:12:09] !log aude@deploy1003 aude, mhorsey: Backport for [[gerrit:1320149|Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854|enable CampaignEvents worklists (T429507 T429508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [13:12:47] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage [13:12:51] (03PS4) 10Clément Goubert: deployment_server: Install make for test suite [puppet] - 10https://gerrit.wikimedia.org/r/1320119 [13:13:39] (03PS1) 10Jgiannelos: mobileapps: Define request template for linked artifacts endpoint [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320156 [13:13:50] HouseOfM can you check that your change looks good? [13:13:57] (03PS2) 10Jgiannelos: mobileapps: Define request template for linked artifacts endpoint [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320156 (https://phabricator.wikimedia.org/T431108) [13:15:42] my patch looks good [13:16:00] mine too, thank you! [13:16:21] thanks [13:16:25] !log aude@deploy1003 aude, mhorsey: Continuing with deployment [13:17:11] (03PS3) 10Jgiannelos: mobileapps: Define request template for linked artifacts endpoint [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320156 (https://phabricator.wikimedia.org/T431108) [13:17:53] (03PS1) 10Ebernhardson: cirus: bump to latest image version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320157 (https://phabricator.wikimedia.org/T431784) [13:18:35] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage [13:19:25] !log mvernon@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage [13:22:47] !log aude@deploy1003 Finished scap sync-world: Backport for [[gerrit:1320149|Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854|enable CampaignEvents worklists (T429507 T429508)]] (duration: 14m 34s) [13:22:55] T433831: Enable Special:ChartWizard on Wikimedia Commons - https://phabricator.wikimedia.org/T433831 [13:22:55] T429507: Enable worklists in beta - https://phabricator.wikimedia.org/T429507 [13:22:56] T429508: Enable worklists in production - https://phabricator.wikimedia.org/T429508 [13:22:58] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage [13:23:36] (03PS17) 10Arnaudb: trafficserver: add a map for gitlab instances as a backend [puppet] - 10https://gerrit.wikimedia.org/r/1290731 (https://phabricator.wikimedia.org/T425441) [13:26:44] (03CR) 10Arnaudb: "the question was about the timings, I provided a more detailed description in the latest change to explain their values a bit better after" [puppet] - 10https://gerrit.wikimedia.org/r/1290731 (https://phabricator.wikimedia.org/T425441) (owner: 10Arnaudb) [13:27:49] (03CR) 10Clément Goubert: [C:03+1] mobileapps: Define request template for linked artifacts endpoint [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320156 (https://phabricator.wikimedia.org/T431108) (owner: 10Jgiannelos) [13:28:29] (03PS31) 10Slyngshede: P:tofurkey Add tofurkey [puppet] - 10https://gerrit.wikimedia.org/r/1260730 (https://phabricator.wikimedia.org/T355446) [13:32:34] (03CR) 10Jgiannelos: [C:03+2] mobileapps: Define request template for linked artifacts endpoint [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320156 (https://phabricator.wikimedia.org/T431108) (owner: 10Jgiannelos) [13:34:05] (03PS1) 10Jforrester: TimedText: Include SRT-only sources when listing playback tracks [extensions/TimedMediaHandler] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1320159 (https://phabricator.wikimedia.org/T433666) [13:35:23] (03Merged) 10jenkins-bot: mobileapps: Define request template for linked artifacts endpoint [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320156 (https://phabricator.wikimedia.org/T431108) (owner: 10Jgiannelos) [13:35:48] (03PS1) 10Majavah: apereo_cas: Pin pac4j session-replication key sizes [puppet] - 10https://gerrit.wikimedia.org/r/1320161 (https://phabricator.wikimedia.org/T402006) [13:36:57] (03PS1) 10Andrew Bogott: wmcs_backup_instances.yaml.erb: move a few projects back to cloudbackup1003 [puppet] - 10https://gerrit.wikimedia.org/r/1320162 (https://phabricator.wikimedia.org/T430018) [13:38:33] (03PS1) 10Milazg: Add configurable RestTermsOfServiceUrl [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320163 (https://phabricator.wikimedia.org/T428147) [13:39:31] (03CR) 10CI reject: [V:04-1] Add configurable RestTermsOfServiceUrl [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320163 (https://phabricator.wikimedia.org/T428147) (owner: 10Milazg) [13:40:12] (03CR) 10Andrew Bogott: [C:03+2] wmcs_backup_instances.yaml.erb: move a few projects back to cloudbackup1003 [puppet] - 10https://gerrit.wikimedia.org/r/1320162 (https://phabricator.wikimedia.org/T430018) (owner: 10Andrew Bogott) [13:40:36] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1090.eqiad.wmnet with OS trixie [13:40:43] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12178799 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-be1090.eqiad.wmnet with OS trixie completed: - ms-be1090 (**PASS*... [13:40:45] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2088.codfw.wmnet with OS trixie [13:40:53] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12178800 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-be2088.codfw.wmnet with OS trixie completed: - ms-be2088 (**PASS*... [13:43:49] (03PS1) 10Slyngshede: P:tofurkey: Tofurkey dummy secret [labs/private] - 10https://gerrit.wikimedia.org/r/1320164 (https://phabricator.wikimedia.org/T355446) [13:44:49] (03CR) 10Slyngshede: [V:03+2 C:03+2] P:tofurkey: Tofurkey dummy secret [labs/private] - 10https://gerrit.wikimedia.org/r/1320164 (https://phabricator.wikimedia.org/T355446) (owner: 10Slyngshede) [13:45:25] FIRING: SystemdUnitFailed: generate_vrts_aliases.service on mx-in2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:47:14] (03PS32) 10Slyngshede: P:tofurkey Add tofurkey [puppet] - 10https://gerrit.wikimedia.org/r/1260730 (https://phabricator.wikimedia.org/T355446) [13:48:03] (03CR) 10Jelto: "I created the `crds.yaml` by running the following command in the `v1.19.6` branch:" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320114 (https://phabricator.wikimedia.org/T427402) (owner: 10Jelto) [13:48:53] (03CR) 10Aghirelli: "recheck" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320163 (https://phabricator.wikimedia.org/T428147) (owner: 10Milazg) [13:49:14] (03PS6) 10Andrew Bogott: Ceph: support setting bdev_enable_discard [puppet] - 10https://gerrit.wikimedia.org/r/1318163 (https://phabricator.wikimedia.org/T429387) [13:49:15] (03PS13) 10Andrew Bogott: ceph.conf: support adjusting slow ops health messages [puppet] - 10https://gerrit.wikimedia.org/r/1313966 (https://phabricator.wikimedia.org/T429387) [13:49:45] (03CR) 10Slyngshede: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1260730 (https://phabricator.wikimedia.org/T355446) (owner: 10Slyngshede) [13:50:32] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host ms-be2089.codfw.wmnet with OS trixie [13:50:41] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12178823 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host ms-be2089.codfw.wmnet with OS trixie [13:53:39] (03PS2) 10Milazg: Add configurable RestTermsOfServiceUrl [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320163 (https://phabricator.wikimedia.org/T428147) [13:53:57] FIRING: [2x] HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-rollback - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [13:54:31] (03CR) 10Slyngshede: P:tofurkey Add tofurkey (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1260730 (https://phabricator.wikimedia.org/T355446) (owner: 10Slyngshede) [14:00:09] (03PS1) 10Kamila Součková: CI: fix comment for _helmfile_build [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320166 (https://phabricator.wikimedia.org/T388390) [14:06:44] (03CR) 10Ladsgroup: [C:03+2] thumbor: Restrict access to thumb file types [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319481 (https://phabricator.wikimedia.org/T430528) (owner: 10Ladsgroup) [14:08:00] FIRING: [2x] HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-rollback - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [14:09:02] (03CR) 10Kamila Součková: [C:03+1] deployment_server: Install make for test suite [puppet] - 10https://gerrit.wikimedia.org/r/1320119 (owner: 10Clément Goubert) [14:09:40] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage [14:09:51] (03Merged) 10jenkins-bot: thumbor: Restrict access to thumb file types [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319481 (https://phabricator.wikimedia.org/T430528) (owner: 10Ladsgroup) [14:10:43] !log ladsgroup@deploy1003 helmfile [staging] START helmfile.d/services/thumbor: apply [14:10:56] !log ladsgroup@deploy1003 helmfile [staging] DONE helmfile.d/services/thumbor: apply [14:12:32] (03CR) 10Bking: [C:03+2] cirrussearch: Replace icinga check with prometheus blackbox check [puppet] - 10https://gerrit.wikimedia.org/r/1318301 (https://phabricator.wikimedia.org/T384998) (owner: 10Bking) [14:14:33] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage [14:16:01] !log ladsgroup@deploy1003 helmfile [codfw] START helmfile.d/services/thumbor: apply [14:16:33] (03PS1) 10Jforrester: Thumbnail: Exclude map figures from multimediaviewer [extensions/Kartographer] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1320171 (https://phabricator.wikimedia.org/T433703) [14:18:00] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [14:18:27] !log ladsgroup@deploy1003 helmfile [codfw] DONE helmfile.d/services/thumbor: apply [14:22:20] 06SRE, 06Infrastructure-Foundations, 10vm-requests: eqiad/codfw: 2 VM request for codesearch - https://phabricator.wikimedia.org/T433316#12178938 (10LSobanski) Approved. [14:22:24] (03CR) 10Slyngshede: [C:03+1] apereo_cas: Pin pac4j session-replication key sizes [puppet] - 10https://gerrit.wikimedia.org/r/1320161 (https://phabricator.wikimedia.org/T402006) (owner: 10Majavah) [14:24:02] (03CR) 10Fabfur: [C:03+1] "lgtm!" [puppet] - 10https://gerrit.wikimedia.org/r/1260730 (https://phabricator.wikimedia.org/T355446) (owner: 10Slyngshede) [14:27:26] !log ladsgroup@deploy1003 helmfile [eqiad] START helmfile.d/services/thumbor: apply [14:28:58] !help [14:28:58] want docs? ask for "!wm-bot". all keywords? try "@regsearch .*" [14:28:58] You're not allowed to perform this action. [14:29:08] !wm-bot [14:29:08] http://meta.wikimedia.org/wiki/WM-Bot [14:29:48] !log ladsgroup@deploy1003 helmfile [eqiad] DONE helmfile.d/services/thumbor: apply [14:30:05] Deploy window Test Kitchen Experiment Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260803T1430) [14:30:47] Results (Found 80): puppet, morebots, bang, nagios, bot, labs-home-wm, labs-nagios-wm, labs-morebots, gerrit-wm, wiki, bastion, extension, wm-bot, putty, gerrit, wikitech, revision, monitor, alert, password, unicorn, help, bz, os-change, leslie's-reset, damianz's-reset, credentials, queue, info, logging, ask, $realm, keys, $site, bug, pageant, blueprint-dns, pxe, ghsh, rt, erb, regsubst, bots, wt, gerrit-search, change, dn, opshelp, testwiki, sal, task, thx, depool, cluster, hiera, infobot, zuul, jouncebot, jenkins, mira, selfie, hug, instance, git, amend, security, labs, projects, instancelist, instance-del, pathconflict, terminology, access, socks-proxy, sudo, ping, tzag, emoji, rain_dance, kuai, [14:30:47] @regsearch .* [14:33:25] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2089.codfw.wmnet with OS trixie [14:33:34] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12179023 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-be2089.codfw.wmnet with OS trixie completed: - ms-be2089 (**PASS*... [14:36:45] (03PS1) 10Slyngshede: Tofurkey: switch to a secrets file [labs/private] - 10https://gerrit.wikimedia.org/r/1320172 [14:37:48] (03CR) 10Scott French: [C:03+2] "Thanks, Ahmon!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319890 (https://phabricator.wikimedia.org/T412941) (owner: 10Ahmon Dancy) [14:37:55] (03PS33) 10Slyngshede: P:tofurkey Add tofurkey [puppet] - 10https://gerrit.wikimedia.org/r/1260730 (https://phabricator.wikimedia.org/T355446) [14:40:09] (03CR) 10CI reject: [V:04-1] P:tofurkey Add tofurkey [puppet] - 10https://gerrit.wikimedia.org/r/1260730 (https://phabricator.wikimedia.org/T355446) (owner: 10Slyngshede) [14:40:35] (03Merged) 10jenkins-bot: services/shellbox: Add traindev environment [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319890 (https://phabricator.wikimedia.org/T412941) (owner: 10Ahmon Dancy) [14:42:02] (03CR) 10Effie Mouzeli: [C:03+1] api-gateway: remove some API Gateway specific template files #1 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319114 (https://phabricator.wikimedia.org/T428625) (owner: 10Aaron Schulz) [14:42:17] (03CR) 10Clément Goubert: [C:03+1] "Diff shows no-op for rest-gateway deployments in production." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319114 (https://phabricator.wikimedia.org/T428625) (owner: 10Aaron Schulz) [14:47:15] jouncebot: now [14:47:15] For the next 0 hour(s) and 12 minute(s): Test Kitchen Experiment Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260803T1430) [14:47:21] jouncebot: next [14:47:21] In 0 hour(s) and 42 minute(s): Wikimedia Portals Update (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260803T1530) [14:47:44] (03CR) 10Effie Mouzeli: [C:03+2] "tx claime!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319114 (https://phabricator.wikimedia.org/T428625) (owner: 10Aaron Schulz) [14:49:46] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-be1091.eqiad.wmnet with OS trixie [14:49:56] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12179135 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-be1091.eqiad.wmnet with OS trixie [14:50:07] (03Merged) 10jenkins-bot: api-gateway: remove some API Gateway specific template files #1 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319114 (https://phabricator.wikimedia.org/T428625) (owner: 10Aaron Schulz) [14:50:18] (03CR) 10Majavah: [V:03+1 C:04-1] "bah, doesn't look like this is quite feasible at this point, per PCC" [puppet] - 10https://gerrit.wikimedia.org/r/1320148 (https://phabricator.wikimedia.org/T433830) (owner: 10Majavah) [14:50:21] (03CR) 10Majavah: [V:03+1 C:04-2] hieradata: Add missing v6 addresses for apt.wikimedia.org [puppet] - 10https://gerrit.wikimedia.org/r/1320148 (https://phabricator.wikimedia.org/T433830) (owner: 10Majavah) [14:50:39] (03CR) 10Majavah: [C:03+2] apereo_cas: Pin pac4j session-replication key sizes [puppet] - 10https://gerrit.wikimedia.org/r/1320161 (https://phabricator.wikimedia.org/T402006) (owner: 10Majavah) [14:52:00] (03PS8) 10Effie Mouzeli: profile::memcached::instance: migrate to PKI [puppet] - 10https://gerrit.wikimedia.org/r/1320123 (https://phabricator.wikimedia.org/T353511) [14:56:15] (03CR) 10Ahmon Dancy: [C:03+1] logging: Remove mention of 'fatal' channel that no longer exists [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319868 (https://phabricator.wikimedia.org/T247113) (owner: 10Krinkle) [14:57:22] (03CR) 10Dpogorzelski: [C:03+1] cirrus: Add a defined type for master-eligibles [puppet] - 10https://gerrit.wikimedia.org/r/1319891 (https://phabricator.wikimedia.org/T433726) (owner: 10Bking) [15:00:56] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12179213 (10Andrew) I put 1048 back in service and it seems to be working fine. [15:03:04] !log mvernon@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage [15:07:28] !log dancy@deploy1003 Installing scap version "4.277.0" for 3 host(s) [15:08:17] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage [15:08:27] (03PS2) 10Majavah: hieradata: Add missing v6 addresses for apt.wikimedia.org [puppet] - 10https://gerrit.wikimedia.org/r/1320148 (https://phabricator.wikimedia.org/T433830) [15:08:28] (03PS1) 10Majavah: P:dns::auth::discovery: Support multiple IPs per service and DC [puppet] - 10https://gerrit.wikimedia.org/r/1320177 (https://phabricator.wikimedia.org/T433830) [15:09:20] (03CR) 10CI reject: [V:04-1] P:dns::auth::discovery: Support multiple IPs per service and DC [puppet] - 10https://gerrit.wikimedia.org/r/1320177 (https://phabricator.wikimedia.org/T433830) (owner: 10Majavah) [15:09:25] !log dancy@deploy1003 Installation of scap version "4.277.0" completed for 3 hosts [15:09:50] !log dancy@deploy1003 Started scap sync-world: testing [15:10:22] 06SRE, 06collaboration-services, 06Release-Engineering-Team (Seen): URL shortener subdomains for useful Wikimedia infrastructure - https://phabricator.wikimedia.org/T223319#12179266 (10LSobanski) 05Open→03Declined I don't see Collab working on this any time soon. [15:11:27] (03PS2) 10Majavah: P:dns::auth::discovery: Support multiple IPs per service and DC [puppet] - 10https://gerrit.wikimedia.org/r/1320177 (https://phabricator.wikimedia.org/T433830) [15:11:27] (03PS3) 10Majavah: hieradata: Add missing v6 addresses for apt.wikimedia.org [puppet] - 10https://gerrit.wikimedia.org/r/1320148 (https://phabricator.wikimedia.org/T433830) [15:12:13] !log marostegui@cumin1003 dbctl commit (dc=all): 'Repool db2245, db2246, db2247 and db2248 T433610', diff saved to https://phabricator.wikimedia.org/P95857 and previous config saved to /var/cache/conftool/dbconfig/20260803-151212-marostegui.json [15:12:18] T433610: FIRING: [2x] SystemdUnitFailed: ifup@eno12399np0.service on db2248:9100 - https://phabricator.wikimedia.org/T433610 [15:12:22] (03CR) 10Majavah: [V:03+1] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9098/co" [puppet] - 10https://gerrit.wikimedia.org/r/1320177 (https://phabricator.wikimedia.org/T433830) (owner: 10Majavah) [15:12:38] FIRING: CertAlmostExpired: gNMI TLS certificate for lsw1-d6-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [15:13:03] (03PS3) 10Milazg: Add configurable RestTermsOfServiceUrl [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320163 (https://phabricator.wikimedia.org/T428147) [15:14:03] (03CR) 10Majavah: [V:03+1 C:04-2] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9099/co" [puppet] - 10https://gerrit.wikimedia.org/r/1320148 (https://phabricator.wikimedia.org/T433830) (owner: 10Majavah) [15:14:42] (03CR) 10Majavah: [V:03+1] hieradata: Add missing v6 addresses for apt.wikimedia.org [puppet] - 10https://gerrit.wikimedia.org/r/1320148 (https://phabricator.wikimedia.org/T433830) (owner: 10Majavah) [15:18:29] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host ms-be2090.codfw.wmnet with OS trixie [15:18:37] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12179347 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host ms-be2090.codfw.wmnet with OS trixie [15:21:23] (03CR) 10David Caro: [C:03+1] "> 1) Is this really how to set this? Online docs barely acknowledge ceph.conf, preferring instead live config with "ceph config"" [puppet] - 10https://gerrit.wikimedia.org/r/1313966 (https://phabricator.wikimedia.org/T429387) (owner: 10Andrew Bogott) [15:23:40] !log dancy@deploy1003 Installing scap version "4.276.1" for 3 host(s) [15:24:33] (03PS1) 10Arlolra: Turn on PRV for all namespaces on enwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320179 (https://phabricator.wikimedia.org/T430194) [15:25:26] !log dancy@deploy1003 Installation of scap version "4.276.1" completed for 3 hosts [15:26:21] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1091.eqiad.wmnet with OS trixie [15:26:29] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12179388 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-be1091.eqiad.wmnet with OS trixie completed: - ms-be1091 (**PASS*... [15:26:40] (03CR) 10BCornwall: [C:03+1] IDM: Update IDM to version 0.1.18 [dns] - 10https://gerrit.wikimedia.org/r/1320124 (owner: 10Slyngshede) [15:30:04] jan_drewniak: That opportune time for a Wikimedia Portals Update deploy is upon us again. Don't be afraid. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260803T1530). [15:30:36] (03CR) 10BCornwall: [C:03+1] lists.wikimedia.org: dmarc, set policy to reject [dns] - 10https://gerrit.wikimedia.org/r/1319154 (https://phabricator.wikimedia.org/T379517) (owner: 10JHathaway) [15:30:59] (03CR) 10JHathaway: [C:03+2] lists.wikimedia.org: dmarc, set policy to reject [dns] - 10https://gerrit.wikimedia.org/r/1319154 (https://phabricator.wikimedia.org/T379517) (owner: 10JHathaway) [15:31:17] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, August 03 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-ite" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319868 (https://phabricator.wikimedia.org/T247113) (owner: 10Krinkle) [15:31:47] !log jhathaway@dns1004 START - running authdns-update [15:33:37] !log jhathaway@dns1004 END - running authdns-update [15:33:49] (03CR) 10Aghirelli: [C:03+1] Remove boilerplate language from wmf-rest and wmf-math API modules [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319906 (https://phabricator.wikimedia.org/T433736) (owner: 10ArielGlenn) [15:34:27] 06SRE, 06collaboration-services, 10Wikimedia-Mailing-lists, 13Patch-For-Review: Further improve DMARC compatibility on lists.wikimedia.org - https://phabricator.wikimedia.org/T379517#12179446 (10jhathaway) I went ahead and merged the patch, without the config enforcement in place. @LSobanski and I agreed t... [15:36:57] (03PS1) 10Andrew Bogott: cloudnfs: replace wikiqlever with project uuid [puppet] - 10https://gerrit.wikimedia.org/r/1320181 [15:39:38] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage [15:43:08] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage [15:47:00] (03CR) 10BCornwall: "Thanks for the detailed explanation! It seems like `systemd-standalone-sysusers` would the intended route in a Debian system. Perhaps we n" [debs/pint] - 10https://gerrit.wikimedia.org/r/1319039 (owner: 10Hnowlan) [15:48:11] 10ops-eqiad, 06SRE, 06collaboration-services, 06DC-Ops: eqiad row A&B host migration details request for Collaboration Services - https://phabricator.wikimedia.org/T432642#12179516 (10LSobanski) @RobH what's the next expected step with this task? [15:49:58] (03PS1) 10Majavah: cloudnfs: Resolve project names to IDs when looking up mounts [puppet] - 10https://gerrit.wikimedia.org/r/1320183 [15:50:55] !log jiji@deploy1003 helmfile [staging] START helmfile.d/services/rest-gateway: apply [15:50:55] (03CR) 10BCornwall: [C:04-1] "Rather, instead of abandon, modify to depend on `systemd-sysusers`, which either `systemd-sysusers` or `systemd-standalone-sysusers` provi" [debs/pint] - 10https://gerrit.wikimedia.org/r/1319039 (owner: 10Hnowlan) [15:51:06] !log jiji@deploy1003 helmfile [staging] DONE helmfile.d/services/rest-gateway: apply [15:51:18] !log jiji@deploy1003 helmfile [codfw] START helmfile.d/services/rest-gateway: apply [15:51:40] !log jiji@deploy1003 helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply [15:52:10] FIRING: BFDdown: BFD session down between cr2-esams and fe80::ee38:7300:17e8:9c56 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-esams:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [15:57:10] RESOLVED: BFDdown: BFD session down between cr2-esams and fe80::ee38:7300:17e8:9c56 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-esams:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [16:00:20] 10ops-eqiad, 06SRE, 06collaboration-services, 06DC-Ops: eqiad row A&B host migration details request for Collaboration Services - https://phabricator.wikimedia.org/T432642#12179593 (10RobH) >>! In T432642#12179516, @LSobanski wrote: > @RobH what's the next expected step with this task? **Projected Target... [16:02:19] 06SRE, 10SRE-swift-storage, 05Goal, 06Machine-Learning-Team (Q1 FY2026-27), 07OKR-Work: Provision and wire object storage for TTS v1 audio artifacts: S3 sink, Swift bucket, and credentials - https://phabricator.wikimedia.org/T432944#12179609 (10kevinbazira) @MatthewVernon following up on the storage buck... [16:02:50] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2090.codfw.wmnet with OS trixie [16:02:59] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12179612 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-be2090.codfw.wmnet with OS trixie completed: - ms-be2090 (**PASS*... [16:03:07] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-be1092.eqiad.wmnet with OS trixie [16:03:18] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12179614 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-be1092.eqiad.wmnet with OS trixie [16:04:47] (03PS1) 10Perryprog: Use os.environ.get throughout docker_settings.py [software/bitu] - 10https://gerrit.wikimedia.org/r/1320185 [16:05:21] 06SRE, 06Infrastructure-Foundations, 10vm-requests: eqiad/codfw: 2 VM request for codesearch - https://phabricator.wikimedia.org/T433316#12179627 (10Dzahn) a:03Dzahn [16:07:23] FIRING: [2x] CertAlmostExpired: gNMI TLS certificate for lsw1-d6-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [16:08:05] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:08:20] (03PS2) 10Perryprog: Use os.environ.get throughout docker_settings.py [software/bitu] - 10https://gerrit.wikimedia.org/r/1320185 [16:11:14] (03CR) 10David Caro: [C:03+1] "LGTM, make sure to test though" [puppet] - 10https://gerrit.wikimedia.org/r/1320183 (owner: 10Majavah) [16:12:26] (03CR) 10Andrew Bogott: [C:03+1] "somewhat surprising but should work!" [puppet] - 10https://gerrit.wikimedia.org/r/1320183 (owner: 10Majavah) [16:14:11] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host ms-be2091.codfw.wmnet with OS trixie [16:14:20] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12179705 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host ms-be2091.codfw.wmnet with OS trixie [16:16:42] (03PS2) 10Majavah: cloudnfs: Resolve project names to IDs when looking up mounts [puppet] - 10https://gerrit.wikimedia.org/r/1320183 [16:17:17] (03CR) 10CI reject: [V:04-1] cloudnfs: Resolve project names to IDs when looking up mounts [puppet] - 10https://gerrit.wikimedia.org/r/1320183 (owner: 10Majavah) [16:18:00] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:18:22] (03PS3) 10Majavah: cloudnfs: Resolve project names to IDs when looking up mounts [puppet] - 10https://gerrit.wikimedia.org/r/1320183 [16:19:19] (03CR) 10Majavah: [C:03+2] cloudnfs: Resolve project names to IDs when looking up mounts [puppet] - 10https://gerrit.wikimedia.org/r/1320183 (owner: 10Majavah) [16:22:03] (03PS3) 10Perryprog: Use os.environ.get throughout docker_settings.py [software/bitu] - 10https://gerrit.wikimedia.org/r/1320185 [16:23:31] !log mvernon@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage [16:24:30] !log jiji@deploy1003 helmfile [eqiad] START helmfile.d/services/rest-gateway: apply [16:24:54] !log jiji@deploy1003 helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply [16:27:10] (03CR) 10Ebernhardson: [C:03+1] opensearch-semantic-search-test: bump to opensearch 3.7.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320153 (https://phabricator.wikimedia.org/T433697) (owner: 10DCausse) [16:27:47] (03CR) 10Ebernhardson: [C:03+2] cirus: bump to latest image version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320157 (https://phabricator.wikimedia.org/T431784) (owner: 10Ebernhardson) [16:27:57] (03Abandoned) 10Effie Mouzeli: profile::pki::root_ca: add new intermediates for memcached [puppet] - 10https://gerrit.wikimedia.org/r/1320104 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [16:28:06] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage [16:28:12] (03PS9) 10Effie Mouzeli: profile::memcached::instance: migrate to PKI [puppet] - 10https://gerrit.wikimedia.org/r/1320123 (https://phabricator.wikimedia.org/T353511) [16:30:17] (03Merged) 10jenkins-bot: cirus: bump to latest image version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320157 (https://phabricator.wikimedia.org/T431784) (owner: 10Ebernhardson) [16:32:03] 10ops-codfw, 06SRE, 06DC-Ops, 10decommission-hardware: Decommission backup2004, backup2005, backup2006 & backup2007 - https://phabricator.wikimedia.org/T431098#12179833 (10Jhancock.wm) 05Open→03Resolved a:05jcrespo→03Jhancock.wm [16:33:11] 06SRE, 10SRE-Access-Requests, 06Release-Engineering-Team: Requesting access to Jenkins server for PWangai-WMF - https://phabricator.wikimedia.org/T433563#12179849 (10brennen) Noting here that I've added this to the RelEng team discussion agenda for Weds this week; we'll get back to you. [16:34:00] 10ops-codfw, 06DC-Ops, 06ServiceOps, 10ServiceOps-Upgrades-Hardware: move wikikube-worker servers in same rack - https://phabricator.wikimedia.org/T431585#12179856 (10Jhancock.wm) @Clement_Goubert I should just need to do an update in netbox. since it won't be leaving the rack, no reimage needed. I'll work... [16:35:17] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage [16:36:00] PROBLEM - PyBal backends health check on lvs2014 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs2015.codfw.wmnet, wdqs2011.codfw.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [16:36:02] PROBLEM - PyBal backends health check on lvs2013 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs2008.codfw.wmnet, wdqs2011.codfw.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [16:37:34] 06SRE, 10SRE-Access-Requests, 06Release-Engineering-Team: Requesting access to Jenkins server for PWangai-WMF - https://phabricator.wikimedia.org/T433563#12179888 (10AMarkossyan-WMF) Thank you, Brennen! On my end, I know that this requires my approval, so here you go: 👍 In general, as this is blocking us... [16:37:45] !log jhancock@cumin2002 START - Cookbook sre.dns.netbox [16:38:01] RECOVERY - PyBal backends health check on lvs2014 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [16:38:01] (03PS1) 10Jforrester: ExtensionDistributor: Drop REL1_44, EOL [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320196 (https://phabricator.wikimedia.org/T428911) [16:38:02] RECOVERY - PyBal backends health check on lvs2013 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [16:38:40] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage [16:40:56] !log jhancock@cumin2002 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [16:41:15] !log jhancock@cumin2002 START - Cookbook sre.network.configure-switch-interfaces for host mc2046 [16:41:24] !log jhancock@cumin2002 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 [16:43:12] !log dzahn@cumin2002 START - Cookbook sre.ganeti.makevm for new host codesearch1001.eqiad.wmnet [16:43:14] !log dzahn@cumin2002 START - Cookbook sre.dns.netbox [16:43:16] !log ebernhardson@deploy1003 helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply [16:43:20] !log ebernhardson@deploy1003 helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply [16:43:28] (03CR) 10Aghirelli: [C:04-1] "CommonSettings.php:647 already sets `$wgRestTermsOfServiceUrl` globally for every wiki — same pattern as `$wgEmergencyContact` right above" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320163 (https://phabricator.wikimedia.org/T428147) (owner: 10Milazg) [16:43:52] 10ops-codfw, 06SRE, 06DBA, 06DC-Ops: db2209 firmware upgrade after a crash - https://phabricator.wikimedia.org/T431952#12179958 (10Jhancock.wm) i gave the dell tech more info but still waiting for the results. will keep bugging them this week. [16:46:10] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1092.eqiad.wmnet with OS trixie [16:46:17] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12179981 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-be1092.eqiad.wmnet with OS trixie completed: - ms-be1092 (**PASS*... [16:48:06] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, August 03 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-ite" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320179 (https://phabricator.wikimedia.org/T430194) (owner: 10Arlolra) [16:49:05] !log ebernhardson@deploy1003 helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply [16:49:12] !log ebernhardson@deploy1003 helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply [16:49:21] dzahn@cumin2002 makevm (PID 3715560) is awaiting input [16:49:56] (03CR) 10BCornwall: [C:03+1] P:puppetserver::volatile add geomap_dump dir [puppet] - 10https://gerrit.wikimedia.org/r/1307217 (https://phabricator.wikimedia.org/T402512) (owner: 10CDobbins) [16:50:15] (03PS10) 10Effie Mouzeli: profile::memcached::instance: migrate to PKI [puppet] - 10https://gerrit.wikimedia.org/r/1320123 (https://phabricator.wikimedia.org/T353511) [16:50:37] (03PS3) 10Effie Mouzeli: hieradata: switch mc1059 to PKI [puppet] - 10https://gerrit.wikimedia.org/r/1320132 (https://phabricator.wikimedia.org/T353511) [16:50:53] (03CR) 10CI reject: [V:04-1] profile::memcached::instance: migrate to PKI [puppet] - 10https://gerrit.wikimedia.org/r/1320123 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [16:51:29] (03CR) 10JHathaway: [C:03+1] hieradata: switch mc1059 to PKI [puppet] - 10https://gerrit.wikimedia.org/r/1320132 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [16:52:02] (03PS11) 10Effie Mouzeli: profile::memcached::instance: migrate to PKI [puppet] - 10https://gerrit.wikimedia.org/r/1320123 (https://phabricator.wikimedia.org/T353511) [16:52:28] (03PS4) 10Effie Mouzeli: hieradata: switch mc1059 to PKI [puppet] - 10https://gerrit.wikimedia.org/r/1320132 (https://phabricator.wikimedia.org/T353511) [16:52:31] (03CR) 10Effie Mouzeli: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1320132 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [16:53:26] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12180033 (10MatthewVernon) [16:54:28] !log ebernhardson@deploy1003 helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply [16:54:32] !log ebernhardson@deploy1003 helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply [16:54:35] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12180035 (10MatthewVernon) [16:56:01] (03CR) 10Effie Mouzeli: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1320132 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [16:57:01] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2091.codfw.wmnet with OS trixie [16:57:05] (03CR) 10Effie Mouzeli: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1320123 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [16:57:08] 06SRE, 10SRE-Access-Requests, 06Release-Engineering-Team: Requesting access to Jenkins server for PWangai-WMF - https://phabricator.wikimedia.org/T433563#12180047 (10Dzahn) I strongly recommend to define the proper solution right now and avoid a diverting "path of least resistence". (Also there is no "just a... [16:57:09] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12180048 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-be2091.codfw.wmnet with OS trixie completed: - ms-be2091 (**PASS*... [16:58:55] !log dzahn@cumin2002 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" [16:59:32] (03CR) 10CDobbins: [C:03+2] P:puppetserver::volatile add geomap_dump dir [puppet] - 10https://gerrit.wikimedia.org/r/1307217 (https://phabricator.wikimedia.org/T402512) (owner: 10CDobbins) [17:00:04] Deploy window MediaWiki infrastructure (UTC late) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260803T1700) [17:00:04] ryankemper: How many deployers does it take to do Wikidata Query Service weekly deploy deploy? (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260803T1700). [17:00:08] o/ [17:00:16] deploying changeprop in the MW infra window [17:00:26] (03CR) 10RLazarus: [C:03+2] Add CacheAbstractContentFragment to high_traffic_jobs with concurrency=1 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1314852 (https://phabricator.wikimedia.org/T430898) (owner: 10Genoveva Galarza) [17:01:43] (03PS12) 10Effie Mouzeli: profile::memcached::instance: migrate to PKI [puppet] - 10https://gerrit.wikimedia.org/r/1320123 (https://phabricator.wikimedia.org/T353511) [17:01:55] (03PS5) 10Effie Mouzeli: hieradata: switch mc1059 to PKI [puppet] - 10https://gerrit.wikimedia.org/r/1320132 (https://phabricator.wikimedia.org/T353511) [17:02:00] dzahn@cumin2002 makevm (PID 3715560) is awaiting input [17:02:02] (03CR) 10Effie Mouzeli: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1320132 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [17:02:55] (03Merged) 10jenkins-bot: Add CacheAbstractContentFragment to high_traffic_jobs with concurrency=1 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1314852 (https://phabricator.wikimedia.org/T430898) (owner: 10Genoveva Galarza) [17:04:04] !log rzl@deploy1003 helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply [17:04:21] !log rzl@deploy1003 helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply [17:04:47] (03PS1) 10Scott French: wmnet: Reduce _etcd._tcp.conftool (R/W) SRV TTL to 10s [dns] - 10https://gerrit.wikimedia.org/r/1319194 (https://phabricator.wikimedia.org/T433554) [17:04:48] (03PS1) 10Scott French: wmnet: Switch _etcd._tcp.conftool (R/W) SRV hosts to codfw [dns] - 10https://gerrit.wikimedia.org/r/1319195 (https://phabricator.wikimedia.org/T433554) [17:04:49] (03PS1) 10Scott French: wmnet: Restore _etcd._tcp.conftool (R/W) SRV TTL to 5M [dns] - 10https://gerrit.wikimedia.org/r/1319196 (https://phabricator.wikimedia.org/T433554) [17:05:46] !log rzl@deploy1003 helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply [17:06:03] !log dzahn@cumin2002 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" [17:06:03] !log dzahn@cumin2002 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [17:06:04] !log dzahn@cumin2002 START - Cookbook sre.dns.wipe-cache codesearch1001.eqiad.wmnet on all recursors [17:06:07] !log dzahn@cumin2002 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch1001.eqiad.wmnet on all recursors [17:06:37] !log dzahn@cumin2002 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" [17:06:40] !log rzl@deploy1003 helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply [17:06:42] !log dzahn@cumin2002 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002" [17:07:40] 10ops-eqiad, 06SRE, 06DBA, 06DC-Ops: db1218 crashed - https://phabricator.wikimedia.org/T433565#12180100 (10VRiley-WMF) a:05Marostegui→03VRiley-WMF [17:08:44] !log dzahn@cumin2002 START - Cookbook sre.hosts.reimage for host codesearch1001.eqiad.wmnet with OS trixie [17:08:56] 06SRE, 06Infrastructure-Foundations, 10vm-requests: eqiad/codfw: 2 VM request for codesearch - https://phabricator.wikimedia.org/T433316#12180102 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by dzahn@cumin2002 for host codesearch1001.eqiad.wmnet with OS trixie [17:09:21] some "broker transport failure" errors in changeprop startup but ISTR we've seen that before, I don't think it's related to this config change -- looking though [17:09:32] (03PS1) 10Jgiannelos: mobileapps: Pin staging to latest master [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320201 [17:10:11] ok.. thanks! crossing fingers [17:11:42] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12180113 (10VRiley-WMF) @Andrew Awesome, sounds good. Would we like to start on 1049? [17:11:56] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12180114 (10VRiley-WMF) [17:12:04] (03PS1) 10Scott French: hieradata: etcd read-only in eqiad [puppet] - 10https://gerrit.wikimedia.org/r/1319191 (https://phabricator.wikimedia.org/T433554) [17:12:05] (03PS1) 10Scott French: hieradata: switch etcd replication from codfw to eqiad [puppet] - 10https://gerrit.wikimedia.org/r/1319192 (https://phabricator.wikimedia.org/T433554) [17:12:06] (03PS1) 10Scott French: hieradata: etcd read-write in codfw [puppet] - 10https://gerrit.wikimedia.org/r/1319193 (https://phabricator.wikimedia.org/T433554) [17:12:20] (03Abandoned) 10Jgiannelos: mobileapps: Pin staging to latest master [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320201 (owner: 10Jgiannelos) [17:13:17] hm, they aren't settling out -- it's the same set of six pods though, wonder if they're all on the same rack or something [17:13:42] (03CR) 10Dzahn: "if the use case is some external user doing ssh to the replica and spare then I think we can't avoid going through CDN" [puppet] - 10https://gerrit.wikimedia.org/r/1314095 (owner: 10Dzahn) [17:14:08] (03CR) 10JHathaway: [C:03+1] profile::memcached::instance: migrate to PKI (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1320123 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [17:14:33] (03CR) 10Dzahn: [C:03+1] "ACK, thanks" [puppet] - 10https://gerrit.wikimedia.org/r/1290731 (https://phabricator.wikimedia.org/T425441) (owner: 10Arnaudb) [17:14:50] no, four in row E and two in row F, and all different racks. didn't love that theory anyway [17:15:02] (03CR) 10Dzahn: [C:03+1] "well, merge whenever but in coordination with Sukhbir" [dns] - 10https://gerrit.wikimedia.org/r/1310135 (owner: 10Dzahn) [17:15:27] ack [17:15:58] I'd like to get this figured out just in case there *is* something about the new config that causes that crash [17:15:59] rzl: do you have an example pod ID? [17:16:10] swfrench-wmf: changeprop-production-697d48f9c7-bfrqx just now [17:16:37] it's not actually on startup, I misspoke at first -- it's any time in the first couple of minutes, so I'm speculating it might be when an event comes in that's now poison [17:16:44] (03PS4) 10Milazg: Add configurable RestTermsOfServiceUrl [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320163 (https://phabricator.wikimedia.org/T428147) [17:17:28] swfrench-wmf: if you `kube-env changeprop-jobqueue codfw; kubectl get pod` you'll see the ones with nonzero restarts [17:17:31] !log dzahn@cumin2002 START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage [17:17:54] * swfrench-wmf thumbs up [17:19:46] trying to sort out what error code -195 in rdkafka [17:20:08] a little tempted to try rolling back the config to see if the errors go with it -- or possibly bouncing one of the bad pods, just to make sure it isn't an effect of the restart, but it's hard to come up with a mechanism for that [17:21:01] there's something very upsetting about the error code being negative. I don't like that. [17:21:16] I would say maybe whackamole first to see if it persists [17:21:20] I can't tell you a specific reason it shouldn't be allowed, but I don't like it [17:22:42] yeah -- I'm going to kubectl delete pod changeprop-production-697d48f9c7-bfrqx and see where it goes [17:22:57] (-195 is not terribly helpful here, it's just the code corresponding to the error message we already have) [17:23:03] thank you for the eyeballs btw [17:23:25] no problem [17:23:47] done, and changeprop-production-697d48f9c7-jdplw is the replacement [17:24:25] and it did just restart [17:24:30] (Apparently they use negative error codes for internal errors and positive ones for broker errors) [17:24:58] !log dzahn@cumin2002 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage [17:25:50] oh interesting, that new pod is giving me something different in the logs [17:25:50] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, August 03 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-ite" [extensions/Kartographer] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1320171 (https://phabricator.wikimedia.org/T433703) (owner: 10Jforrester) [17:25:57] "message":"Broker: Unknown topic or partition" [17:26:19] ah! that's a better one ... I think? or is that fatal as well? [17:26:26] that's ERROR yeah [17:26:31] awesome [17:26:34] and under rule_name: low_traffic_jobs [17:26:43] ah, that tracks [17:27:00] that pod did also crash once though, which is why I'm confused, but I'm getting the logs from the new iteration [17:27:05] if this job is new and has never been produced while codfw is the primary dc [17:27:17] the topic won't exist [17:27:22] (in codfw that is) [17:27:59] nod [17:28:06] hm I think this job is new from last quarter [17:28:22] that tracks, then [17:28:25] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, August 03 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-ite" [extensions/TimedMediaHandler] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1320159 (https://phabricator.wikimedia.org/T433666) (owner: 10Jforrester) [17:28:35] * genocation sweats [17:28:59] seeing it under low_traffic_jobs makes me think it's some other job though, and maybe pre-existing? since we shouldn't be consumed there anymore [17:29:39] I didn't think to kubectl logs before deploying for comparison, checking logstash to see what I can find [17:29:42] rzl: it's in the exclusion list in low-traffic [17:29:49] ohh okay [17:30:23] (03CR) 10TheDJ: [C:04-1] "you are right, wikimania still mindwarps my perception of time" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319945 (https://phabricator.wikimedia.org/T152294) (owner: 10TheDJ) [17:31:24] still only the one crash for the new pod -jdplw, and the others seem to be slowing down -- I still wish I knew what was going on there but it seems like maybe something converged gradually? [17:31:32] alright, wild theory: I just checked for admin_ng diffs in codfw, and I'm not seeing anything, so scratch that. [17:31:47] I mean, two replicas still in CrashLoopBackOff, not like we're done [17:32:08] I'm tempted to whack the remaining moles all at once just to see what happens [17:32:28] * swfrench-wmf endorses this course of action [17:33:34] (I say "all at once," I'm going two at a time) [17:34:56] okay, both of the first two crashed early on [17:35:38] and both with the same "message":"broker transport failure" -- I also just caught "msg":"Message not supplied" in there [17:36:14] so it feels more and more like a real config bug [17:36:39] I'd like to identify what it is, before we roll back, just so we know what to fix though [17:37:24] (03CR) 10Bking: [C:03+2] cirrus: Add a defined type for master-eligibles [puppet] - 10https://gerrit.wikimedia.org/r/1319891 (https://phabricator.wikimedia.org/T433726) (owner: 10Bking) [17:37:26] agreed, yeah [17:37:52] !log dzahn@cumin2002 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch1001.eqiad.wmnet with OS trixie [17:37:53] !log dzahn@cumin2002 END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch1001.eqiad.wmnet [17:38:01] 06SRE, 06Infrastructure-Foundations, 10vm-requests: eqiad/codfw: 2 VM request for codesearch - https://phabricator.wikimedia.org/T433316#12180256 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by dzahn@cumin2002 for host codesearch1001.eqiad.wmnet with OS trixie completed: - codesearch10... [17:38:05] if the jobs are malformed I'd have thought this would just move the errors from one consumer to another [17:38:38] this feels squarely like a networking problem [17:39:37] yeah, I'm just thrown because afaik the network situation shouldn't have changed [17:39:50] 10ops-eqiad, 06SRE, 06DC-Ops: Unresponsive management for an-worker1147.mgmt:22 - https://phabricator.wikimedia.org/T433348#12180259 (10VRiley-WMF) a:03VRiley-WMF [17:41:38] rzl: let me know if there's anything I can do to help diagnose [17:41:44] !log dzahn@cumin2002 START - Cookbook sre.ganeti.makevm for new host codesearch2001.codfw.wmnet [17:41:47] !log dzahn@cumin2002 START - Cookbook sre.dns.netbox [17:42:51] 10ops-codfw, 06SRE, 06DC-Ops, 13Patch-For-Review: Q4:rack/setup/install sretest2010 Config J 1P test host - https://phabricator.wikimedia.org/T394357#12180267 (10Jhancock.wm) well i thought i had a drive i could swap. i do not have a SAS drive. just sata. i'll do option 3 in on tuesday morning [17:43:47] alright, I can't seem to find any network policy issue that might explain this [17:45:04] the callstack leads here https://github.com/Blizzard/node-rdkafka/blob/v3.0.0/lib/client.js#L350 (for the version of node-rdkafka listed in changeprop's package.json) so it looks like this *is* associated with having to create the topic [17:45:40] FIRING: SystemdUnitFailed: generate_vrts_aliases.service on mx-in2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [17:45:54] but unclear why that crashes sometimes instead of just the "hey, no topic" ERROR [17:46:36] and we don't get a full stack, just error.js and that line [17:48:25] !log dzahn@cumin2002 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" [17:50:51] that would be absolutely wild if this gets surfaced as "broker transport failure" [17:50:52] hm, with ten minutes left in the window I'm inclined to start rolling back [17:51:12] that sounds reasonable - I am, alas, out of ideas [17:51:30] dzahn@cumin2002 makevm (PID 3729444) is awaiting input [17:51:30] yeah, right? I don't think the root cause actually is just that there's no topic -- but I'd believe there's an edge case somewhere that we're only tickling because there isn't one [17:52:06] but I don't have any smarter or more specific speculations about it unfortunately [17:52:43] just to confirm -- I've done a plain old "helmfile apply" in staging and then codfw, but nothing else -- any changeprop-jobqueue specific steps that you know about that I don't? [17:53:03] nothing that I'm aware of, no! [17:53:11] (and, I was waiting for codfw to stabilize before starting eqiad, any reason that's actually a huge mistake that's causing all this?) [17:53:22] like "oh they actually need to match" [17:53:27] !log dzahn@cumin2002 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002" [17:53:28] !log dzahn@cumin2002 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [17:53:28] !log dzahn@cumin2002 START - Cookbook sre.dns.wipe-cache codesearch2001.codfw.wmnet on all recursors [17:53:32] !log dzahn@cumin2002 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch2001.codfw.wmnet on all recursors [17:53:39] no, since the codfw instance will only ever touch the codfw topics [17:53:52] yeah [17:53:57] just wanted to rule it out [17:54:04] !log dzahn@cumin2002 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" [17:54:10] !log dzahn@cumin2002 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002" [17:54:11] okay I'm going to start rolling back. sorry genocation :( [17:54:41] v_v no worries, thanks!! [17:54:58] https://github.com/Blizzard/node-rdkafka/blob/v3.0.0/lib/client.js#L336 - this is where I'm squinting a bit. I wonder if this is timing out, which is in turn surfaced as "transport failure" [17:54:59] (03PS1) 10RLazarus: Revert "Add CacheAbstractContentFragment to high_traffic_jobs with concurrency=1" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320210 [17:55:13] !log dzahn@cumin2002 START - Cookbook sre.hosts.reimage for host codesearch2001.codfw.wmnet with OS trixie [17:55:28] 06SRE, 06Infrastructure-Foundations, 10vm-requests: eqiad/codfw: 2 VM request for codesearch - https://phabricator.wikimedia.org/T433316#12180280 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by dzahn@cumin2002 for host codesearch2001.codfw.wmnet with OS trixie [17:55:37] (03CR) 10Scott French: [C:03+1] Revert "Add CacheAbstractContentFragment to high_traffic_jobs with concurrency=1" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320210 (owner: 10RLazarus) [17:55:58] swfrench-wmf: huh, yeah [17:56:08] I'd still have so many questions, but [17:56:39] well, if there's like a socket timeout under the hood ... [17:58:02] (03CR) 10RLazarus: [C:03+2] Revert "Add CacheAbstractContentFragment to high_traffic_jobs with concurrency=1" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320210 (owner: 10RLazarus) [17:58:34] yeah [18:00:45] (03Merged) 10jenkins-bot: Revert "Add CacheAbstractContentFragment to high_traffic_jobs with concurrency=1" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320210 (owner: 10RLazarus) [18:01:15] !log rzl@deploy1003 helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply [18:01:26] !log rzl@deploy1003 helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply [18:02:15] !log rzl@deploy1003 helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply [18:02:35] okay let's see [18:02:55] !log rzl@deploy1003 helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply [18:04:19] okay well, good news bad news [18:04:25] good news, the config wasn't at fault after all [18:04:25] lol [18:04:48] lol [18:04:52] bad news, wow that's a lot more pods affected this time. I'm going to see if they're all still in E and F [18:05:34] D: [18:05:43] ack [18:07:59] ... and the answer is yes actually [18:08:00] FIRING: HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-rollback - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs-next - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [18:08:15] wow, seriously? [18:08:29] https://www.irccloud.com/pastebin/zKRmW5Tv/ [18:08:34] look how satisfying that is [18:08:46] changeprop-jobqueue is allergic to node numbers greater than 2,331 [18:09:19] and (https://netbox.wikimedia.org/search/?q=wikikube-worker2) they just so happen to be mostly correlated with location [18:09:53] ... [18:09:57] this is not a type of cursed I expected to encounter [18:10:09] I'm not crazy though right? you see it too? [18:10:18] 100% [18:10:23] I'm here right with you [18:10:43] so, on the one hand, I guess we have an easy workaround which is just a scheduling constraint [18:10:58] !log jhancock@cumin2002 START - Cookbook sre.network.configure-switch-interfaces for host mc2046 [18:11:01] AND a pretty good straightforward repro case when we have a fix for... whatever this is [18:11:14] on the other hand I have no idea what this is [18:11:16] !log jhancock@cumin2002 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 [18:11:25] !log jhancock@cumin2002 START - Cookbook sre.dns.netbox [18:12:17] well, AFAIK E&F still have a slightly different network design than the other rows, right? [18:12:38] !log dzahn@cumin2002 START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage [18:12:56] rzl: what would be amazing is if there's something cursed firewall-config wise that depends on these hosts being in conftool as workers. since they never will be for the reason bblack mentions. [18:13:11] i.e., they have no L2 adjacency with LVS [18:13:16] oh man [18:13:32] ... so we don't include them in kubesvc [18:14:00] yeah - I'm still missing the part of the story where that only gets triggered on the missing-topic path though [18:14:14] like I could easily believe "always works" or "never works" [18:14:31] does the missing-topic path imply reaching out to a different service than others? [18:14:38] !log jhancock@cumin2002 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [18:14:48] !log jhancock@cumin2002 START - Cookbook sre.network.configure-switch-interfaces for host mc2046 [18:15:43] I wouldn't have thought so, but if so it would fit yeah [18:17:54] jhancock@cumin2002 configure-switch-interfaces (PID 3736094) is awaiting input [18:18:29] !log dzahn@cumin2002 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage [18:18:41] so, fun fact: the topic does exist, at least as of now: https://www.irccloud.com/pastebin/tnj7Ij1q/ [18:19:42] hm! [18:19:58] !log jhancock@cumin2002 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 [18:20:00] which is to say, auto-creation on topic subscription might have "worked" as expected here [18:20:23] so that could have been a red herring. but man it's a great answer for "why is this an ERROR sometimes and FATAL other times" [18:20:39] because we do see ERRORs on the replicas on lower-numbered hosts [18:20:52] but those were from the low-traffic consumer, right? [18:21:00] hm. yeah [18:21:35] e.g., those would emit errors because they included an exclusion for a topic that did not yet exist in their subscription configuration [18:21:49] oh I see what you're saying [18:21:59] while the new dedicated subscriber would eventually get around to doing so(?) via auto-create [18:22:17] and you're right, those ERRORs stopped on the rollback, I missed that in the timestamps [18:22:47] so we're back to "wtf, network" :) [18:23:02] PROBLEM - PyBal backends health check on lvs2014 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs2014.codfw.wmnet, wdqs2022.codfw.wmnet, wdqs2010.codfw.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [18:23:02] PROBLEM - PyBal backends health check on lvs2013 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs2013.codfw.wmnet, wdqs2010.codfw.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [18:24:02] RECOVERY - PyBal backends health check on lvs2014 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [18:24:02] RECOVERY - PyBal backends health check on lvs2013 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [18:24:32] yeah [18:25:06] fun additional fact: kafka-main2006-2010 (i.e., the whole cluster) are in racks A-D [18:25:28] (03CR) 10TrainBranchBot: [C:03+2] "Approved by krinkle@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319868 (https://phabricator.wikimedia.org/T247113) (owner: 10Krinkle) [18:25:42] hm. [18:27:08] (03Merged) 10jenkins-bot: logging: Remove mention of 'fatal' channel that no longer exists [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319868 (https://phabricator.wikimedia.org/T247113) (owner: 10Krinkle) [18:27:22] !log krinkle@deploy1003 Started scap sync-world: Backport for [[gerrit:1319868|logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] [18:27:26] T247113: Standardise on Logstash channel for fatal exceptions (fatal, exception, DeferredUpdates) - https://phabricator.wikimedia.org/T247113 [18:29:10] !log krinkle@deploy1003 krinkle: Backport for [[gerrit:1319868|logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [18:30:13] rzl: I might nsenter into changeprop-production-6fbff89654-rsm4s and try something funny [18:30:18] go for it [18:30:20] mind leaving that one intact [18:30:22] awesome [18:33:27] !log krinkle@deploy1003 krinkle: Continuing with deployment [18:34:19] !log dzahn@cumin2002 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch2001.codfw.wmnet with OS trixie [18:34:20] !log dzahn@cumin2002 END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch2001.codfw.wmnet [18:34:26] 06SRE, 06Infrastructure-Foundations, 10vm-requests: eqiad/codfw: 2 VM request for codesearch - https://phabricator.wikimedia.org/T433316#12180335 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by dzahn@cumin2002 for host codesearch2001.codfw.wmnet with OS trixie completed: - codesearch20... [18:34:46] rzl: I'm signing off for the day, will catch up with you tomorrow, thanks you!! [18:35:04] 06SRE, 10LDAP-Access-Requests: Grant Access to NDA for jaleman-vdr-wmf - https://phabricator.wikimedia.org/T433417#12180336 (10KFrancis) Hi all, waiting on this one. If this person is currently a contractor with the WMF, they probably don't need an separate NDA, but I'll need their actual name to verify. Ple... [18:35:22] rzl: I only have ipv4 connectivity to the kafka brokers. going to cross-check against a healthy pod in the other racks. [18:35:25] RESOLVED: SystemdUnitFailed: generate_vrts_aliases.service on mx-in2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [18:36:17] 10ops-codfw, 06SRE, 06DC-Ops, 06ServiceOps: Rerack mc-gp2006 and mc2046 - https://phabricator.wikimedia.org/T432716#12180337 (10Jhancock.wm) it's magical to me. Updated the server location. I was gonna give the reimage a try but i didn't know which distro was on it. @jijiki lmk which distro or you can tak... [18:37:35] !log krinkle@deploy1003 Finished scap sync-world: Backport for [[gerrit:1319868|logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] (duration: 10m 13s) [18:37:36] rzl: aaaaand I have dual-stack connectivity on wikikube-worker2284 [18:37:39] T247113: Standardise on Logstash channel for fatal exceptions (fatal, exception, DeferredUpdates) - https://phabricator.wikimedia.org/T247113 [18:37:46] seems the network is not hap [18:38:38] (03CR) 10TrainBranchBot: [C:03+2] "Approved by krinkle@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311041 (https://phabricator.wikimedia.org/T271001) (owner: 10Krinkle) [18:38:51] (03CR) 10CI reject: [V:04-1] MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311041 (https://phabricator.wikimedia.org/T271001) (owner: 10Krinkle) [18:39:44] (03PS1) 10Ahmon Dancy: update_version: use the ruamel.yaml YAML() instance API [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320211 (https://phabricator.wikimedia.org/T433875) [18:40:58] rzl: for the record: https://www.irccloud.com/pastebin/fsrUaYrj/ [18:41:07] (03PS5) 10Krinkle: MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311041 (https://phabricator.wikimedia.org/T271001) [18:41:16] (03CR) 10CI reject: [V:04-1] MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311041 (https://phabricator.wikimedia.org/T271001) (owner: 10Krinkle) [18:41:48] (03PS6) 10Krinkle: MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311041 (https://phabricator.wikimedia.org/T271001) [18:43:52] swfrench-wmf: ok to deploy? Happy to wait until the next window. [18:44:14] noticed you mention wikikube being unhappy [18:44:53] swfrench-wmf: oh, nice find [18:45:17] Krinkle: I think you're good to go. we have something kinda tricky happening with changeprop, but I don't think we have any reason to hold backports etc. [18:45:18] (03CR) 10Ahmon Dancy: "Just noticed that there was a lingering fix for this already in Id9c1da6cf10c0f7cac8061abb8268091c1a87ce1." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320211 (https://phabricator.wikimedia.org/T433875) (owner: 10Ahmon Dancy) [18:45:27] (03Abandoned) 10Ahmon Dancy: update_version: use the ruamel.yaml YAML() instance API [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320211 (https://phabricator.wikimedia.org/T433875) (owner: 10Ahmon Dancy) [18:46:22] so, I'm tempted to add a nodeAffinity for cpjq for now sticking it to the other rows -- we do have a node label for rack at least [18:46:25] k [18:46:39] and obviously hang that off a task for "figure this out and fix it" [18:47:04] (03CR) 10TrainBranchBot: [C:03+2] "Approved by krinkle@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311041 (https://phabricator.wikimedia.org/T271001) (owner: 10Krinkle) [18:47:30] (03PS8) 10Ahmon Dancy: Update update_version.py to be compatible with ruamel >=0.15.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310090 (https://phabricator.wikimedia.org/T433875) (owner: 10Mvolz) [18:47:55] (03CR) 10Ahmon Dancy: [C:03+2] Update update_version.py to be compatible with ruamel >=0.15.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310090 (https://phabricator.wikimedia.org/T433875) (owner: 10Mvolz) [18:47:57] rzl: sounds good to me! my wild theory is that we're not announcing the v6 pod IPs to the ToRs properly, but that will require netops to confirm what the ToRs are seeing (and what they're in turn advertising onward) [18:48:29] (03Merged) 10jenkins-bot: MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311041 (https://phabricator.wikimedia.org/T271001) (owner: 10Krinkle) [18:48:42] !log krinkle@deploy1003 Started scap sync-world: Backport for [[gerrit:1311041|MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] [18:48:46] T271001: Transition to MathML rendering as default - https://phabricator.wikimedia.org/T271001 [18:49:36] nod [18:49:51] (03Merged) 10jenkins-bot: Update update_version.py to be compatible with ruamel >=0.15.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310090 (https://phabricator.wikimedia.org/T433875) (owner: 10Mvolz) [18:50:04] jasmine@cumin2002 roll-restart-reboot-brokers (PID 3744314) is awaiting input [18:50:29] !log krinkle@deploy1003 krinkle: Backport for [[gerrit:1311041|MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [18:53:05] !log jasmine@cumin2002 START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-codfw [18:53:59] (03PS1) 10PipelineBot: mobileapps: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320213 [18:53:59] !log krinkle@deploy1003 krinkle: Continuing with deployment [18:54:42] (03CR) 10Ahmon Dancy: [C:03+2] "Thanks for the fix mvolz!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310090 (https://phabricator.wikimedia.org/T433875) (owner: 10Mvolz) [18:56:25] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [18:58:05] !log krinkle@deploy1003 Finished scap sync-world: Backport for [[gerrit:1311041|MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] (duration: 09m 23s) [18:58:10] T271001: Transition to MathML rendering as default - https://phabricator.wikimedia.org/T271001 [18:58:44] swfrench, rzl: apologies didn't see the hold in time, should I terminate the cookbook following the first broker reboot? [18:59:32] I think it won't materially change the situation for changeprop-jobqueue but I definitely wanted to coordinate it in case swfrench-wmf is still collecting data [19:01:17] thanks! jasmine_, rzl: I think I have enough to vaguely gesture toward the problem at this point. I'd say feel free to move ahead with the restarts. [19:01:42] i.e., there's no longer a risk of me confusing a restart with a reachability issue [19:03:30] 06SRE, 06Infrastructure-Foundations, 10vm-requests: eqiad/codfw: 2 VM request for codesearch - https://phabricator.wikimedia.org/T433316#12180406 (10Dzahn) 05Open→03Resolved thanks! VMs created. [19:03:47] (03PS1) 10Clare Ming: eventLogUtils: Set schema on Test Kitchen instrument before send [extensions/WikiLambda] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1320215 (https://phabricator.wikimedia.org/T433550) [19:05:52] swfrench: makes sense, thanks! will keep an eye here if anything changes) [19:09:31] 10ops-codfw, 06SRE, 06DBA, 06DC-Ops: db2209 firmware upgrade after a crash - https://phabricator.wikimedia.org/T431952#12180433 (10Jhancock.wm) Got something back. No findings on their part. He acknowledged that the idrac and bios were update but said to "update the other components" without going into spe... [19:11:02] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, August 03 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-ite" [extensions/WikiLambda] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1320215 (https://phabricator.wikimedia.org/T433550) (owner: 10Clare Ming) [19:14:50] (03PS1) 10AOkoth: site: apply migration role to phab1005 [puppet] - 10https://gerrit.wikimedia.org/r/1320217 (https://phabricator.wikimedia.org/T377889) [19:15:57] (03CR) 10David Martin: [C:03+1] eventLogUtils: Set schema on Test Kitchen instrument before send [extensions/WikiLambda] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1320215 (https://phabricator.wikimedia.org/T433550) (owner: 10Clare Ming) [19:32:48] (03PS1) 10RLazarus: changeprop: Add a nodeAffinity rule against codfw rows E and F [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320218 (https://phabricator.wikimedia.org/T433881) [19:37:11] (03CR) 10RLazarus: changeprop: Add a nodeAffinity rule against codfw rows E and F (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320218 (https://phabricator.wikimedia.org/T433881) (owner: 10RLazarus) [19:39:20] (03CR) 10RLazarus: [C:03+2] "Post-revert update: The problem actually turned out to be unrelated to this change; see T433881. Once we've worked around it with I1fcd7d8" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320210 (owner: 10RLazarus) [19:43:07] (03CR) 10Scott French: [C:03+1] changeprop: Add a nodeAffinity rule against codfw rows E and F [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320218 (https://phabricator.wikimedia.org/T433881) (owner: 10RLazarus) [19:43:44] FIRING: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [19:45:25] !log jasmine@cumin2002 END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-codfw [19:47:56] (03PS2) 10Arlolra: Turn on PRV for all namespaces on enwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320179 (https://phabricator.wikimedia.org/T430194) [19:48:44] RESOLVED: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [19:49:11] (03CR) 10Subramanya Sastry: [C:03+1] Turn on PRV for all namespaces on enwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320179 (https://phabricator.wikimedia.org/T430194) (owner: 10Arlolra) [20:00:05] RoanKattouw, urbanecm, TheresNoTime, kindrobot, and cjming: May I have your attention please! UTC late backport window. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260803T2000) [20:00:05] thedj, Krinkle, arlolra, and cjming: A patch you scheduled for UTC late backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [20:00:10] o/ [20:00:31] o/ [20:00:36] i can deploy for anyone who needs a deployer [20:01:21] o/ [20:03:00] FIRING: [2x] HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-rollback - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [20:03:07] not sure thedj is around [20:03:16] Krinkle: do you want to go first? [20:06:00] sure [20:06:09] all you [20:06:18] (03CR) 10TrainBranchBot: [C:03+2] "Approved by krinkle@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319903 (https://phabricator.wikimedia.org/T414338) (owner: 10Krinkle) [20:07:29] (03Merged) 10jenkins-bot: Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319903 (https://phabricator.wikimedia.org/T414338) (owner: 10Krinkle) [20:07:38] FIRING: [2x] CertAlmostExpired: gNMI TLS certificate for lsw1-d6-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [20:07:40] !log krinkle@deploy1003 Started scap sync-world: Backport for [[gerrit:1319903|Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] [20:07:44] T414338: FY25-26 WE5.4.12: Identify the provenance of image requests - https://phabricator.wikimedia.org/T414338 [20:08:57] FIRING: [2x] HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-rollback - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [20:09:29] !log krinkle@deploy1003 krinkle: Backport for [[gerrit:1319903|Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [20:10:09] (03PS7) 10Ahmon Dancy: wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1315968 (https://phabricator.wikimedia.org/T433104) [20:10:23] (03CR) 10Ahmon Dancy: wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (031 comment) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1315968 (https://phabricator.wikimedia.org/T433104) (owner: 10Ahmon Dancy) [20:12:02] !log krinkle@deploy1003 krinkle: Continuing with deployment [20:16:06] !log krinkle@deploy1003 Finished scap sync-world: Backport for [[gerrit:1319903|Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] (duration: 08m 26s) [20:16:10] T414338: FY25-26 WE5.4.12: Identify the provenance of image requests - https://phabricator.wikimedia.org/T414338 [20:18:00] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [20:23:00] FIRING: [3x] HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-rollback - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [20:24:16] cjming: Is it me now? [20:24:27] Krinkle: are you done? [20:25:04] Arlolra: yes if Krinkle is finished - are you ok self-deploying or do you need a deployer? [20:25:17] I can deploy [20:26:39] I'll give Krinkle a couple minutes to chime in [20:26:50] cool [20:29:52] (03PS2) 10Scott French: hieradata: etcd read-write in codfw [puppet] - 10https://gerrit.wikimedia.org/r/1319193 (https://phabricator.wikimedia.org/T433554) [20:31:53] Arlolra: i think it's safe to proceed - lmk when you're done [20:31:59] ok [20:32:22] arlolra: yes [20:32:41] (I'm done) [20:32:49] thanks! [20:34:32] (03Merged) 10jenkins-bot: Turn on PRV for all namespaces on enwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320179 (https://phabricator.wikimedia.org/T430194) (owner: 10Arlolra) [20:34:47] !log arlolra@deploy1003 Started scap sync-world: Backport for [[gerrit:1320179|Turn on PRV for all namespaces on enwiki (T430194)]] [20:34:51] T430194: Parsoid Read Views deploy to English Wikipedia (enwiki) June 25-June 30 - https://phabricator.wikimedia.org/T430194 [20:36:35] !log arlolra@deploy1003 arlolra: Backport for [[gerrit:1320179|Turn on PRV for all namespaces on enwiki (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [20:38:19] !log arlolra@deploy1003 arlolra: Continuing with deployment [20:39:49] (03CR) 10Scott French: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1319193 (https://phabricator.wikimedia.org/T433554) (owner: 10Scott French) [20:42:23] !log arlolra@deploy1003 Finished scap sync-world: Backport for [[gerrit:1320179|Turn on PRV for all namespaces on enwiki (T430194)]] (duration: 07m 36s) [20:42:28] T430194: Parsoid Read Views deploy to English Wikipedia (enwiki) June 25-June 30 - https://phabricator.wikimedia.org/T430194 [20:42:35] cjming: done [20:42:40] great - thanks [20:42:47] thedj: are you around? [20:42:57] i'm going to start my backport now [20:43:46] (03CR) 10TrainBranchBot: [C:03+2] "Approved by cjming@deploy1003 using scap backport" [extensions/WikiLambda] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1320215 (https://phabricator.wikimedia.org/T433550) (owner: 10Clare Ming) [20:45:50] (03Merged) 10jenkins-bot: eventLogUtils: Set schema on Test Kitchen instrument before send [extensions/WikiLambda] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1320215 (https://phabricator.wikimedia.org/T433550) (owner: 10Clare Ming) [20:46:06] !log cjming@deploy1003 Started scap sync-world: Backport for [[gerrit:1320215|eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] [20:46:10] T433550: [Abstract Wikipedia] Events are not being sent from migrated wikifunctions-ui-actions instrument - https://phabricator.wikimedia.org/T433550 [20:47:52] !log cjming@deploy1003 cjming: Backport for [[gerrit:1320215|eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [20:48:16] !log cjming@deploy1003 cjming: Continuing with deployment [20:52:29] !log cjming@deploy1003 Finished scap sync-world: Backport for [[gerrit:1320215|eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] (duration: 06m 23s) [20:52:42] T433550: [Abstract Wikipedia] Events are not being sent from migrated wikifunctions-ui-actions instrument - https://phabricator.wikimedia.org/T433550 [20:53:28] (03PS1) 10Ahmon Dancy: Bump buildkit image to wmf-v0.32.1 [puppet] - 10https://gerrit.wikimedia.org/r/1320229 (https://phabricator.wikimedia.org/T433879) [20:54:28] (03CR) 10RLazarus: [C:03+1] wmnet: Reduce _etcd._tcp.conftool (R/W) SRV TTL to 10s [dns] - 10https://gerrit.wikimedia.org/r/1319194 (https://phabricator.wikimedia.org/T433554) (owner: 10Scott French) [20:54:29] Are there any more deployments pending? [20:54:33] (03CR) 10RLazarus: [C:03+1] wmnet: Switch _etcd._tcp.conftool (R/W) SRV hosts to codfw [dns] - 10https://gerrit.wikimedia.org/r/1319195 (https://phabricator.wikimedia.org/T433554) (owner: 10Scott French) [20:54:37] i think we're done [20:54:38] (03CR) 10RLazarus: [C:03+1] wmnet: Restore _etcd._tcp.conftool (R/W) SRV TTL to 5M [dns] - 10https://gerrit.wikimedia.org/r/1319196 (https://phabricator.wikimedia.org/T433554) (owner: 10Scott French) [20:54:47] Great. I'm going to deploy a change. [20:55:11] (03CR) 10TrainBranchBot: [C:03+2] "Approved by dancy@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1315968 (https://phabricator.wikimedia.org/T433104) (owner: 10Ahmon Dancy) [20:56:37] (03Merged) 10jenkins-bot: wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1315968 (https://phabricator.wikimedia.org/T433104) (owner: 10Ahmon Dancy) [20:56:49] !log dancy@deploy1003 Started scap sync-world: Backport for [[gerrit:1315968|wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] [20:56:53] T433104: l10n rebuild errors are hidden - https://phabricator.wikimedia.org/T433104 [20:58:34] !log dancy@deploy1003 dancy: Backport for [[gerrit:1315968|wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [20:59:05] !log dancy@deploy1003 dancy: Continuing with deployment [21:00:05] alexsanford, Reedy, sbassett, Maryum, and manfredi: I seem to be stuck in Groundhog week. Sigh. Time for (yet another) Weekly Security deployment window deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260803T2100). [21:02:45] (03PS1) 10Ebernhardson: cirrus: Enable building redirect documents [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320230 (https://phabricator.wikimedia.org/T204089) [21:03:18] (03PS1) 10Ahmon Dancy: Merge remote-tracking branch 'origin/master' into train-dev [mediawiki-config] (train-dev) - 10https://gerrit.wikimedia.org/r/1320231 [21:03:23] !log dancy@deploy1003 Finished scap sync-world: Backport for [[gerrit:1315968|wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] (duration: 06m 34s) [21:03:27] T433104: l10n rebuild errors are hidden - https://phabricator.wikimedia.org/T433104 [21:03:40] (03CR) 10CI reject: [V:04-1] cirrus: Enable building redirect documents [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320230 (https://phabricator.wikimedia.org/T204089) (owner: 10Ebernhardson) [21:04:00] (03CR) 10Ahmon Dancy: [C:03+2] Merge remote-tracking branch 'origin/master' into train-dev [mediawiki-config] (train-dev) - 10https://gerrit.wikimedia.org/r/1320231 (owner: 10Ahmon Dancy) [21:04:33] (03PS2) 10Ebernhardson: cirrus: Enable building redirect documents [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320230 (https://phabricator.wikimedia.org/T204089) [21:06:24] (03Merged) 10jenkins-bot: Merge remote-tracking branch 'origin/master' into train-dev [mediawiki-config] (train-dev) - 10https://gerrit.wikimedia.org/r/1320231 (owner: 10Ahmon Dancy) [21:10:03] (03Restored) 10Dzahn: gerrit: limit the volume of returnable pages [puppet] - 10https://gerrit.wikimedia.org/r/1308081 (https://phabricator.wikimedia.org/T431398) (owner: 10Arnaudb) [21:10:52] (03CR) 10Dzahn: [C:03+2] Bump buildkit image to wmf-v0.32.1 [puppet] - 10https://gerrit.wikimedia.org/r/1320229 (https://phabricator.wikimedia.org/T433879) (owner: 10Ahmon Dancy) [21:11:35] (03CR) 10RLazarus: [C:03+1] hieradata: etcd read-only in eqiad [puppet] - 10https://gerrit.wikimedia.org/r/1319191 (https://phabricator.wikimedia.org/T433554) (owner: 10Scott French) [21:11:39] (03CR) 10RLazarus: [C:03+1] hieradata: switch etcd replication from codfw to eqiad (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1319192 (https://phabricator.wikimedia.org/T433554) (owner: 10Scott French) [21:11:44] (03CR) 10RLazarus: [C:03+1] hieradata: etcd read-write in codfw [puppet] - 10https://gerrit.wikimedia.org/r/1319193 (https://phabricator.wikimedia.org/T433554) (owner: 10Scott French) [21:12:12] (03PS1) 10Ahmon Dancy: mergeMessageFileList: Suppress "doesn't exist" error when --quiet [core] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1320233 (https://phabricator.wikimedia.org/T125678) [21:14:58] (03CR) 10Dzahn: [C:03+1] "Sukhbir, I think Arnaudb answered the context and you guys already agreed on this in private chat." [dns] - 10https://gerrit.wikimedia.org/r/1310135 (owner: 10Dzahn) [21:18:47] (03CR) 10TrainBranchBot: [C:03+2] "Approved by dancy@deploy1003 using scap backport" [core] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1320233 (https://phabricator.wikimedia.org/T125678) (owner: 10Ahmon Dancy) [21:21:12] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host search-loader2002.codfw.wmnet with OS trixie [21:28:57] FIRING: [2x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [21:30:46] (03Merged) 10jenkins-bot: mergeMessageFileList: Suppress "doesn't exist" error when --quiet [core] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1320233 (https://phabricator.wikimedia.org/T125678) (owner: 10Ahmon Dancy) [21:31:00] !log dancy@deploy1003 Started scap sync-world: Backport for [[gerrit:1320233|mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] [21:31:06] T125678: Scap reliance on extension-list in turn requires extensions to be deployed to WMF production before being deployed to the Beta Cluster - https://phabricator.wikimedia.org/T125678 [21:31:06] T433104: l10n rebuild errors are hidden - https://phabricator.wikimedia.org/T433104 [21:32:47] !log dancy@deploy1003 dancy: Backport for [[gerrit:1320233|mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [21:33:09] !log dancy@deploy1003 dancy: Continuing with deployment [21:37:13] !log dancy@deploy1003 Finished scap sync-world: Backport for [[gerrit:1320233|mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] (duration: 06m 13s) [21:37:19] T125678: Scap reliance on extension-list in turn requires extensions to be deployed to WMF production before being deployed to the Beta Cluster - https://phabricator.wikimedia.org/T125678 [21:37:20] T433104: l10n rebuild errors are hidden - https://phabricator.wikimedia.org/T433104 [21:37:53] !log dancy@deploy1003 Installing scap version "4.277.0" for 3 host(s) [21:39:39] !log dancy@deploy1003 Installation of scap version "4.277.0" completed for 3 hosts [21:41:51] !log dancy@deploy1003 Started scap sync-world: testing [21:42:22] !log dancy@deploy1003 Stopping before sync operations [21:42:47] Done messing with scap [21:42:57] !log bking@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage [21:43:52] 10ops-eqiad, 06SRE, 10Ceph, 06cloud-services-team, and 3 others: cloudcephosd1044 boot issues - https://phabricator.wikimedia.org/T429267#12181024 (10VRiley-WMF) a:03VRiley-WMF [21:43:55] 06SRE, 10Ganeti, 06Infrastructure-Foundations: SSH host key verification failures in Ganeti intra node SSH calls after Bullseye update - https://phabricator.wikimedia.org/T309724#12181028 (10bking) Ran into the issue again today while working thru T433890. Are there any workarounds for accessing the console... [21:48:43] (03PS1) 10Jforrester: Revert^2 "Add CacheAbstractContentFragment to high_traffic_jobs with concurrency=1" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320238 (https://phabricator.wikimedia.org/T430898) [21:48:55] (03PS2) 10Jforrester: Revert^2 "Add CacheAbstractContentFragment to high_traffic_jobs with concurrency=1" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320238 (https://phabricator.wikimedia.org/T430898) [21:48:59] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage [21:51:29] oh lol.. i can't read clocks apparently. [22:03:00] FIRING: [2x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [22:08:46] 06SRE, 10Observability-Metrics, 07Wikimedia-production-error: PHP Warning: RedisException: read error on connection (arclamp1001.eqiad.wmnet) - https://phabricator.wikimedia.org/T432730#12181085 (10colewhite) 05Open→03Resolved Seems the memory upgrade has had a positive effect. Optimistically resolv... [22:08:56] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader2002.codfw.wmnet with OS trixie [22:09:01] (03CR) 10Aaron Schulz: api-gateway: further simplify chart by assuming REST Gateway mode #2 (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319122 (https://phabricator.wikimedia.org/T428625) (owner: 10Aaron Schulz) [22:11:37] (03CR) 10Aaron Schulz: api-gateway: further simplify chart by assuming REST Gateway mode #2 (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319122 (https://phabricator.wikimedia.org/T428625) (owner: 10Aaron Schulz) [22:15:04] (03PS6) 10Aaron Schulz: api-gateway: further simplify chart by assuming REST Gateway mode #2 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319122 (https://phabricator.wikimedia.org/T428625) [22:16:02] (03CR) 10Aaron Schulz: api-gateway: further simplify chart by assuming REST Gateway mode #2 (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319122 (https://phabricator.wikimedia.org/T428625) (owner: 10Aaron Schulz) [22:16:31] (03CR) 10BCornwall: [C:03+1] Revert^2 "gitlab: lower TTL for CNAMEs" [dns] - 10https://gerrit.wikimedia.org/r/1310135 (owner: 10Dzahn) [22:23:44] (03CR) 10RLazarus: [C:03+2] changeprop: Add a nodeAffinity rule against codfw rows E and F [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320218 (https://phabricator.wikimedia.org/T433881) (owner: 10RLazarus) [22:26:05] (03Merged) 10jenkins-bot: changeprop: Add a nodeAffinity rule against codfw rows E and F [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320218 (https://phabricator.wikimedia.org/T433881) (owner: 10RLazarus) [22:32:20] (03PS2) 10Scott French: hieradata: switch etcd replication from codfw to eqiad [puppet] - 10https://gerrit.wikimedia.org/r/1319192 (https://phabricator.wikimedia.org/T433554) [22:32:20] (03PS3) 10Scott French: hieradata: etcd read-write in codfw [puppet] - 10https://gerrit.wikimedia.org/r/1319193 (https://phabricator.wikimedia.org/T433554) [22:36:33] !log rzl@deploy1003 helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply [22:36:45] !log rzl@deploy1003 helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply [22:38:24] * swfrench-wmf excitedly watches pods reschedule in places they should be happier [22:38:56] yeah, I have to bump them with kubectl delete pod since requiredDuringSchedulingIgnoredDuringExecution is accurately names [22:39:08] but one just landed on wikikube-worker2357 which I didn't expect [22:39:18] accurately *named [22:39:34] so I'm checking to see what I missed [22:40:38] it ... doesn't have the scheduling constraints?? [22:40:55] yeah! hm. [22:41:03] I specifically asked for no pickles [22:41:16] (03PS6) 10KineticPelagic: rest: Add test server option to REST Sandbox for Wikipedia projects [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1313827 (https://phabricator.wikimedia.org/T408816) [22:42:01] * swfrench-wmf now realizes these pods are still called 6fbff89654, which is suspicious [22:45:09] so ... the live changeprop-production Deployment resource doesn't appear to have the nodeAffinity I was promised by CI :) [22:47:22] rzl: s/nodeAffinity/podAffinity/ ? [22:47:35] yeah, and it doesn't have the same nodeAffinity I was promised in helmfile template either [22:47:46] wait really? [22:48:11] no I think podAffinity is for when you want to be scheduled near some other pod [22:48:33] ah, you're right - I agrred [22:48:33] d [22:48:35] (and likewise podAntiAffinity is for when you *don't* want to be, like for row redundancy) [22:49:49] yeah, no I just completely inverted them in my head for a minute :) [22:50:16] mrm. I'm tempted to try a "helmfile sync" just because I never really comprehended when "helmfile apply" isn't sufficient [22:50:31] but I think that's wrong, because I can't imagine the first apply wouldn't have synced [22:50:43] since it did highlight the diffs, I mean [22:51:12] what happens if you run the diffs now? [22:51:19] nothing, it's clean [22:53:53] talk me out of reaching in there with kubectl edit [22:53:59] lol [22:54:21] I'm just really puzzled as to why we would not be able to read this back [22:54:32] partly just because I'm a little curious if that would make a diff appear -- but I don't really know what I'd do with that information either way, I guess [22:55:26] I guess I would maybe reach for sync first, but that's entirely based on vibes [22:56:30] okay, here comes that [22:56:34] !log rzl@deploy1003 helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: sync [22:56:39] !log rzl@deploy1003 helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: sync [22:56:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [22:57:04] no effect, which I guess is a relief on that front I guess [22:57:18] ... strike one "I guess". or leave them both, I'm guessing HRAD [22:57:20] hard [22:57:41] wait ... there's ... an `affinity:` missing? [22:58:11] oh wow, there is [22:58:11] e.g., compare CI, where nodeAffinity is directly nested under the pod spec [22:58:16] w/ the docs [22:58:25] how did this ever work? [22:58:28] first of all, great eye, second of all, I don't even get an error message for that? [22:59:08] apparently "something" in the path from "steaming pile of values and templates" to "actual k8s API" (inclusive) is highly forgiving [23:00:04] Deploy window Readers deployment window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260803T2300) [23:00:07] man, that part of the template is straight out of _skel but I guess a lot of the examples in there say "affinity: affinity: nodeAffinity" and I didn't [23:00:12] man I got opinions about that [23:00:56] * swfrench-wmf hears you like affinity, so I put an affinity in your affinity [23:01:39] (that's https://gerrit.wikimedia.org/r/plugins/gitiles/operations/deployment-charts/+/refs/heads/master/_scaffold/service/_skel/templates/deployment.yaml.skel#17 and https://gerrit.wikimedia.org/r/plugins/gitiles/operations/deployment-charts/+/refs/heads/master/_scaffold/service/_skel/values.yaml#24, and I guess I misinterpreted what "affinity: {}" followed by that comment meant [23:01:40] ) [23:02:00] okay. well. this whole day is officially cursed. thanks for being on this ride with me. fix incoming [23:02:30] wow ... I ... yeah, I can see how that comment can be interpreted two ways [23:02:54] git: ssh: Could not resolve hostname gerrit.wikimedia.org: Temporary failure in name resolution [23:03:00] it worked on a retry, but get AWAY from me [23:03:12] lol [23:03:30] maybe we need a linter that knows these things? [23:05:06] that would be quite nice [23:05:23] so I thought we had kubeconform for that [23:06:07] I was about to ask whether we had something now or in the past that's "supposed" to do schema validation on the rendered resources [23:06:16] sounds like the answer is "maybe" [23:07:17] (03PS1) 10RLazarus: changeprop: Apply the correct (if surprising) number of "affinity:"s [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320275 (https://phabricator.wikimedia.org/T433881) [23:10:31] (03CR) 10Scott French: [C:03+1] changeprop: Apply the correct (if surprising) number of "affinity:"s [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320275 (https://phabricator.wikimedia.org/T433881) (owner: 10RLazarus) [23:10:57] that looks like the correct number of `affinity:`s to me [23:12:25] (03CR) 10RLazarus: [C:03+2] changeprop: Apply the correct (if surprising) number of "affinity:"s [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320275 (https://phabricator.wikimedia.org/T433881) (owner: 10RLazarus) [23:12:44] sorry I got sidetracked trying to understand deployment-charts CI [23:12:58] that's quite a side quest [23:13:05] there *is* a kubeconform in there somewhere but I don't fully understand what changes it triggers on yet. maybe this one, maybe not? [23:13:20] you'd certainly like it to [23:15:05] (03Merged) 10jenkins-bot: changeprop: Apply the correct (if surprising) number of "affinity:"s [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320275 (https://phabricator.wikimedia.org/T433881) (owner: 10RLazarus) [23:15:11] (03CR) 10Scott French: "Thanks for the review!" [puppet] - 10https://gerrit.wikimedia.org/r/1319192 (https://phabricator.wikimedia.org/T433554) (owner: 10Scott French) [23:17:22] okay, "helmfile diff" does show a diff, let's see what happens [23:17:27] !log rzl@deploy1003 helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply [23:17:51] oho! [23:17:58] this time we did get a new pod spec hash [23:18:20] I've been following this from a distance, and what the fudge affinity: affinity: is just insane [23:18:40] !log rzl@deploy1003 helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply [23:19:25] well so, in defense of my coworkers who are both smart and well-intentioned -- the idea is that under affinity you can also specify a nodeSelector [23:19:46] which is definitely affinity-related even though it goes into a higher level of the Pod spec proto [23:20:24] oh I was assuming the fault was on helm, oops [23:20:27] it's a product of a reasonable series of choices, it just landed in a place I wasn't expecting [23:21:32] alright, we are rows E/F free [23:22:02] where I get a little mad is that helmfile took the resulting yaml and shipped it off to kubernetes, and kubernetes just quietly ignored the part it didn't know what to do with, and nobody saw fit to mention it to the user (me) along the way [23:22:45] Yes that is also the part that perplexes me. [23:22:47] and, yeah! looks like that worked [23:23:06] the best kind of backward and forward API compatibility: silently ignore thing you don't understand [23:23:31] I'm tempted to roll forwad with https://gerrit.wikimedia.org/r/1320238 but I think try it tomorrow is the better part of valor [23:23:59] as illustrated by my decreasingly hinged narration of that whole sitch [23:24:07] :) [23:24:47] and yes, I think that sounds reasonable [23:25:32] (deferring to tomorrow so as not to tempt fate further) [23:26:56] !log robh@cumin1003 START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet [23:27:05] !log robh@cumin1003 END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet [23:29:43] !log robh@cumin1003 START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet [23:29:52] !log robh@cumin1003 END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet [23:38:09] (03CR) 10Cwhite: [C:03+2] profile: add initial opensearch security plugin config [puppet] - 10https://gerrit.wikimedia.org/r/1306792 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [23:41:53] (03PS1) 10TrainBranchBot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1320288 [23:41:53] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1320288 (owner: 10TrainBranchBot) [23:42:41] (03CR) 10RLazarus: [V:03+1] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9101/co" [puppet] - 10https://gerrit.wikimedia.org/r/1272633 (https://phabricator.wikimedia.org/T428490) (owner: 10Clare Ming) [23:44:02] moving in here from where I've been coordinating with cjming -- we'll go ahead with https://gerrit.wikimedia.org/r/1272633 [23:44:11] ty ty !! [23:44:28] (straightforward maintenance script update, not expecting any impact on anything) [23:44:42] 🤞 [23:44:48] (03CR) 10RLazarus: [V:03+1 C:03+2] Add script to get constructive edits for all wikis [puppet] - 10https://gerrit.wikimedia.org/r/1272633 (https://phabricator.wikimedia.org/T428490) (owner: 10Clare Ming) [23:46:00] puppet-merge is done, running puppet on the deploy host now [23:48:56] (03CR) 10Cwhite: [C:03+2] opensearch: add security-plugin specific configuration to opensearch.yml [puppet] - 10https://gerrit.wikimedia.org/r/1318792 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [23:51:51] (03CR) 10CI reject: [V:04-1] Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1320288 (owner: 10TrainBranchBot) [23:55:16] done, deploying mw-cron in eqiad [23:55:26] tysm! [23:55:41] looks like we'll pick up a couple of other undeployed changes, just giving those a quick look [23:57:41] !log rzl@deploy1003 helmfile [eqiad] START helmfile.d/services/mw-cron: apply [23:58:06] !log rzl@deploy1003 helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply [23:58:21] (03PS1) 10Cwhite: logstash: enable reversing security plugin enable [puppet] - 10https://gerrit.wikimedia.org/r/1320291 (https://phabricator.wikimedia.org/T350516) [23:58:52] cjming: there it is https://www.irccloud.com/pastebin/YfAjaRwq/ [23:59:07] so that'll next run in 20 minutes -- do you want to kick off a test run right away? [23:59:17] you're the best - thanks so much!