[00:00:40] FIRING: KubernetesRsyslogDown: rsyslog on wikikube-worker1260:9105 is missing kubernetes logs - https://wikitech.wikimedia.org/wiki/Kubernetes/Logging#Common_issues - https://grafana.wikimedia.org/d/OagQjQmnk?var-server=wikikube-worker1260 - https://alerts.wikimedia.org/?q=alertname%3DKubernetesRsyslogDown [00:05:40] RESOLVED: KubernetesRsyslogDown: rsyslog on wikikube-worker1260:9105 is missing kubernetes logs - https://wikitech.wikimedia.org/wiki/Kubernetes/Logging#Common_issues - https://grafana.wikimedia.org/d/OagQjQmnk?var-server=wikikube-worker1260 - https://alerts.wikimedia.org/?q=alertname%3DKubernetesRsyslogDown [00:15:02] (03CR) 10Ladsgroup: "I need to backport a couple of patches and make sure the Persian translation is there. Need a bit of time." [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308671 (https://phabricator.wikimedia.org/T431514) (owner: 10Jdlrobson) [00:24:12] (03CR) 10RLazarus: "recheck" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308788 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [00:57:31] (03CR) 10RLazarus: "recheck" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308788 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [01:00:05] (03Abandoned) 10TrainBranchBot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1308244 (owner: 10TrainBranchBot) [01:09:22] RESOLVED: [2x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-1/0/2 (Transport: cr1-eqiad:et-1/1/2 (Arelion, IC-374549) {#12267}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [01:12:39] (03PS1) 10TrainBranchBot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1308797 [01:12:39] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1308797 (owner: 10TrainBranchBot) [01:20:56] (03Merged) 10jenkins-bot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1308797 (owner: 10TrainBranchBot) [01:26:19] (03PS1) 10RLazarus: Revert "admin_ng: Install coredns-internalonly in wikikube" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308798 (https://phabricator.wikimedia.org/T427864) [01:29:06] (03CR) 10Scott French: [C:03+1] coredns-internalonly: Delete nodePort [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308788 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [01:35:53] (03CR) 10CI reject: [V:04-1] Revert "admin_ng: Install coredns-internalonly in wikikube" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308798 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [01:37:56] (03CR) 10RLazarus: [V:03+2 C:03+2] "`" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308798 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [01:47:00] (03CR) 10CI reject: [V:04-1] Revert "admin_ng: Install coredns-internalonly in wikikube" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308798 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [01:56:40] FIRING: SystemdUnitFailed: fetch-rings-thanos.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [01:57:01] (03CR) 10RLazarus: [V:03+2] Revert "admin_ng: Install coredns-internalonly in wikikube" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308798 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [01:57:04] (03CR) 10RLazarus: [V:03+2 C:03+2] Revert "admin_ng: Install coredns-internalonly in wikikube" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308798 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [02:00:54] !log mwpresync@deploy2003 Started scap build-images: Publishing wmf/next image [02:05:10] FIRING: [2x] SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [02:06:49] (03Merged) 10jenkins-bot: Revert "admin_ng: Install coredns-internalonly in wikikube" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308798 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [02:07:25] !log mwpresync@deploy2003 Finished scap build-images: Publishing wmf/next image (duration: 06m 31s) [02:09:22] FIRING: [2x] JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:14:22] RESOLVED: [2x] JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:40:25] (03PS1) 10RLazarus: admin_ng: Install coredns-internalonly in wikikube [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308802 (https://phabricator.wikimedia.org/T427864) [02:45:28] (03Abandoned) 10RLazarus: coredns-internalonly: Delete nodePort [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308788 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [02:49:13] (03CR) 10CI reject: [V:04-1] admin_ng: Install coredns-internalonly in wikikube [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308802 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [02:50:00] (03CR) 10RLazarus: "recheck" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308802 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [03:01:21] (03CR) 10RLazarus: "The flaky CI failure (different environments both times...) is unrelated." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308802 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [03:04:45] PROBLEM - Host cp6008 is DOWN: PING CRITICAL - Packet loss = 100% [03:21:15] FIRING: [4x] MediaWikiHighErrorRate: Elevated rate of MediaWiki errors - kube-mw-api-int - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiHighErrorRate [03:26:15] RESOLVED: [4x] MediaWikiHighErrorRate: Elevated rate of MediaWiki errors - kube-mw-api-int - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiHighErrorRate [03:27:28] !log ryankemper@cumin2002 conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad [03:27:29] !log ryankemper@cumin2002 conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad [03:27:30] !log ryankemper@cumin2002 conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad [03:29:21] !log T431311 Repooled eqiad cirrussearch clusters (`chi/omega/psi`) following completion of OpenSearch 2.19 migration [03:29:23] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [03:29:24] T431311: Migrate production OpenSearch clusters from 1.x-2.x - EQIAD - https://phabricator.wikimedia.org/T431311 [03:39:06] (03PS10) 10Cwhite: opensearch_dashboards: add settings required to enable the security plugin [puppet] - 10https://gerrit.wikimedia.org/r/1305774 (https://phabricator.wikimedia.org/T350516) [03:55:34] !log brett@puppetserver1001 conftool action : set/pooled=no; selector: name=cp6008.* [03:55:54] (03CR) 10Cwhite: "PCC OK: https://puppet-compiler.wmflabs.org/output/1305774/8958/" [puppet] - 10https://gerrit.wikimedia.org/r/1305774 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [04:04:30] 10ops-drmrs, 06DC-Ops: hw troubleshooting: Memory failure for cp6008.drmrs.wmnet - https://phabricator.wikimedia.org/T431651 (10BCornwall) 03NEW [04:05:46] 10ops-drmrs, 06DC-Ops, 06Traffic: hw troubleshooting: Memory failure for cp6008.drmrs.wmnet - https://phabricator.wikimedia.org/T431651#12102916 (10BCornwall) [04:09:59] !log brett@cumin2002 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 14 days, 0:00:00 on cp6008.drmrs.wmnet with reason: Hardware failure - T431651 [04:10:02] T431651: hw troubleshooting: Memory failure for cp6008.drmrs.wmnet - https://phabricator.wikimedia.org/T431651 [04:13:21] 10ops-drmrs, 06DC-Ops, 06Traffic: hw troubleshooting: Memory failure for cp6008.drmrs.wmnet - https://phabricator.wikimedia.org/T431651#12102935 (10BCornwall) [04:56:42] (03PS2) 10Ryan Kemper: pontoon: add rkemper-kafka stack [puppet] - 10https://gerrit.wikimedia.org/r/1308355 (https://phabricator.wikimedia.org/T276088) [04:58:03] 06SRE, 06Data-Engineering, 10Kafka-Infrastructure, 06serviceops-radar, and 3 others: Configuration Management for Kafka settings - https://phabricator.wikimedia.org/T276088#12102950 (10RKemper) We've got the third broker up now that some old instances have been cleaned up, so we're all set to start on the... [05:16:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 15.36% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [05:21:15] FIRING: [2x] PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at codfw: 23.73% idle - https://bit.ly/wmf-fpmsat - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [05:22:15] FIRING: MediaWikiLatencyExceeded: p75 latency high: eqiad mw-web releases routed via main (k8s) 2.5s - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook#Average_latency_exceeded - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=55&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiLatencyExceeded [05:26:15] RESOLVED: [2x] PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at codfw: 23.38% idle - https://bit.ly/wmf-fpmsat - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [05:27:15] RESOLVED: MediaWikiLatencyExceeded: p75 latency high: eqiad mw-web releases routed via main (k8s) 868.7ms - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook#Average_latency_exceeded - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=55&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiLatencyExceeded [05:47:14] (03CR) 10Marostegui: tables-catalog: set betafeatures_user_counts to public visibility (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1298329 (https://phabricator.wikimedia.org/T402145) (owner: 10SD0001) [05:56:40] FIRING: SystemdUnitFailed: fetch-rings-thanos.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [06:00:05] Deploy window MediaWiki infrastructure (UTC early) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260709T0600) [06:00:05] marostegui, Amir1, and federico3: #bothumor When your hammer is PHP, everything starts looking like a thumb. Rise for Primary database switchover. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260709T0600). [06:00:09] (03CR) 10Marostegui: tables-catalog: set betafeatures_user_counts to public visibility (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1298329 (https://phabricator.wikimedia.org/T402145) (owner: 10SD0001) [06:05:10] FIRING: [2x] SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [06:17:43] (03Abandoned) 10Mvolz: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1268047 (owner: 10PipelineBot) [06:47:07] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie [06:57:01] !log rebalance thanos swift rings after previous re-image of thanos-fe1004 to trixie [06:57:02] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [06:58:42] mvernon@cumin2003 reimage (PID 185636) is awaiting input [06:59:45] !log mvernon@cumin2003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1008.eqiad.wmnet with OS trixie [07:00:05] Amir1, urbanecm, and awight: How many deployers does it take to do UTC morning backport window deploy? (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260709T0700). [07:00:05] WMDE-Fisch: A patch you scheduled for UTC morning backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [07:00:14] !log mvernon@cumin2003 START - Cookbook sre.hosts.provision for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [07:01:25] RESOLVED: SystemdUnitFailed: fetch-rings-thanos.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [07:10:00] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [07:12:04] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie [07:12:46] (03CR) 10Ethan.bathon: "recheck" [puppet] - 10https://gerrit.wikimedia.org/r/757622 (https://phabricator.wikimedia.org/T299416) (owner: 10Ladsgroup) [07:25:28] \o [07:25:43] I totally missed deployment window start this morning. [07:26:13] But if nobody complains I would just self serve since there's nothing else in the calendar 🤔 [07:31:32] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage [07:32:48] (03PS1) 10Ayounsi: depool-rack: various fix [cookbooks] - 10https://gerrit.wikimedia.org/r/1309030 (https://phabricator.wikimedia.org/T327300) [07:35:28] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage [07:38:45] (03CR) 10TrainBranchBot: [C:03+2] "Approved by wmde-fisch@deploy2003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308687 (https://phabricator.wikimedia.org/T430941) (owner: 10WMDE-Fisch) [07:40:10] (03Merged) 10jenkins-bot: Enable sub-references on more group2 wikis (batch2) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308687 (https://phabricator.wikimedia.org/T430941) (owner: 10WMDE-Fisch) [07:41:03] !log wmde-fisch@deploy2003 Started scap sync-world: Backport for [[gerrit:1308687|Enable sub-references on more group2 wikis (batch2) (T430941)]] [07:41:06] T430941: Global rollout - Sub-ref deployments to group 2 wikis (batch 2) - https://phabricator.wikimedia.org/T430941 [07:43:21] !log wmde-fisch@deploy2003 wmde-fisch: Backport for [[gerrit:1308687|Enable sub-references on more group2 wikis (batch2) (T430941)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [07:44:50] !log wmde-fisch@deploy2003 wmde-fisch: Continuing with deployment [07:49:39] !log wmde-fisch@deploy2003 Finished scap sync-world: Backport for [[gerrit:1308687|Enable sub-references on more group2 wikis (batch2) (T430941)]] (duration: 08m 36s) [07:49:42] T430941: Global rollout - Sub-ref deployments to group 2 wikis (batch 2) - https://phabricator.wikimedia.org/T430941 [07:51:23] I'm done! [07:54:06] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1008.eqiad.wmnet with OS trixie [07:56:16] !log ayounsi@cumin1003 START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B4 [07:57:56] (03PS2) 10JavierMonton: streams: webrequest - pageview - trending [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308656 (https://phabricator.wikimedia.org/T430675) [07:59:16] ayounsi@cumin1003 depool-rack (PID 3083788) is awaiting input [07:59:45] FIRING: SwiftLowContainerAvailability: Swift eqiad container availability low - https://wikitech.wikimedia.org/wiki/Swift/How_To - https://grafana.wikimedia.org/d/OPgmB1Eiz/swift?panelId=8&fullscreen&orgId=1&var-DC=eqiad - https://alerts.wikimedia.org/?q=alertname%3DSwiftLowContainerAvailability [08:00:05] thcipriani and thcipriani: #bothumor When your hammer is PHP, everything starts looking like a thumb. Rise for MediaWiki train - Utc morning Version. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260709T0800). [08:03:07] !log jmm@cumin2002 START - Cookbook sre.hosts.reboot-single for host cumin2003.codfw.wmnet [08:04:23] thcipriani and thcipriani hehe [08:04:38] the deployment page still has some breakage I did not fix yesterday (poke andre) [08:04:56] WMDE-Fisch: well done, thank you for having notified you are done with the patch deployment [08:05:13] I am going to look at the logs a bit then will promote the rest of wikis [08:08:00] !log ayounsi@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 27 hosts with reason: codfw rack B4 depool for maintenance [08:08:06] !log ayounsi@cumin1003 START - Cookbook sre.mysql.depool depool db2204: codfw rack B4 depool for maintenance [08:08:47] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2204: codfw rack B4 depool for maintenance [08:08:56] !log jmm@cumin2002 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2003.codfw.wmnet [08:08:57] !log ayounsi@cumin1003 START - Cookbook sre.mysql.depool depool db2205: codfw rack B4 depool for maintenance [08:09:45] RESOLVED: SwiftLowContainerAvailability: Swift eqiad container availability low - https://wikitech.wikimedia.org/wiki/Swift/How_To - https://grafana.wikimedia.org/d/OPgmB1Eiz/swift?panelId=8&fullscreen&orgId=1&var-DC=eqiad - https://alerts.wikimedia.org/?q=alertname%3DSwiftLowContainerAvailability [08:09:49] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2205: codfw rack B4 depool for maintenance [08:09:54] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host build2002.codfw.wmnet [08:10:28] !log ayounsi@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet [08:10:32] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host build2004.codfw.wmnet [08:12:37] PROBLEM - Check unit status of httpbb_kubernetes_mw-api-int_hourly on cumin2003 is CRITICAL: CRITICAL: Status of the systemd unit httpbb_kubernetes_mw-api-int_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [08:13:52] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet [08:13:52] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B4 [08:15:28] !log ayounsi@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b4-codfw,lsw1-b4-codfw IPv6,lsw1-b4-codfw.mgmt with reason: Switch maintenance [08:15:33] !log lsw1-b4-codfw> request system reboot - T430910 [08:15:35] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [08:15:36] T430910: codfw: rack B4 maintenance - https://phabricator.wikimedia.org/T430910 [08:15:57] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2002.codfw.wmnet [08:16:21] !log klausman@cumin1003 START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P{ml-serve1003.eqiad.wmnet} and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) [08:16:25] !log klausman@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet [08:16:26] !log klausman@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet [08:16:32] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2004.codfw.wmnet [08:16:46] !log klausman@cumin1003 START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm [08:17:13] PROBLEM - Host ml-staging-etcd2002 is DOWN: PING CRITICAL - Packet loss = 100% [08:17:49] PROBLEM - BFD status on ssw1-a8-codfw.mgmt is CRITICAL: Down: 1 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [08:17:49] PROBLEM - BFD status on ssw1-a1-codfw.mgmt is CRITICAL: Down: 1 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [08:18:39] FIRING: CoreBGPDown: Core BGP session down between ssw1-a8-codfw and lsw1-b4-codfw (10.192.252.13) - group EVPN_IBGP - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=ssw1-a8-codfw:9804&var-bgp_group=EVPN_IBGP&var-bgp_neighbor=lsw1-b4-codfw - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [08:19:51] FIRING: [2x] SwitchCoreInterfaceDown: Switch core interface down - ssw1-a1-codfw:et-0/0/10 (Core: lsw1-b4-codfw:et-0/0/55 {#230403800001}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Switch_interface_down - https://alerts.wikimedia.org/?q=alertname%3DSwitchCoreInterfaceDown [08:21:19] lets roll the train [08:21:51] (03PS1) 10TrainBranchBot: group2 to 1.47.0-wmf.10 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309092 (https://phabricator.wikimedia.org/T430829) [08:21:53] (03CR) 10TrainBranchBot: [C:03+2] "Initiated by hashar@deploy2003" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309092 (https://phabricator.wikimedia.org/T430829) (owner: 10TrainBranchBot) [08:22:31] FIRING: [2x] RedisReplicaDown: Redis replica down rdb2014:16378 redis_misc - https://wikitech.wikimedia.org/wiki/Redis#Cluster_redis_misc - https://alerts.wikimedia.org/?q=alertname%3DRedisReplicaDown [08:22:52] (03Merged) 10jenkins-bot: group2 to 1.47.0-wmf.10 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309092 (https://phabricator.wikimedia.org/T430829) (owner: 10TrainBranchBot) [08:23:39] FIRING: [2x] CoreBGPDown: Core BGP session down between ssw1-a1-codfw and lsw1-b4-codfw (10.192.252.13) - group EVPN_IBGP - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [08:24:25] PROBLEM - Host sretest2010 is DOWN: PING CRITICAL - Packet loss = 100% [08:24:25] PROBLEM - Host sretest2004 is DOWN: PING CRITICAL - Packet loss = 100% [08:24:25] PROBLEM - Host sretest2009 is DOWN: PING CRITICAL - Packet loss = 100% [08:24:25] PROBLEM - Host sretest2001 is DOWN: PING CRITICAL - Packet loss = 100% [08:24:25] PROBLEM - Host sretest2003 is DOWN: PING CRITICAL - Packet loss = 100% [08:24:29] PROBLEM - Host sretest2006 is DOWN: PING CRITICAL - Packet loss = 100% [08:24:55] RECOVERY - Host sretest2001 is UP: PING OK - Packet loss = 0%, RTA = 31.67 ms [08:24:55] RECOVERY - Host sretest2009 is UP: PING OK - Packet loss = 0%, RTA = 30.76 ms [08:25:08] ^ expected [08:25:53] RECOVERY - Host sretest2004 is UP: PING OK - Packet loss = 0%, RTA = 31.73 ms [08:25:55] RECOVERY - Host sretest2003 is UP: PING OK - Packet loss = 0%, RTA = 31.68 ms [08:25:57] RECOVERY - Host sretest2006 is UP: PING OK - Packet loss = 0%, RTA = 31.67 ms [08:26:09] ^ expected [08:26:37] !log failover Ganeti master in codfw to ganeti2048 T430928 [08:26:40] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [08:26:41] T430928: codfw: rack B7 maintenance - https://phabricator.wikimedia.org/T430928 [08:26:55] RECOVERY - Host sretest2010 is UP: PING OK - Packet loss = 0%, RTA = 31.67 ms [08:27:31] FIRING: [5x] RedisReplicaDown: Redis replica down rdb2014:16378 redis_misc - https://wikitech.wikimedia.org/wiki/Redis#Cluster_redis_misc - https://alerts.wikimedia.org/?q=alertname%3DRedisReplicaDown [08:27:51] RECOVERY - BFD status on ssw1-a1-codfw.mgmt is OK: UP: 17 AdminDown: 0 Down: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [08:28:37] RECOVERY - Host ml-staging-etcd2002 is UP: PING OK - Packet loss = 0%, RTA = 32.07 ms [08:28:39] FIRING: [2x] CoreBGPDown: Core BGP session down between ssw1-a1-codfw and lsw1-b4-codfw (10.192.252.13) - group EVPN_IBGP - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [08:28:51] RECOVERY - BFD status on ssw1-a8-codfw.mgmt is OK: UP: 17 AdminDown: 0 Down: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [08:28:55] PROBLEM - ganeti-wconfd running on ganeti2032 is CRITICAL: PROCS CRITICAL: 0 processes with UID = 110 (gnt-masterd), command name ganeti-wconfd https://wikitech.wikimedia.org/wiki/Ganeti [08:29:34] Production checks failed. [08:29:35] fun times [08:29:51] RESOLVED: [2x] SwitchCoreInterfaceDown: Switch core interface down - ssw1-a1-codfw:et-0/0/10 (Core: lsw1-b4-codfw:et-0/0/55 {#230403800001}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Switch_interface_down - https://alerts.wikimedia.org/?q=alertname%3DSwitchCoreInterfaceDown [08:30:13] 272 errors in the last errors. Top one being: 135 hits for "[136 hits] Error: Call to a member function getDifficulty() on null" [08:30:15] I'll check [08:30:51] and 135 hits PHP Warning: Undefined array key "copyedit" [08:31:15] FIRING: [2x] MediaWikiHighErrorRate: Elevated rate of MediaWiki errors - kube-mw-web - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiHighErrorRate [08:31:48] !log hashar@deploy2003 Rolling back deployment [08:32:05] !log ayounsi@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet [08:32:09] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet [08:32:26] I have confirmed the errors show up in the OpenSearch dashboard and rollback of deployment is ongoing [08:32:31] RESOLVED: [5x] RedisReplicaDown: Redis replica down rdb2014:16378 redis_misc - https://wikitech.wikimedia.org/wiki/Redis#Cluster_redis_misc - https://alerts.wikimedia.org/?q=alertname%3DRedisReplicaDown [08:33:01] I am not sure why the MediaWikiHighErrorRate triggered given the new code was only on canaries, but maybe the rate of errors coming from canaries was sufficient to trigger the limit [08:33:39] RESOLVED: [2x] CoreBGPDown: Core BGP session down between ssw1-a1-codfw and lsw1-b4-codfw (10.192.252.13) - group EVPN_IBGP - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [08:35:15] 10SRE-swift-storage, 06Commons, 10MediaWiki-File-management, 07Wikimedia-production-error: DBTransactionSizeError when trying to move file - https://phabricator.wikimedia.org/T431616#12103413 (10Aklapper) [08:35:31] !log hashar@deploy2003 rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.10 refs T430829 [08:35:34] T430829: 1.47.0-wmf.10 deployment blockers - https://phabricator.wikimedia.org/T430829 [08:36:13] !log klausman@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage [08:36:15] RESOLVED: [2x] MediaWikiHighErrorRate: Elevated rate of MediaWiki errors - kube-mw-web - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiHighErrorRate [08:36:20] FIRING: CirrusSearchFullTextLatencyTooHigh: CirrusSearch full_text 95th percentiles latency is too high (mw@codfw to dnsdisc) - https://wikitech.wikimedia.org/wiki/Search#Health/Activity_Monitoring - https://grafana.wikimedia.org/d/dc04b9f2-b8d5-4ab6-9482-5d9a75728951/elasticsearch-percentiles?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchFullTextLatencyTooHigh [08:36:26] FIRING: CirrusSearchMoreLikeLatencyTooHigh: CirrusSearch more_like 95th percentiles latency is too high (mw@codfw to dnsdisc) - https://wikitech.wikimedia.org/wiki/Search#Health/Activity_Monitoring - https://grafana.wikimedia.org/d/dc04b9f2-b8d5-4ab6-9482-5d9a75728951/elasticsearch-percentiles?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchMoreLikeLatencyTooHigh [08:37:41] !log ayounsi@cumin1003 START - Cookbook sre.mysql.pool pool db2204: codfw rack B4 repool after maintenance [08:37:50] (03CR) 10JavierMonton: streams: webrequest - pageview - trending (034 comments) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308656 (https://phabricator.wikimedia.org/T430675) (owner: 10JavierMonton) [08:38:16] group2 rolled back to wmf.9 [08:38:17] !log ayounsi@cumin1003 START - Cookbook sre.mysql.pool pool db2205: codfw rack B4 repool after maintenance [08:38:43] (03CR) 10Klausman: [C:03+1] "I know it's outsider perspective in a way, but LGTM!" [puppet] - 10https://gerrit.wikimedia.org/r/1308149 (https://phabricator.wikimedia.org/T406746) (owner: 10Btullis) [08:38:53] FIRING: DDoSDetected: FastNetMon has detected an attack on esams #page - https://bit.ly/wmf-fastnetmon - https://w.wiki/8oU - https://alerts.wikimedia.org/?q=alertname%3DDDoSDetected [08:38:59] 10SRE-swift-storage, 06Commons, 06DBA, 10MediaWiki-File-management, 07Wikimedia-production-error: DBTransactionSizeError when trying to move file - https://phabricator.wikimedia.org/T431616#12103431 (10MatthewVernon) [08:39:18] !log klausman@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage [08:40:23] James_F: Dreamy_Jazz: you guys file and investigate a train blocker task before I had even had a chance to look at the logs :] that is great. (re: https://phabricator.wikimedia.org/T431668 ) [08:40:51] (not me this time :D ) [08:41:08] the "PHP Warning: Undefined array key "copyedit"" one get filtered out in the opensearch dashboard for "mediawiki-NEW-errors", so I guess I will investigate the filters [08:41:20] RESOLVED: CirrusSearchFullTextLatencyTooHigh: CirrusSearch full_text 95th percentiles latency is too high (mw@codfw to dnsdisc) - https://wikitech.wikimedia.org/wiki/Search#Health/Activity_Monitoring - https://grafana.wikimedia.org/d/dc04b9f2-b8d5-4ab6-9482-5d9a75728951/elasticsearch-percentiles?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchFullTextLatencyTooHigh [08:41:25] cause last time I ran the train, some blocker error was left unnoticed by me due to a wrong ilter [08:41:26] RESOLVED: CirrusSearchMoreLikeLatencyTooHigh: CirrusSearch more_like 95th percentiles latency is too high (mw@codfw to dnsdisc) - https://wikitech.wikimedia.org/wiki/Search#Health/Activity_Monitoring - https://grafana.wikimedia.org/d/dc04b9f2-b8d5-4ab6-9482-5d9a75728951/elasticsearch-percentiles?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchMoreLikeLatencyTooHigh [08:42:02] (and because I did not learn yet about the /MediaWiki Error Logs/ tab in spiderpig ( https://spiderpig.wikimedia.org/mediawiki/logs [private] ) [08:43:53] RESOLVED: DDoSDetected: FastNetMon has detected an attack on esams #page - https://bit.ly/wmf-fastnetmon - https://w.wiki/8oU - https://alerts.wikimedia.org/?q=alertname%3DDDoSDetected [08:45:41] If someone from Growth could advise as to what the expected prod config is that'd be great. [08:46:02] urbanecm: Are you around? [08:46:40] James_F: yeah, what's up? [08:46:45] T431668 [08:46:45] T431668: German Wikipedia broken! Fatal exception of type "Error": Call to a member function getDifficulty() on null - https://phabricator.wikimedia.org/T431668 [08:46:52] Doesnt sound good... [08:46:54] Looking [08:46:56] urbanecm: revise-tone is live in CC but not config. [08:52:37] urbanecm: I've pushed a longer-term code fix to avoid this regressing, but we should probably change in-prod or on-wiki config? [08:54:32] (03CR) 10Federico Ceratto: sre.mysql: split pool/depool (031 comment) [cookbooks] - 10https://gerrit.wikimedia.org/r/1295480 (https://phabricator.wikimedia.org/T422361) (owner: 10Federico Ceratto) [08:54:39] the CC config is expected to have config even for tasks that are not enabled... and the feature is not expected at dewiki atm [08:54:59] 06SRE, 10SRE-Access-Requests: Grant shell access (fr-tech-devs) for laurabarluzzi - https://phabricator.wikimedia.org/T431338#12103506 (10cmooney) [08:55:10] !log klausman@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm [08:55:17] !log klausman@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet [08:55:19] !log klausman@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet [08:55:20] !log klausman@cumin1003 END (PASS) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=0) rolling reimage on P{ml-serve1003.eqiad.wmnet} and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) [08:55:51] FIRING: ATSBackendErrorsHigh: ATS: elevated 5xx errors from swift.discovery.wmnet in eqsin #page - https://wikitech.wikimedia.org/wiki/Apache_Traffic_Server#Debugging - https://grafana.wikimedia.org/d/1T_4O08Wk/ats-backends-origin-servers-overview?orgId=1&viewPanel=12&var-site=eqsin&var-cluster=upload&var-origin=swift.discovery.wmnet - https://alerts.wikimedia.org/?q=alertname%3DATSBackendErrorsHigh [08:57:32] !log klausman@cumin1003 START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm [08:57:48] !log klausman@cumin1003 START - Cookbook sre.hosts.move-vlan for host ml-serve1003 [08:58:32] !log klausman@cumin1003 START - Cookbook sre.dns.netbox [08:59:45] (03PS1) 10Mszwarc: Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups [core] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1309101 (https://phabricator.wikimedia.org/T431176) [09:00:34] James_F: your patch seems reasonable, I added a +2. I'll look for a more permanent fix in the meantime. [09:01:01] (but the fact we're just 2 engineers ATM isn't helping :/) [09:04:30] klausman@cumin1003 reimage (PID 3103779) is awaiting input [09:05:22] (03PS1) 10Urbanecm: NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback [extensions/GrowthExperiments] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1309102 (https://phabricator.wikimedia.org/T431668) [09:06:10] jouncebot: nowandnexr [09:06:12] jouncebot: nowandnext [09:06:12] For the next 0 hour(s) and 53 minute(s): MediaWiki train - Utc morning Version (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260709T0800) [09:06:12] In 0 hour(s) and 53 minute(s): MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260709T1000) [09:06:22] seems like using train window to unbreak train is a good idea [09:06:23] !log klausman@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" [09:06:26] (03CR) 10Urbanecm: [C:03+2] NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback [extensions/GrowthExperiments] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1309102 (https://phabricator.wikimedia.org/T431668) (owner: 10Urbanecm) [09:07:16] !log klausman@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003" [09:07:16] !log klausman@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [09:07:16] !log klausman@cumin1003 START - Cookbook sre.dns.wipe-cache ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [09:07:20] !log klausman@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [09:07:21] !log klausman@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 [09:08:08] !log klausman@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 [09:08:08] !log klausman@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1003 [09:08:48] (03CR) 10Brouberol: [C:03+1] Switch new an-test-master and an-test-coord nodes into their role [puppet] - 10https://gerrit.wikimedia.org/r/1308149 (https://phabricator.wikimedia.org/T406746) (owner: 10Btullis) [09:10:51] FIRING: [3x] ATSBackendErrorsHigh: ATS: elevated 5xx errors from swift.discovery.wmnet in eqiad #page - https://wikitech.wikimedia.org/wiki/Apache_Traffic_Server#Debugging - https://alerts.wikimedia.org/?q=alertname%3DATSBackendErrorsHigh [09:12:37] RECOVERY - Check unit status of httpbb_kubernetes_mw-api-int_hourly on cumin2003 is OK: OK: Status of the systemd unit httpbb_kubernetes_mw-api-int_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [09:12:45] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for rscout - https://phabricator.wikimedia.org/T431633#12103594 (10jcrespo) p:05Triage→03High a:03jcrespo [09:12:46] (03CR) 10Federico Ceratto: "Tested successfully on a core host and on pc1023, both pool and depool." [cookbooks] - 10https://gerrit.wikimedia.org/r/1295480 (https://phabricator.wikimedia.org/T422361) (owner: 10Federico Ceratto) [09:13:15] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for rscout - https://phabricator.wikimedia.org/T431633#12103597 (10jcrespo) [09:13:59] (03CR) 10TrainBranchBot: [C:03+2] "Approved by urbanecm@deploy2003 using scap backport" [extensions/GrowthExperiments] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1309102 (https://phabricator.wikimedia.org/T431668) (owner: 10Urbanecm) [09:14:21] (03CR) 10Muehlenhoff: [C:03+1] "Looks good!" [puppet] - 10https://gerrit.wikimedia.org/r/1308585 (https://phabricator.wikimedia.org/T431010) (owner: 10Cathal Mooney) [09:14:40] urbanecm: James_F: thx for the patch! [09:14:53] thanks goes to James :) i merely hit the big blue button [09:15:09] Ha. [09:15:11] TEAM WORK!!! [09:15:33] (03PS1) 10Arthur taylor: Upgrade to the newly-released version of the Queryservice GUI [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309103 (https://phabricator.wikimedia.org/T431080) [09:15:51] FIRING: [3x] ATSBackendErrorsHigh: ATS: elevated 5xx errors from swift.discovery.wmnet in eqiad #page - https://wikitech.wikimedia.org/wiki/Apache_Traffic_Server#Debugging - https://alerts.wikimedia.org/?q=alertname%3DATSBackendErrorsHigh [09:17:14] (03Merged) 10jenkins-bot: NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback [extensions/GrowthExperiments] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1309102 (https://phabricator.wikimedia.org/T431668) (owner: 10Urbanecm) [09:17:25] !log urbanecm@deploy2003 Started scap sync-world: Backport for [[gerrit:1309102|NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] [09:17:29] T431668: Fatal exception of type "Error": Call to a member function getDifficulty() on null - https://phabricator.wikimedia.org/T431668 [09:17:36] (03CR) 10Federico Ceratto: sre.mysql: split pool/depool (031 comment) [cookbooks] - 10https://gerrit.wikimedia.org/r/1295480 (https://phabricator.wikimedia.org/T422361) (owner: 10Federico Ceratto) [09:18:58] !log jmm@cumin2003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: reboot [09:19:04] !log urbanecm@deploy2003 urbanecm: Backport for [[gerrit:1309102|NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [09:19:36] !log urbanecm@deploy2003 urbanecm: Continuing with deployment [09:23:09] 06SRE, 06Infrastructure-Foundations, 10netops: Faulty port cr2-eqord:xe-0/1/1 - https://phabricator.wikimedia.org/T252988#12103639 (10cmooney) 05Open→03Resolved a:03cmooney Port is currently not in use, router seems healthy. [09:23:20] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2204: codfw rack B4 repool after maintenance [09:23:28] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host sretest1006.eqiad.wmnet [09:23:45] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2205: codfw rack B4 repool after maintenance [09:23:53] !log urbanecm@deploy2003 Finished scap sync-world: Backport for [[gerrit:1309102|NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] (duration: 06m 27s) [09:23:57] T431668: Fatal exception of type "Error": Call to a member function getDifficulty() on null - https://phabricator.wikimedia.org/T431668 [09:25:26] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host debmonitor2003.codfw.wmnet [09:25:55] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host debmonitor-dev2001.codfw.wmnet [09:26:36] (03CR) 10Btullis: [C:03+2] Switch new an-test-master and an-test-coord nodes into their role [puppet] - 10https://gerrit.wikimedia.org/r/1308149 (https://phabricator.wikimedia.org/T406746) (owner: 10Btullis) [09:27:09] James_F: hashar. backported, dewiki seems to be up [09:27:15] !log klausman@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage [09:27:18] and for some reason on wmf.10 already? [09:27:45] (03PS1) 10Btullis: opensearch-cluster: render nodePool affinity through tpl [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309107 (https://phabricator.wikimedia.org/T402408) [09:27:47] (03PS1) 10Btullis: opensearch: match pod anti-affinity to the actual cluster name [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309108 (https://phabricator.wikimedia.org/T402408) [09:27:58] hmmm [09:28:22] urbanecm: I definitely rolled back group2 to wmf.9 [09:28:38] but maybe you unexpectedly promoted them back as part of deploying the backport [09:28:46] i see `group2 to 1.47.0-wmf.10` by `SpiderPig jobrunner/apiserver` on deployhost [09:28:52] but...how was that created, i don't know [09:29:08] https://tools.wmflabs.org/versions/ shows all wikis are on wmf.10 [09:29:10] i definitely did not expect to promote anything [09:29:14] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest1006.eqiad.wmnet [09:29:19] * hashar nods [09:29:26] that is a glitch in spiderpig/scap I guess [09:29:26] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor2003.codfw.wmnet [09:29:37] I rolledback using the "rollback all stages" button in SpiderPig [09:29:55] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor-dev2001.codfw.wmnet [09:30:02] the positive note is that there are no errors happening anymore which confirms your patches fixed it [09:30:06] 06SRE, 06Infrastructure-Foundations: Integrate Bookworm 12.14 point update - https://phabricator.wikimedia.org/T426759#12103687 (10MoritzMuehlenhoff) [09:30:19] so phabricator.wikimedia.org/T431668 can be marked resolved / VERIFIED :] [09:30:28] I'll file a task for my team to investigate [09:30:37] it is entirely I screwed up something while clicking buttons [09:30:51] RESOLVED: [2x] ATSBackendErrorsHigh: ATS: elevated 5xx errors from swift.discovery.wmnet in eqsin #page - https://wikitech.wikimedia.org/wiki/Apache_Traffic_Server#Debugging - https://alerts.wikimedia.org/?q=alertname%3DATSBackendErrorsHigh [09:31:15] logged T431677 for my team to ensure nothing else needs to be done (I'm not fully up to speed to revise tone expectations, so i might not see everything) [09:31:16] T431677: Decide whether any follow up is needed regarding Revise Tone UBN - https://phabricator.wikimedia.org/T431677 [09:31:58] (03CR) 10Brouberol: [C:03+1] Puppet 8: Replace legacy facts in role kafka [puppet] - 10https://gerrit.wikimedia.org/r/1308759 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [09:32:42] (03CR) 10Brouberol: [C:03+1] Puppet 8: Replace legacy facts in role druid [puppet] - 10https://gerrit.wikimedia.org/r/1308756 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [09:33:02] (03CR) 10Brouberol: [C:03+1] Puppet 8: Replace legacy facts in role analytics_test_cluster [puppet] - 10https://gerrit.wikimedia.org/r/1308755 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [09:33:05] !log klausman@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage [09:33:16] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host debmonitor1003.eqiad.wmnet [09:33:20] (03CR) 10Brouberol: [C:03+1] Puppet 8: Replace legacy facts in profile zookeeper [puppet] - 10https://gerrit.wikimedia.org/r/1308753 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [09:33:24] anyway, this seems to be done-ish [09:33:25] * urbanecm hides [09:33:36] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for rscout - https://phabricator.wikimedia.org/T431633#12103736 (10jcrespo) [09:33:47] (03CR) 10Brouberol: [C:03+1] Puppet 8: Replace legacy facts in role analytics_cluster [puppet] - 10https://gerrit.wikimedia.org/r/1308754 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [09:34:01] (03CR) 10Brouberol: [C:03+1] Puppet 8: Replace legacy facts in profile presto [puppet] - 10https://gerrit.wikimedia.org/r/1308751 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [09:34:20] (03CR) 10Brouberol: [C:03+1] Puppet 8: Replace legacy facts in profile query_service [puppet] - 10https://gerrit.wikimedia.org/r/1308752 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [09:34:45] (03CR) 10Brouberol: [C:03+1] Puppet 8: Replace legacy facts in profile hadoop [puppet] - 10https://gerrit.wikimedia.org/r/1308761 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [09:34:55] !log jmm@cumin2003 START - Cookbook sre.hosts.decommission for hosts testvm2007.codfw.wmnet [09:35:07] (03CR) 10Brouberol: [C:03+1] Puppet 8: Replace legacy facts in profile kafkatee [puppet] - 10https://gerrit.wikimedia.org/r/1308749 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [09:35:21] (03CR) 10Brouberol: [C:03+1] Puppet 8: Replace legacy facts in profile elasticsearch [puppet] - 10https://gerrit.wikimedia.org/r/1308732 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [09:35:34] (03CR) 10Brouberol: [C:03+1] Puppet 8: Replace legacy facts in profile hive [puppet] - 10https://gerrit.wikimedia.org/r/1308747 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [09:35:57] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host puppetboard2003.codfw.wmnet [09:36:23] (03CR) 10Brouberol: [C:03+1] Puppet 8: Replace legacy facts in profile kafka [puppet] - 10https://gerrit.wikimedia.org/r/1308748 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [09:36:25] 06SRE, 06Infrastructure-Foundations, 10netops: Cookbook sre.network.configure-switch-interfaces failing on upgraded Juniper switch - https://phabricator.wikimedia.org/T428071#12103739 (10ayounsi) 05Open→03Resolved [09:36:56] (03PS1) 10Majavah: hieradata: cloudinfra: Make cloudinfra-db-5 the primary one [puppet] - 10https://gerrit.wikimedia.org/r/1309109 (https://phabricator.wikimedia.org/T402005) [09:37:04] (03CR) 10Brouberol: [C:03+1] Puppet 8: Replace legacy facts in profile druid [puppet] - 10https://gerrit.wikimedia.org/r/1308730 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [09:37:11] (03CR) 10Brouberol: [C:03+1] Puppet 8: Replace legacy facts in module zookeeper [puppet] - 10https://gerrit.wikimedia.org/r/1308724 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [09:37:14] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor1003.eqiad.wmnet [09:37:21] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host puppetboard1003.eqiad.wmnet [09:37:24] (03CR) 10Brouberol: [C:03+1] Puppet 8: Replace legacy facts in profile analytics [puppet] - 10https://gerrit.wikimedia.org/r/1308725 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [09:37:46] 06SRE, 06MediaWiki-Media-Platform-Team, 06ServiceOps new, 10ServiceOps-Services-Oids, 10Thumbor: Thumbor-k8s performance improvements - https://phabricator.wikimedia.org/T333445#12103748 (10JTweed-WMF) p:05Triage→03High [09:37:53] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for rscout - https://phabricator.wikimedia.org/T431633#12103749 (10jcrespo) [09:38:01] (03CR) 10Brouberol: [C:03+1] Puppet 8: Replace legacy facts in module statistics [puppet] - 10https://gerrit.wikimedia.org/r/1308722 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [09:38:03] (03PS2) 10Btullis: opensearch-cluster: render nodePool affinity through tpl [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309107 (https://phabricator.wikimedia.org/T402408) [09:38:04] (03PS2) 10Btullis: opensearch: match pod anti-affinity to the actual cluster name [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309108 (https://phabricator.wikimedia.org/T402408) [09:38:08] (03CR) 10Brouberol: [C:03+1] Puppet 8: Replace legacy facts in module kafkatee [puppet] - 10https://gerrit.wikimedia.org/r/1308720 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [09:38:15] 10ops-eqiad, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 10netops: eqiad row A/B switch upgrade - https://phabricator.wikimedia.org/T418012#12103752 (10cmooney) p:05Medium→03High [09:38:24] (03CR) 10Brouberol: [C:03+1] Puppet 8: Replace legacy facts in module elasticsearch [puppet] - 10https://gerrit.wikimedia.org/r/1308719 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [09:38:38] (03CR) 10Brouberol: [C:03+1] Puppet 8: Replace legacy facts in module confluent [puppet] - 10https://gerrit.wikimedia.org/r/1308718 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [09:38:48] (03CR) 10Majavah: [C:03+2] hieradata: cloudinfra: Make cloudinfra-db-5 the primary one [puppet] - 10https://gerrit.wikimedia.org/r/1309109 (https://phabricator.wikimedia.org/T402005) (owner: 10Majavah) [09:38:53] (03CR) 10Brouberol: [C:03+1] Puppet 8: Replace legacy facts in profile bigtop [puppet] - 10https://gerrit.wikimedia.org/r/1308727 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [09:38:53] 10ops-eqiad, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 10netops: eqiad row A/B switch refresh prep - https://phabricator.wikimedia.org/T418012#12103758 (10cmooney) [09:39:05] (03CR) 10Brouberol: [C:03+1] Puppet 8: Replace legacy facts in module airflow [puppet] - 10https://gerrit.wikimedia.org/r/1308716 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [09:39:20] (03CR) 10Brouberol: [C:03+1] Puppet 8: Replace legacy facts in module bigtop [puppet] - 10https://gerrit.wikimedia.org/r/1308717 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [09:39:33] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for rscout - https://phabricator.wikimedia.org/T431633#12103765 (10jcrespo) I just realized this is an extension of existing production access, which was granted at T430594, so I could have sped up part of the process. [09:39:37] !log jmm@cumin2003 START - Cookbook sre.dns.netbox [09:39:48] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard2003.codfw.wmnet [09:39:56] (03CR) 10Brouberol: [C:03+1] Puppet 8: Replace legacy facts in role wdqs [puppet] - 10https://gerrit.wikimedia.org/r/1308757 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [09:41:11] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard1003.eqiad.wmnet [09:41:43] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for rscout - https://phabricator.wikimedia.org/T431633#12103771 (10jcrespo) [09:41:59] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for rscout - https://phabricator.wikimedia.org/T431633#12103772 (10jcrespo) //Explicit approval is not required for WMF or WMDE Staff.// [09:43:13] (03CR) 10Blake: "recheck" [alerts] - 10https://gerrit.wikimedia.org/r/1307787 (https://phabricator.wikimedia.org/T430290) (owner: 10Blake) [09:43:41] 10SRE-swift-storage, 06MediaWiki-Media-Platform-Team, 10MediaWiki-Uploading, 07Wikimedia-production-error: MediaWiki\Upload\Exception\UploadChunkFileException: Error storing file in '{chunkPath}': backend-fail-internal; local-swift-codfw - https://phabricator.wikimedia.org/T430986#12103790 (10JTweed-WMF) p... [09:43:58] 10SRE-swift-storage, 06MediaWiki-Media-Platform-Team, 10MediaWiki-Uploading, 07Wikimedia-production-error: MediaWiki\Upload\Exception\UploadChunkFileException: Error storing file in '{chunkPath}': backend-fail-internal; local-swift-codfw - https://phabricator.wikimedia.org/T430986#12103791 (10JTweed-WMF) p... [09:45:02] (03CR) 10CWilliams: sre.mysql: split pool/depool (032 comments) [cookbooks] - 10https://gerrit.wikimedia.org/r/1295480 (https://phabricator.wikimedia.org/T422361) (owner: 10Federico Ceratto) [09:45:23] 10ops-eqsin, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 10netops: EQSIN:New switch setup/configuration - https://phabricator.wikimedia.org/T418439#12103797 (10cmooney) p:05Low→03Medium [09:45:37] jmm@cumin2003 decommission (PID 21434) is awaiting input [09:48:35] (03PS1) 10Majavah: hieradata: cloudinfra: Replace cloudinfra-db03 with cloudinfra-db-6 [puppet] - 10https://gerrit.wikimedia.org/r/1309110 (https://phabricator.wikimedia.org/T402005) [09:49:01] !log jmm@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" [09:49:20] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet [09:49:29] !log jmm@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" [09:49:30] !log jmm@cumin2003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [09:49:32] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2007.codfw.wmnet [09:49:43] !log jmm@cumin2003 START - Cookbook sre.hosts.decommission for hosts testvm2008.wikimedia.org [09:49:49] !log elukey@cumin1003 START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet [09:50:00] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet [09:50:02] !log klausman@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm [09:50:16] !log klausman@cumin1003 START - Cookbook sre.hosts.reimage for host ml-serve1004.eqiad.wmnet with OS bookworm [09:50:28] (03PS1) 10Jcrespo: admin: Add rscout to analytics-privatedata-users group [puppet] - 10https://gerrit.wikimedia.org/r/1309111 (https://phabricator.wikimedia.org/T431633) [09:50:44] !log klausman@cumin1003 START - Cookbook sre.hosts.move-vlan for host ml-serve1004 [09:50:50] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host install7002.wikimedia.org [09:50:58] (03CR) 10Majavah: [C:03+2] hieradata: cloudinfra: Replace cloudinfra-db03 with cloudinfra-db-6 [puppet] - 10https://gerrit.wikimedia.org/r/1309110 (https://phabricator.wikimedia.org/T402005) (owner: 10Majavah) [09:52:00] (03PS1) 10Jforrester: abstractwiki: Don't run our jobs too quickly, they'll collide and error [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309112 (https://phabricator.wikimedia.org/T430898) [09:52:07] !log klausman@cumin1003 START - Cookbook sre.dns.netbox [09:52:23] 06SRE, 06Infrastructure-Foundations, 10netops: OSPF metrics - https://phabricator.wikimedia.org/T200277#12103837 (10ayounsi) 05Open→03Declined Not relevant soon with the new transport links. [09:52:35] 06SRE, 06Infrastructure-Foundations, 10netops: Cleanup confed BGP peerings and policies - https://phabricator.wikimedia.org/T167841#12103839 (10cmooney) 05Open→03Declined This work is going to be done, but we are tracking the overall WAN configuration in task T297355 we don't need this separate one. [09:52:50] !log elukey@cumin1003 START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet [09:52:51] !log elukey@cumin1003 END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet [09:55:25] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet [09:55:38] !log elukey@cumin1003 START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet [09:55:50] !log elukey@cumin1003 END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet [09:55:56] (03PS2) 10Jcrespo: admin: Add rscout to analytics-privatedata-users group [puppet] - 10https://gerrit.wikimedia.org/r/1309111 (https://phabricator.wikimedia.org/T431633) [09:55:58] (03CR) 10Jcrespo: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1309111 (https://phabricator.wikimedia.org/T431633) (owner: 10Jcrespo) [09:56:30] !log klausman@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" [09:57:06] !log klausman@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003" [09:57:06] !log klausman@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [09:57:06] !log klausman@cumin1003 START - Cookbook sre.dns.wipe-cache ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [09:57:10] !log klausman@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [09:57:10] !log klausman@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1004 [09:57:18] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install7002.wikimedia.org [09:57:28] !log jmm@cumin2003 START - Cookbook sre.dns.netbox [09:57:48] !log klausman@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1004 [09:57:48] !log klausman@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1004 [09:58:26] 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to analytics-privatedata-users for rscout - https://phabricator.wikimedia.org/T431633#12103873 (10jcrespo) [09:58:30] (03CR) 10Genoveva Galarza: [C:03+1] "LGTM" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309112 (https://phabricator.wikimedia.org/T430898) (owner: 10Jforrester) [09:59:05] jouncebot: nowandnext [09:59:05] For the next 0 hour(s) and 0 minute(s): MediaWiki train - Utc morning Version (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260709T0800) [09:59:05] In 0 hour(s) and 0 minute(s): MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260709T1000) [09:59:25] Bah, let's not trip over SRE's window. :-) [10:00:05] Deploy window MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260709T1000) [10:00:32] !log jmm@cumin2003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [10:00:34] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2008.wikimedia.org [10:01:16] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, 06cloud-services-team (FY2025/2026-Q3-Q4): Rebalance cloudvirts out of E4 - https://phabricator.wikimedia.org/T431682 (10fgiunchedi) 03NEW [10:03:05] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, 06cloud-services-team (FY2025/2026-Q3-Q4): Ensure cloudvirt capacity is more evenly spread out among racks - https://phabricator.wikimedia.org/T424658#12103917 (10fgiunchedi) >>! In T424658#12101239, @VRiley-WMF wrote: > @fgiunchedi I was curious to know if ther... [10:05:07] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, 06cloud-services-team (FY2025/2026-Q3-Q4): Rebalance cloudvirts out of E4 - https://phabricator.wikimedia.org/T431682#12103927 (10fgiunchedi) [10:05:10] FIRING: [2x] SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [10:06:05] (03PS3) 10Btullis: opensearch-cluster: render nodePool affinity through tpl [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309107 (https://phabricator.wikimedia.org/T402408) [10:06:06] (03PS3) 10Btullis: opensearch: match pod anti-affinity to the actual cluster name [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309108 (https://phabricator.wikimedia.org/T402408) [10:11:33] (03CR) 10Jcrespo: [C:03+1] "LGTM, let me know if you are going to merge or prefer me to apply it." [puppet] - 10https://gerrit.wikimedia.org/r/1308585 (https://phabricator.wikimedia.org/T431010) (owner: 10Cathal Mooney) [10:11:47] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Thursday, July 09 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployca" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309112 (https://phabricator.wikimedia.org/T430898) (owner: 10Jforrester) [10:12:41] 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to cloudcontrol servers for tlepage - https://phabricator.wikimedia.org/T431010#12103958 (10jcrespo) [10:14:10] !log elukey@cumin1003 START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet [10:14:26] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet [10:15:04] !log klausman@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage [10:15:59] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host install6003.wikimedia.org [10:16:53] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host install5004.wikimedia.org [10:17:26] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install6003.wikimedia.org [10:17:34] 10SRE-Access-Requests, 10LDAP-Access-Requests: Superset data access request for abibendall - https://phabricator.wikimedia.org/T430938#12104004 (10jcrespo) @ABendall-WMF Friendly ping to enquire if you followed the instructions above that my colleague suggested to request the wmf ldap group; or in case you fou... [10:17:58] !log jmm@cumin2003 START - Cookbook sre.ganeti.addnode for new host ganeti2031.codfw.wmnet to cluster codfw and group B [10:18:09] (03CR) 10Jforrester: "Note: Ia5e6f0a6 isn't pushed to gerrit yet." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308788 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [10:18:29] !log klausman@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003 [10:18:38] !log klausman@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003 [10:18:39] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install5004.wikimedia.org [10:18:39] (03CR) 10Federico Ceratto: sre.mysql: split pool/depool (032 comments) [cookbooks] - 10https://gerrit.wikimedia.org/r/1295480 (https://phabricator.wikimedia.org/T422361) (owner: 10Federico Ceratto) [10:18:50] !log klausman@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage [10:19:23] !log readded ganeti2031 to the codfw Ganeti cluster T430910 [10:19:26] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [10:19:26] T430910: codfw: rack B4 maintenance - https://phabricator.wikimedia.org/T430910 [10:19:39] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.addnode (exit_code=0) for new host ganeti2031.codfw.wmnet to cluster codfw and group B [10:21:13] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Thursday, July 09 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployca" [core] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1309101 (https://phabricator.wikimedia.org/T431176) (owner: 10Mszwarc) [10:21:19] 06SRE, 06Infrastructure-Foundations, 10netops: Create alerting for saturation on sub-rated interfaces - https://phabricator.wikimedia.org/T374614#12104022 (10ayounsi) a:03ayounsi [10:21:42] !log failover Ganeti master in codfw/routed to ganeti2034 T430928 [10:21:44] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [10:21:45] T430928: codfw: rack B7 maintenance - https://phabricator.wikimedia.org/T430928 [10:22:37] (03PS1) 10Elukey: redfish: skip setting AccountTypes if the admin user doesn't expose it [software/spicerack] - 10https://gerrit.wikimedia.org/r/1309120 (https://phabricator.wikimedia.org/T426180) [10:22:50] (03PS1) 10Btullis: kubernetes: run lvmd as a systemd service on selected dse-k8s workers [puppet] - 10https://gerrit.wikimedia.org/r/1309121 (https://phabricator.wikimedia.org/T429325) [10:23:51] PROBLEM - ganeti-wconfd running on ganeti2033 is CRITICAL: PROCS CRITICAL: 0 processes with UID = 111 (gnt-masterd), command name ganeti-wconfd https://wikitech.wikimedia.org/wiki/Ganeti [10:23:59] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet [10:24:12] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host install4004.wikimedia.org [10:24:17] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host install3004.wikimedia.org [10:24:30] 06SRE, 10SRE-Access-Requests: Grant shell access (fr-tech-devs) for laurabarluzzi - https://phabricator.wikimedia.org/T431338#12104034 (10jcrespo) Taking over access request, summary before this can be deployed: the blockers on this ticket to move forward: * Formal ok from direct manager to aprove production... [10:24:50] 10SRE-Access-Requests: Grant shell access (fr-tech-devs) for laurabarluzzi - https://phabricator.wikimedia.org/T431338#12104036 (10jcrespo) [10:25:26] (03PS1) 10Mszwarc: SuggestedInvestigations: Instrument link clicks in the cases table [extensions/CheckUser] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1309122 (https://phabricator.wikimedia.org/T429320) [10:25:47] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Thursday, July 09 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployca" [extensions/CheckUser] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1309122 (https://phabricator.wikimedia.org/T429320) (owner: 10Mszwarc) [10:25:53] (03CR) 10Jcrespo: [C:03+1] "Waiting for an SRE review to proceed." [puppet] - 10https://gerrit.wikimedia.org/r/1309111 (https://phabricator.wikimedia.org/T431633) (owner: 10Jcrespo) [10:27:00] jmm@cumin2003 drain-node (PID 32125) is awaiting input [10:28:04] 06SRE, 06Infrastructure-Foundations, 10netops, 13Patch-For-Review: Netbox network report failing - timeout errors - https://phabricator.wikimedia.org/T321704#12104049 (10cmooney) 05Open→03Resolved Closing this task, the issue was not in the test_matching_vlan() function. Thankfully right now we do... [10:29:56] !log aokoth@cumin1003 START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - T431656 [10:30:05] 06SRE, 06Infrastructure-Foundations, 10netops, 13Patch-Needs-Improvement: test_matching_vlan() function crashing in Netbox network report - https://phabricator.wikimedia.org/T339133#12104051 (10cmooney) 05Open→03Resolved a:03cmooney [10:30:43] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install3004.wikimedia.org [10:30:48] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install4004.wikimedia.org [10:30:52] 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to cloudcontrol servers for tlepage - https://phabricator.wikimedia.org/T431010#12104053 (10jcrespo) 05Open→03In progress p:05Triage→03High [10:31:02] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host install2005.wikimedia.org [10:31:06] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host install1005.wikimedia.org [10:32:43] (03PS2) 10Btullis: kubernetes: run lvmd as a systemd service on selected dse-k8s workers [puppet] - 10https://gerrit.wikimedia.org/r/1309121 (https://phabricator.wikimedia.org/T429325) [10:35:21] !log klausman@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1004.eqiad.wmnet with OS bookworm [10:37:30] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install1005.wikimedia.org [10:37:35] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install2005.wikimedia.org [10:38:45] 06SRE, 06Infrastructure-Foundations, 10netops: Nokia: how to approach schema differences in SR-Linux versions - https://phabricator.wikimedia.org/T412157#12104102 (10cmooney) 05Open→03Resolved a:03cmooney Gonna close this for now. We've tackled it by just having different files that get added for... [10:38:53] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2033.codfw.wmnet [10:39:00] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.5 point update - https://phabricator.wikimedia.org/T427072#12104109 (10MoritzMuehlenhoff) [10:39:33] !log aokoth@cumin1003 END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - T431656 [10:39:42] (03PS3) 10Btullis: kubernetes: run lvmd as a systemd service on selected dse-k8s workers [puppet] - 10https://gerrit.wikimedia.org/r/1309121 (https://phabricator.wikimedia.org/T429325) [10:40:07] !log cmooney@cumin1003 START - Cookbook sre.network.host-bgp for host ml-serve1003 [10:40:18] !log cmooney@cumin1003 END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1003 [10:41:57] !log aokoth@cumin1003 START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - T431656 [10:47:23] (03CR) 10Cathal Mooney: [C:03+2] Add shell access to group 'wmcs-roots' for tlepage [puppet] - 10https://gerrit.wikimedia.org/r/1308585 (https://phabricator.wikimedia.org/T431010) (owner: 10Cathal Mooney) [10:47:50] (03PS4) 10Btullis: kubernetes: run lvmd as a systemd service on selected dse-k8s workers [puppet] - 10https://gerrit.wikimedia.org/r/1309121 (https://phabricator.wikimedia.org/T429325) [10:48:48] (03CR) 10Muehlenhoff: [C:03+1] "Looks good" [puppet] - 10https://gerrit.wikimedia.org/r/1308695 (https://phabricator.wikimedia.org/T424055) (owner: 10AOkoth) [10:50:31] !log jmm@cumin2003 START - Cookbook sre.hosts.decommission for hosts testvm2005.codfw.wmnet [10:50:46] (03PS5) 10Btullis: kubernetes: run lvmd as a systemd service on selected dse-k8s workers [puppet] - 10https://gerrit.wikimedia.org/r/1309121 (https://phabricator.wikimedia.org/T429325) [10:51:04] (03CR) 10Btullis: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1309121 (https://phabricator.wikimedia.org/T429325) (owner: 10Btullis) [10:51:38] !log aokoth@cumin1003 END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - T431656 [10:51:42] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org [10:52:48] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.4 point update - https://phabricator.wikimedia.org/T420240#12104191 (10MoritzMuehlenhoff) [10:53:33] (03PS1) 10Giuseppe Lavagetto: haproxy: modularize hiddenparma support [puppet] - 10https://gerrit.wikimedia.org/r/1309128 (https://phabricator.wikimedia.org/T422235) [10:55:13] !log jmm@cumin2003 START - Cookbook sre.dns.netbox [10:57:28] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org [10:57:43] FIRING: KeyholderUnarmed: 1 unarmed Keyholder key(s) on netmon2002:9100 - https://wikitech.wikimedia.org/wiki/Keyholder - TODO - https://alerts.wikimedia.org/?q=alertname%3DKeyholderUnarmed [10:57:47] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host netmon1003.wikimedia.org [10:58:13] (03PS13) 10Federico Ceratto: cookbooks/sre/mysql/decommission: add cookbook [cookbooks] - 10https://gerrit.wikimedia.org/r/1291952 (https://phabricator.wikimedia.org/T426613) [10:58:17] (03PS2) 10Giuseppe Lavagetto: haproxy: modularize hiddenparma support [puppet] - 10https://gerrit.wikimedia.org/r/1309128 (https://phabricator.wikimedia.org/T422235) [10:58:35] (03CR) 10Federico Ceratto: cookbooks/sre/mysql/decommission: add cookbook (033 comments) [cookbooks] - 10https://gerrit.wikimedia.org/r/1291952 (https://phabricator.wikimedia.org/T426613) (owner: 10Federico Ceratto) [10:59:22] FIRING: JobUnavailable: Reduced availability for job atlas_exporter in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [10:59:44] !log jmm@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" [11:00:49] !log jmm@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" [11:00:49] !log jmm@cumin2003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [11:00:51] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2005.codfw.wmnet [11:02:16] 10SRE-Access-Requests: Grant shell access (fr-tech-devs) for laurabarluzzi - https://phabricator.wikimedia.org/T431338#12104220 (10cmooney) To record here Laura confirmed her ssh key to me over Slack and it matches the one in the task description. So this should be good to go once the approvers respond. [11:03:03] (03PS1) 10Btullis: Disable systemd timers for refine jobs on an-test-coord1002 [puppet] - 10https://gerrit.wikimedia.org/r/1309130 (https://phabricator.wikimedia.org/T406746) [11:03:28] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon1003.wikimedia.org [11:03:47] 10SRE-Access-Requests, 10LDAP-Access-Requests: Superset data access request for abibendall - https://phabricator.wikimedia.org/T430938#12104226 (10ABendall-WMF) Hello, Thanks for getting in touch. Yes I followed the link your colleague sent me and the subsequent page instructions - I think I followed the corr... [11:04:00] (03CR) 10Btullis: [C:03+2] Disable systemd timers for refine jobs on an-test-coord1002 [puppet] - 10https://gerrit.wikimedia.org/r/1309130 (https://phabricator.wikimedia.org/T406746) (owner: 10Btullis) [11:04:14] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ldap-rw1001.wikimedia.org [11:04:22] FIRING: [2x] JobUnavailable: Reduced availability for job atlas_exporter in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [11:04:23] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ldap-rw2001.wikimedia.org [11:06:59] (03CR) 10Hashar: "I explain it in the commit message:" [puppet] - 10https://gerrit.wikimedia.org/r/1308659 (https://phabricator.wikimedia.org/T431582) (owner: 10Hashar) [11:07:18] (03CR) 10Muehlenhoff: "You also need "krb: present" for the user entry" [puppet] - 10https://gerrit.wikimedia.org/r/1309111 (https://phabricator.wikimedia.org/T431633) (owner: 10Jcrespo) [11:08:08] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw1001.wikimedia.org [11:08:16] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw2001.wikimedia.org [11:08:59] (03PS6) 10Btullis: kubernetes: run lvmd as a systemd service on selected dse-k8s workers [puppet] - 10https://gerrit.wikimedia.org/r/1309121 (https://phabricator.wikimedia.org/T429325) [11:09:15] (03CR) 10Btullis: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1309121 (https://phabricator.wikimedia.org/T429325) (owner: 10Btullis) [11:09:22] RESOLVED: [2x] JobUnavailable: Reduced availability for job atlas_exporter in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [11:09:56] (03CR) 10Jcrespo: [C:04-1] "Good catch, thank you and sorry." [puppet] - 10https://gerrit.wikimedia.org/r/1309111 (https://phabricator.wikimedia.org/T431633) (owner: 10Jcrespo) [11:12:04] (03CR) 10JMeybohm: "recheck" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308802 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [11:12:14] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ldap-maint1001.eqiad.wmnet [11:12:23] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ldap-maint2001.codfw.wmnet [11:12:55] (03PS3) 10Jcrespo: admin: Add rscout to analytics-privatedata-users group [puppet] - 10https://gerrit.wikimedia.org/r/1309111 (https://phabricator.wikimedia.org/T431633) [11:13:33] (03CR) 10Jcrespo: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1309111 (https://phabricator.wikimedia.org/T431633) (owner: 10Jcrespo) [11:14:38] (03CR) 10JMeybohm: [C:04-1] "This requires a chart version bump" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306165 (https://phabricator.wikimedia.org/T427400) (owner: 10Jelto) [11:15:07] (03CR) 10JMeybohm: "LGTM, but we need to make sure to pin the calico version fleet wide before merging this!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306307 (https://phabricator.wikimedia.org/T427400) (owner: 10Jelto) [11:15:23] (03CR) 10Muehlenhoff: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1309111 (https://phabricator.wikimedia.org/T431633) (owner: 10Jcrespo) [11:16:03] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint1001.eqiad.wmnet [11:16:13] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint2001.codfw.wmnet [11:16:49] (03CR) 10JMeybohm: [C:03+1] bgp: Remove port from instance and remote_instance names. [alerts] - 10https://gerrit.wikimedia.org/r/1307787 (https://phabricator.wikimedia.org/T430290) (owner: 10Blake) [11:19:53] 10SRE-Access-Requests, 10LDAP-Access-Requests: Superset data access request for abibendall - https://phabricator.wikimedia.org/T430938#12104340 (10jcrespo) >>! In T430938#12104226, @ABendall-WMF wrote: > Hello, > > Thanks for getting in touch. Yes I followed the link your colleague sent me > and the subsequen... [11:21:41] (03CR) 10Blake: [C:03+2] bgp: Remove port from instance and remote_instance names. [alerts] - 10https://gerrit.wikimedia.org/r/1307787 (https://phabricator.wikimedia.org/T430290) (owner: 10Blake) [11:22:20] 10SRE-Access-Requests, 10LDAP-Access-Requests: Superset data access request for abibendall - https://phabricator.wikimedia.org/T430938#12104355 (10ABendall-WMF) Amazing thank you so much! And please just let me know if I've missed something or need to do anything else. Thanks again! [11:23:27] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host irc2003.wikimedia.org [11:24:11] (03Merged) 10jenkins-bot: bgp: Remove port from instance and remote_instance names. [alerts] - 10https://gerrit.wikimedia.org/r/1307787 (https://phabricator.wikimedia.org/T430290) (owner: 10Blake) [11:24:52] (03CR) 10Jcrespo: [C:03+2] "Thank you a lot for the quick review and correction!" [puppet] - 10https://gerrit.wikimedia.org/r/1309111 (https://phabricator.wikimedia.org/T431633) (owner: 10Jcrespo) [11:25:55] (03PS1) 10Bartosz Wójtowicz: ml-services: Update dtype to auto for qwen36-27b. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309133 (https://phabricator.wikimedia.org/T425680) [11:25:58] (03CR) 10AOkoth: [C:03+2] fpm: add templates for version 8.4 [puppet] - 10https://gerrit.wikimedia.org/r/1308695 (https://phabricator.wikimedia.org/T424055) (owner: 10AOkoth) [11:27:03] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc2003.wikimedia.org [11:27:44] 10SRE-Access-Requests, 10LDAP-Access-Requests: Superset data access request for abibendall - https://phabricator.wikimedia.org/T430938#12104380 (10jcrespo) 05Open→03In progress p:05Triage→03High [11:27:52] 10SRE-Access-Requests, 10LDAP-Access-Requests: Superset data access request for abibendall - https://phabricator.wikimedia.org/T430938#12104382 (10jcrespo) a:03jcrespo [11:30:25] (03CR) 10Clément Goubert: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308604 (https://phabricator.wikimedia.org/T431108) (owner: 10Clément Goubert) [11:31:27] 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to cloudcontrol servers for tlepage - https://phabricator.wikimedia.org/T431010#12104420 (10jcrespo) 05In progress→03Resolved a:03cmooney @TLepage-WMF Your access has been granted and deployed, you can follow the section https://wikitech.wi... [11:33:22] Gerrit seems down for me [11:33:31] confirm [11:33:45] back for me [11:34:21] Back for me too, but loading very slowly [11:34:59] Gone again [11:35:00] FIRING: [2x] NodeBGPSessionStatusNotEstablished: Kubernetes node ml-serve1003:0 has a BGP session which is not in the 'established' state. - https://wikitech.wikimedia.org/wiki/Kubernetes/Administration#NodeBGPSessionStatusNotEstablished - https://alerts.wikimedia.org/?q=alertname%3DNodeBGPSessionStatusNotEstablished [11:35:57] (03PS1) 10Hashar: docker_registry: allow push from new contint hosts [puppet] - 10https://gerrit.wikimedia.org/r/1309135 (https://phabricator.wikimedia.org/T418521) [11:36:45] 06SRE, 10Maps, 07affects-Kiwix-and-openZIM, 06Content-Transform-Team (Work In Progress), and 2 others: Wikipedia wikis have broken maps URLs in infobox: "Bad GeoJSON - unknown \"type\" property \"ExternalData\"" - https://phabricator.wikimedia.org/T424046#12104476 (10MSantos) [11:38:24] (03PS1) 10Cathal Mooney: Add user account for Abi Bendall to analytics-privatedata-users [puppet] - 10https://gerrit.wikimedia.org/r/1309136 (https://phabricator.wikimedia.org/T430938) [11:39:13] (03CR) 10Atsuko: [C:03+1] kubernetes: run lvmd as a systemd service on selected dse-k8s workers [puppet] - 10https://gerrit.wikimedia.org/r/1309121 (https://phabricator.wikimedia.org/T429325) (owner: 10Btullis) [11:39:15] (03CR) 10CI reject: [V:04-1] Add user account for Abi Bendall to analytics-privatedata-users [puppet] - 10https://gerrit.wikimedia.org/r/1309136 (https://phabricator.wikimedia.org/T430938) (owner: 10Cathal Mooney) [11:40:33] (03CR) 10Jcrespo: [C:04-1] "Based on self reporting, we may likely need to set an expiration time, managers will confirm when asked." [puppet] - 10https://gerrit.wikimedia.org/r/1309136 (https://phabricator.wikimedia.org/T430938) (owner: 10Cathal Mooney) [11:41:48] (03CR) 10Jcrespo: [C:04-1] "I can take it from here and amend your patch so you can be free to focus on other stuff, ldap is still pending anyway." [puppet] - 10https://gerrit.wikimedia.org/r/1309136 (https://phabricator.wikimedia.org/T430938) (owner: 10Cathal Mooney) [11:42:48] (03PS2) 10Cathal Mooney: Add user account for Abi Bendall to analytics-privatedata-users [puppet] - 10https://gerrit.wikimedia.org/r/1309136 (https://phabricator.wikimedia.org/T430938) [11:43:23] (03CR) 10Lucas Werkmeister (WMDE): [C:03+1] Upgrade to the newly-released version of the Queryservice GUI [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309103 (https://phabricator.wikimedia.org/T431080) (owner: 10Arthur taylor) [11:44:07] (03CR) 10Cathal Mooney: [C:04-1] "Thanks Jaime. I'm voting -1 for this for now as it should not be merged until the LDAP access is confirmed, managers have approved and gi" [puppet] - 10https://gerrit.wikimedia.org/r/1309136 (https://phabricator.wikimedia.org/T430938) (owner: 10Cathal Mooney) [11:44:41] (03CR) 10Clément Goubert: [C:03+2] docker_registry: allow push from new contint hosts [puppet] - 10https://gerrit.wikimedia.org/r/1309135 (https://phabricator.wikimedia.org/T418521) (owner: 10Hashar) [11:45:09] (03CR) 10Jcrespo: "That looks ok, but not voting +1 until the proper approvals happen on ticket." [puppet] - 10https://gerrit.wikimedia.org/r/1309136 (https://phabricator.wikimedia.org/T430938) (owner: 10Cathal Mooney) [11:45:22] (03PS1) 10Muehlenhoff: Failover irc.w.o for reboot [dns] - 10https://gerrit.wikimedia.org/r/1309141 [11:48:22] (03CR) 10Muehlenhoff: [C:03+2] Failover irc.w.o for reboot [dns] - 10https://gerrit.wikimedia.org/r/1309141 (owner: 10Muehlenhoff) [11:48:29] !log jmm@dns1004 START - running authdns-update [11:49:10] (03CR) 10CWilliams: cookbooks/sre/mysql/decommission: add cookbook (031 comment) [cookbooks] - 10https://gerrit.wikimedia.org/r/1291952 (https://phabricator.wikimedia.org/T426613) (owner: 10Federico Ceratto) [11:50:14] !log jmm@dns1004 END - running authdns-update [11:51:22] (03CR) 10CWilliams: sre.mysql: split pool/depool (031 comment) [cookbooks] - 10https://gerrit.wikimedia.org/r/1295480 (https://phabricator.wikimedia.org/T422361) (owner: 10Federico Ceratto) [11:53:36] !log jmm@dns1004 START - running authdns-update [11:55:18] !log jmm@dns1004 END - running authdns-update [11:58:23] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host irc1003.wikimedia.org [11:58:56] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host failoid2003.codfw.wmnet [12:00:00] 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for rscout - https://phabricator.wikimedia.org/T431633#12104672 (10jcrespo) 05Open→03Resolved @Rscout Access has been deployed (It could take up to 30 minutes to be applied cluster-wide). You should be able now to private data dashbo... [12:00:04] (03PS3) 10Clément Goubert: linked-artifacts: Add service catalog and mesh entries [puppet] - 10https://gerrit.wikimedia.org/r/1308604 (https://phabricator.wikimedia.org/T431108) [12:00:05] Deploy window Mobileapps/RESTBase/Wikifeeds (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260709T1200) [12:00:08] (03PS3) 10Clément Goubert: linked-artifacts: service::catalog to production [puppet] - 10https://gerrit.wikimedia.org/r/1308605 (https://phabricator.wikimedia.org/T431108) [12:00:18] (03CR) 10Clément Goubert: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308604 (https://phabricator.wikimedia.org/T431108) (owner: 10Clément Goubert) [12:02:12] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc1003.wikimedia.org [12:02:36] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid2003.codfw.wmnet [12:03:33] (03PS3) 10Kamila Součková: P:mediawiki::php: Support debian bookworm [puppet] - 10https://gerrit.wikimedia.org/r/1295057 (owner: 10Scott French) [12:03:38] !log jynus@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[1003,1014].eqiad.wmnet with reason: reboot & upgrade [12:05:49] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host failoid1003.eqiad.wmnet [12:06:09] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host apt-staging2001.codfw.wmnet [12:07:35] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.5 point update - https://phabricator.wikimedia.org/T427072#12104731 (10MoritzMuehlenhoff) [12:08:24] 06SRE, 06Infrastructure-Foundations: Integrate Bookworm 12.14 point update - https://phabricator.wikimedia.org/T426759#12104733 (10MoritzMuehlenhoff) [12:09:41] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid1003.eqiad.wmnet [12:10:10] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt-staging2001.codfw.wmnet [12:10:48] !log jynus@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[2003,2014].codfw.wmnet with reason: reboot & upgrade [12:11:21] (03CR) 10Bartosz Wójtowicz: [C:03+2] ml-services: declare ephemeral-storage for all MI300 LLM predictors [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308563 (https://phabricator.wikimedia.org/T431089) (owner: 10Bartosz Wójtowicz) [12:14:36] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host bast7002.wikimedia.org [12:14:42] 06SRE, 10Page Content Service, 06Content-Transform-Team (Work In Progress), 07Essential-Work, 13Patch-For-Review: Failure to decode mobileapps parameters should return a 404, not a 503 - https://phabricator.wikimedia.org/T394582#12104782 (10MSantos) [12:14:59] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host apt2002.wikimedia.org [12:15:18] (03CR) 10Clément Goubert: [C:03+2] linked-artifacts: Add service catalog and mesh entries [puppet] - 10https://gerrit.wikimedia.org/r/1308604 (https://phabricator.wikimedia.org/T431108) (owner: 10Clément Goubert) [12:15:39] (03Merged) 10jenkins-bot: ml-services: declare ephemeral-storage for all MI300 LLM predictors [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308563 (https://phabricator.wikimedia.org/T431089) (owner: 10Bartosz Wójtowicz) [12:16:07] (03PS1) 10Kevin Bazira: ml-services: deploy TTS image that sets explicit intra_op_num_threads on Kokoro's ONNX session [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309157 (https://phabricator.wikimedia.org/T430536) [12:18:16] !log cmooney@cumin1003 START - Cookbook sre.network.host-bgp for host ml-serve1004 [12:18:26] !log cmooney@cumin1003 END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1004 [12:18:32] (03CR) 10CI reject: [V:04-1] Localisation updates from https://translatewiki.net. [phabricator/translations] (wmf/stable) - 10https://gerrit.wikimedia.org/r/1309158 (owner: 10L10n-bot) [12:18:33] !log bwojtowicz@deploy2003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . [12:18:36] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host bast4006.wikimedia.org [12:20:00] FIRING: [4x] NodeBGPSessionStatusNotEstablished: Kubernetes node ml-serve1003 has a BGP session which is not in the 'established' state. [12:20:15] (03PS4) 10Clément Goubert: linked-artifacts: service::catalog to production [puppet] - 10https://gerrit.wikimedia.org/r/1308605 (https://phabricator.wikimedia.org/T431108) [12:20:22] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast7002.wikimedia.org [12:20:24] (03CR) 10Clément Goubert: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308605 (https://phabricator.wikimedia.org/T431108) (owner: 10Clément Goubert) [12:20:49] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt2002.wikimedia.org [12:24:18] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast4006.wikimedia.org [12:25:00] RESOLVED: [4x] NodeBGPSessionStatusNotEstablished: Kubernetes node ml-serve1003 has a BGP session which is not in the 'established' state. [12:25:02] (03CR) 10Kamila Součková: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1295057 (owner: 10Scott French) [12:25:28] (03CR) 10Clément Goubert: "recheck" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308609 (https://phabricator.wikimedia.org/T431108) (owner: 10Clément Goubert) [12:25:33] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, 06cloud-services-team (FY2025/2026-Q3-Q4): Rebalance cloudvirts out of E4 - https://phabricator.wikimedia.org/T431682#12104867 (10fgiunchedi) [12:26:03] 10ops-eqiad, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 10netops: eqiad row A/B switch refresh prep - https://phabricator.wikimedia.org/T418012#12104878 (10VRiley-WMF) @cmooney Yes, the majority of them should be done, but I'll be double checking them as well. A8 willl need more time for the racking.... [12:27:39] FIRING: RejectingBGPPrefixes: ... [12:27:39] Rejecting 1 BGP prefixes - ml-serve1003 (10.64.141.8 - group k8s_mlserve4) - https://wikitech.wikimedia.org/wiki/Network_monitoring#RejectingBGPPrefixes - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=eqiad&var-device=lsw1-c3-eqiad:9804&var-bgp_group=k8s_mlserve4&var-bgp_neighbor=ml-serve1003 - https://alerts.wikimedia.org/?q=alertname%3DRejectingBGPPrefixes [12:28:16] (03PS1) 10Bartosz Wójtowicz: ml-services: Update ephemeral storage for gpt-oss. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309172 (https://phabricator.wikimedia.org/T431089) [12:31:18] (03PS14) 10Federico Ceratto: cookbooks/sre/mysql/decommission: add cookbook [cookbooks] - 10https://gerrit.wikimedia.org/r/1291952 (https://phabricator.wikimedia.org/T426613) [12:32:39] FIRING: [4x] RejectingBGPPrefixes: Rejecting 1 BGP prefixes - ml-serve1003 (10.64.141.8 - group k8s_mlserve4) - https://wikitech.wikimedia.org/wiki/Network_monitoring#RejectingBGPPrefixes - https://alerts.wikimedia.org/?q=alertname%3DRejectingBGPPrefixes [12:33:13] (03CR) 10Clément Goubert: [C:03+2] mobileapps: Add linked-artifacts mesh listener [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308609 (https://phabricator.wikimedia.org/T431108) (owner: 10Clément Goubert) [12:33:37] (03CR) 10Federico Ceratto: cookbooks/sre/mysql/decommission: add cookbook (031 comment) [cookbooks] - 10https://gerrit.wikimedia.org/r/1291952 (https://phabricator.wikimedia.org/T426613) (owner: 10Federico Ceratto) [12:33:52] (03CR) 10Bartosz Wójtowicz: [C:03+1] ml-services: deploy TTS image that sets explicit intra_op_num_threads on Kokoro's ONNX session [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309157 (https://phabricator.wikimedia.org/T430536) (owner: 10Kevin Bazira) [12:34:38] (03CR) 10Kevin Bazira: [C:03+1] ml-services: Update dtype to auto for qwen36-27b. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309133 (https://phabricator.wikimedia.org/T425680) (owner: 10Bartosz Wójtowicz) [12:35:25] (03Merged) 10jenkins-bot: mobileapps: Add linked-artifacts mesh listener [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308609 (https://phabricator.wikimedia.org/T431108) (owner: 10Clément Goubert) [12:36:07] (03CR) 10Kevin Bazira: [C:03+1] ml-services: Update ephemeral storage for gpt-oss. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309172 (https://phabricator.wikimedia.org/T431089) (owner: 10Bartosz Wójtowicz) [12:36:23] (03CR) 10Kevin Bazira: [C:03+2] ml-services: deploy TTS image that sets explicit intra_op_num_threads on Kokoro's ONNX session [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309157 (https://phabricator.wikimedia.org/T430536) (owner: 10Kevin Bazira) [12:36:34] (03CR) 10Kamila Součková: [V:03+1] "PCC SUCCESS (NOOP 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/8961/console" [puppet] - 10https://gerrit.wikimedia.org/r/1295057 (owner: 10Scott French) [12:38:37] (03CR) 10Bartosz Wójtowicz: [C:03+2] ml-services: Update ephemeral storage for gpt-oss. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309172 (https://phabricator.wikimedia.org/T431089) (owner: 10Bartosz Wójtowicz) [12:38:43] (03Merged) 10jenkins-bot: ml-services: deploy TTS image that sets explicit intra_op_num_threads on Kokoro's ONNX session [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309157 (https://phabricator.wikimedia.org/T430536) (owner: 10Kevin Bazira) [12:40:47] (03Merged) 10jenkins-bot: ml-services: Update ephemeral storage for gpt-oss. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309172 (https://phabricator.wikimedia.org/T431089) (owner: 10Bartosz Wójtowicz) [12:42:33] (03CR) 10Nikerabbit: [V:03+2] Localisation updates from https://translatewiki.net. [phabricator/translations] (wmf/stable) - 10https://gerrit.wikimedia.org/r/1309158 (owner: 10L10n-bot) [12:42:41] !log kevinbazira@deploy2003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . [12:43:44] !log cgoubert@deploy2003 helmfile [staging] START helmfile.d/services/mobileapps: apply [12:43:58] !log cgoubert@deploy2003 helmfile [staging] DONE helmfile.d/services/mobileapps: apply [12:44:30] !log cgoubert@deploy2003 helmfile [codfw] START helmfile.d/services/mobileapps: apply [12:44:40] !log bwojtowicz@deploy2003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . [12:45:41] !log cgoubert@deploy2003 helmfile [codfw] DONE helmfile.d/services/mobileapps: apply [12:47:37] (03PS3) 10Giuseppe Lavagetto: haproxy: modularize hiddenparma support [puppet] - 10https://gerrit.wikimedia.org/r/1309128 (https://phabricator.wikimedia.org/T422235) [12:50:26] (03CR) 10Bartosz Wójtowicz: [C:03+2] ml-services: Update dtype to auto for qwen36-27b. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309133 (https://phabricator.wikimedia.org/T425680) (owner: 10Bartosz Wójtowicz) [12:51:20] (03PS4) 10Giuseppe Lavagetto: haproxy: modularize hiddenparma support [puppet] - 10https://gerrit.wikimedia.org/r/1309128 (https://phabricator.wikimedia.org/T422235) [12:52:20] (03CR) 10Giuseppe Lavagetto: [V:03+1] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/8963/co" [puppet] - 10https://gerrit.wikimedia.org/r/1309128 (https://phabricator.wikimedia.org/T422235) (owner: 10Giuseppe Lavagetto) [12:52:38] !log cgoubert@deploy2003 helmfile [eqiad] START helmfile.d/services/mobileapps: apply [12:52:50] (03Merged) 10jenkins-bot: ml-services: Update dtype to auto for qwen36-27b. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309133 (https://phabricator.wikimedia.org/T425680) (owner: 10Bartosz Wójtowicz) [12:54:03] (03CR) 10Clément Goubert: [C:03+2] linked-artifacts: service::catalog to production [puppet] - 10https://gerrit.wikimedia.org/r/1308605 (https://phabricator.wikimedia.org/T431108) (owner: 10Clément Goubert) [12:54:12] !log bwojtowicz@deploy2003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . [12:54:13] (03CR) 10Brouberol: [C:03+1] opensearch: match pod anti-affinity to the actual cluster name [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309108 (https://phabricator.wikimedia.org/T402408) (owner: 10Btullis) [12:54:19] !log cgoubert@deploy2003 helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply [12:54:28] (03CR) 10Brouberol: [C:03+1] opensearch-cluster: render nodePool affinity through tpl [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309107 (https://phabricator.wikimedia.org/T402408) (owner: 10Btullis) [12:57:11] (03PS1) 10Brouberol: Define public DNS records for WDQS/qlever [dns] - 10https://gerrit.wikimedia.org/r/1309183 (https://phabricator.wikimedia.org/T429150) [12:58:13] (03PS5) 10Giuseppe Lavagetto: haproxy: modularize hiddenparma support [puppet] - 10https://gerrit.wikimedia.org/r/1309128 (https://phabricator.wikimedia.org/T422235) [12:59:12] I have patches scheduled in this window, but I'll be back in around 15 minutes (I'm a deployer, there's no need for anyone to wait for me, I'll deploy them myself). [12:59:25] (03CR) 10Giuseppe Lavagetto: [V:03+1] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/8964/co" [puppet] - 10https://gerrit.wikimedia.org/r/1309128 (https://phabricator.wikimedia.org/T422235) (owner: 10Giuseppe Lavagetto) [13:00:04] Lucas_WMDE, urbanecm, and TheresNoTime: How many deployers does it take to do UTC afternoon backport window deploy? (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260709T1300). [13:00:05] James_F and Msz2001: A patch you scheduled for UTC afternoon backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [13:00:11] * James_F waves. [13:00:27] o/ [13:01:16] I’m guessing you can both self-service? [13:01:49] I can. Msz2001: I can do yours too? [13:02:44] (03CR) 10TrainBranchBot: [C:03+2] "Approved by jforrester@deploy2003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309112 (https://phabricator.wikimedia.org/T430898) (owner: 10Jforrester) [13:02:45] You can, there's nothing to verify really. I can do mine in about 15 minutes, but if you'd like to deploy that, it's okay [13:02:58] I'll just do mine and then hand over to you. [13:03:05] Okay [13:04:55] RESOLVED: [2x] SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:05:40] (03PS2) 10Brouberol: Define public DNS records for WDQS/qlever [dns] - 10https://gerrit.wikimedia.org/r/1309183 (https://phabricator.wikimedia.org/T429150) [13:06:10] (03Merged) 10jenkins-bot: abstractwiki: Don't run our jobs too quickly, they'll collide and error [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309112 (https://phabricator.wikimedia.org/T430898) (owner: 10Jforrester) [13:06:11] (03PS3) 10Brouberol: Define public/discovery DNS records for WDQS/qlever [dns] - 10https://gerrit.wikimedia.org/r/1309183 (https://phabricator.wikimedia.org/T429150) [13:06:19] !log jforrester@deploy2003 Started scap sync-world: Backport for [[gerrit:1309112|abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] [13:06:22] T430898: AW now calls for all fragments at once and overloads WF - https://phabricator.wikimedia.org/T430898 [13:08:02] !log jforrester@deploy2003 jforrester: Backport for [[gerrit:1309112|abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [13:13:26] !log jforrester@deploy2003 jforrester: Continuing with deployment [13:17:39] RESOLVED: [4x] RejectingBGPPrefixes: Rejecting 1 BGP prefixes - ml-serve1003 (10.64.141.8 - group k8s_mlserve4) - https://wikitech.wikimedia.org/wiki/Network_monitoring#RejectingBGPPrefixes - https://alerts.wikimedia.org/?q=alertname%3DRejectingBGPPrefixes [13:17:45] !log jforrester@deploy2003 Finished scap sync-world: Backport for [[gerrit:1309112|abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] (duration: 11m 26s) [13:17:49] T430898: AW now calls for all fragments at once and overloads WF - https://phabricator.wikimedia.org/T430898 [13:18:03] Msz2001: Over to you. [13:21:37] (03CR) 10CWilliams: cookbooks/sre/mysql/decommission: add cookbook (031 comment) [cookbooks] - 10https://gerrit.wikimedia.org/r/1291952 (https://phabricator.wikimedia.org/T426613) (owner: 10Federico Ceratto) [13:24:28] Okay, I'm back [13:24:39] James_F: Thanks, Proceeding to deploy mine [13:25:18] (03CR) 10TrainBranchBot: [C:03+2] "Approved by mszwarc@deploy2003 using scap backport" [core] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1309101 (https://phabricator.wikimedia.org/T431176) (owner: 10Mszwarc) [13:25:18] (03CR) 10TrainBranchBot: [C:03+2] "Approved by mszwarc@deploy2003 using scap backport" [extensions/CheckUser] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1309122 (https://phabricator.wikimedia.org/T429320) (owner: 10Mszwarc) [13:28:29] (03CR) 10Brouberol: [C:03+1] Puppet 8: Replace legacy facts in profile opensearch [puppet] - 10https://gerrit.wikimedia.org/r/1308750 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [13:28:47] (03CR) 10Brouberol: [C:03+1] Puppet 8: Replace legacy facts in module opensearch [puppet] - 10https://gerrit.wikimedia.org/r/1308721 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [13:29:04] (03CR) 10Brouberol: [C:03+1] Puppet 8: Replace legacy facts in misc kafka [puppet] - 10https://gerrit.wikimedia.org/r/1308760 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [13:29:43] !log blake@cumin1003 START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1052.eqiad.wmnet [13:29:47] !log blake@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1052.eqiad.wmnet [13:30:22] !log blake@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1052.eqiad.wmnet [13:30:46] !log blake@cumin1003 START - Cookbook sre.hosts.reimage for host wikikube-worker1052.eqiad.wmnet with OS trixie [13:31:03] !log blake@cumin1003 START - Cookbook sre.hosts.move-vlan for host wikikube-worker1052 [13:31:14] !log blake@cumin1003 START - Cookbook sre.dns.netbox [13:31:45] (03PS6) 10Filippo Giunchedi: dumps: add nfs_client for dumps-nfs.w.o [puppet] - 10https://gerrit.wikimedia.org/r/1308126 (https://phabricator.wikimedia.org/T411248) [13:31:45] (03PS6) 10Filippo Giunchedi: statistics: optional support for dumps-nfs.w.o [puppet] - 10https://gerrit.wikimedia.org/r/1308127 (https://phabricator.wikimedia.org/T411248) [13:31:45] (03PS6) 10Filippo Giunchedi: wmcs: add dumps-nfs.w.o support [puppet] - 10https://gerrit.wikimedia.org/r/1308128 (https://phabricator.wikimedia.org/T411248) [13:31:46] (03PS1) 10Filippo Giunchedi: labstore: don't check after umount [puppet] - 10https://gerrit.wikimedia.org/r/1309200 (https://phabricator.wikimedia.org/T411248) [13:32:20] (03CR) 10CI reject: [V:04-1] dumps: add nfs_client for dumps-nfs.w.o [puppet] - 10https://gerrit.wikimedia.org/r/1308126 (https://phabricator.wikimedia.org/T411248) (owner: 10Filippo Giunchedi) [13:32:24] (03PS1) 10Cathal Mooney: Update BGP policy naming: k8s-ml-serve and related [homer/public] - 10https://gerrit.wikimedia.org/r/1309201 (https://phabricator.wikimedia.org/T371088) [13:32:50] (03Merged) 10jenkins-bot: Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups [core] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1309101 (https://phabricator.wikimedia.org/T431176) (owner: 10Mszwarc) [13:33:11] (03CR) 10CI reject: [V:04-1] wmcs: add dumps-nfs.w.o support [puppet] - 10https://gerrit.wikimedia.org/r/1308128 (https://phabricator.wikimedia.org/T411248) (owner: 10Filippo Giunchedi) [13:33:44] (03Merged) 10jenkins-bot: SuggestedInvestigations: Instrument link clicks in the cases table [extensions/CheckUser] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1309122 (https://phabricator.wikimedia.org/T429320) (owner: 10Mszwarc) [13:33:59] !log mszwarc@deploy2003 Started scap sync-world: Backport for [[gerrit:1309101|Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122|SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] [13:34:02] T431176: User right log doesn't tally on a wiki and Meta - https://phabricator.wikimedia.org/T431176 [13:34:26] (03CR) 10JMeybohm: [C:03+1] admin_ng: Install coredns-internalonly in wikikube [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308802 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [13:35:26] !log blake@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" [13:35:31] !log blake@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003" [13:35:31] !log blake@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [13:35:31] !log blake@cumin1003 START - Cookbook sre.dns.wipe-cache wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [13:35:35] !log blake@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [13:35:35] !log blake@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1052 [13:35:42] !log mszwarc@deploy2003 mszwarc: Backport for [[gerrit:1309101|Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122|SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [13:36:10] !log blake@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1052 [13:36:10] !log blake@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1052 [13:37:12] !log mszwarc@deploy2003 mszwarc: Continuing with deployment [13:38:23] (03CR) 10Cathal Mooney: [C:03+2] Nokia SR-Linux: fix issues with BGP policies matching communities [homer/public] - 10https://gerrit.wikimedia.org/r/1308689 (https://phabricator.wikimedia.org/T430810) (owner: 10Cathal Mooney) [13:39:50] (03Merged) 10jenkins-bot: Nokia SR-Linux: fix issues with BGP policies matching communities [homer/public] - 10https://gerrit.wikimedia.org/r/1308689 (https://phabricator.wikimedia.org/T430810) (owner: 10Cathal Mooney) [13:40:23] (03PS7) 10Filippo Giunchedi: dumps: add nfs_client for dumps-nfs.w.o [puppet] - 10https://gerrit.wikimedia.org/r/1308126 (https://phabricator.wikimedia.org/T411248) [13:40:23] (03PS7) 10Filippo Giunchedi: statistics: optional support for dumps-nfs.w.o [puppet] - 10https://gerrit.wikimedia.org/r/1308127 (https://phabricator.wikimedia.org/T411248) [13:40:23] (03PS7) 10Filippo Giunchedi: wmcs: add dumps-nfs.w.o support [puppet] - 10https://gerrit.wikimedia.org/r/1308128 (https://phabricator.wikimedia.org/T411248) [13:40:23] (03PS2) 10Filippo Giunchedi: labstore: don't check after umount [puppet] - 10https://gerrit.wikimedia.org/r/1309200 (https://phabricator.wikimedia.org/T411248) [13:41:11] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host apt1002.wikimedia.org [13:41:28] !log mszwarc@deploy2003 Finished scap sync-world: Backport for [[gerrit:1309101|Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122|SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] (duration: 07m 30s) [13:41:32] T431176: User right log doesn't tally on a wiki and Meta - https://phabricator.wikimedia.org/T431176 [13:41:55] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt1002.wikimedia.org [13:43:29] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply [13:43:40] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply [13:43:52] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host netboxdb2003.codfw.wmnet [13:44:01] !log Updated `logging` on `metawiki` to fix log performers, T431176#12105297 [13:44:03] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [13:44:22] !log UTC afternoon config+backport window is done [13:44:23] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [13:47:51] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb2003.codfw.wmnet [13:49:00] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.4 point update - https://phabricator.wikimedia.org/T420240#12105309 (10MoritzMuehlenhoff) [13:50:17] !log installing python-cryptography security updates [13:50:20] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [13:51:17] (03PS1) 10Aqu: airflow: add gcs-token-search-console secret [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309204 (https://phabricator.wikimedia.org/T427457) [13:52:40] (03CR) 10Cathal Mooney: [C:03+2] "Sorry. I self-merged this in error, meant to do the other one. Given it's done now I'm gonna roll out, feel free to review we can unwind" [homer/public] - 10https://gerrit.wikimedia.org/r/1308689 (https://phabricator.wikimedia.org/T430810) (owner: 10Cathal Mooney) [13:53:03] !log blake@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage [13:53:17] (03CR) 10Btullis: [C:03+1] "Looks good to me." [dns] - 10https://gerrit.wikimedia.org/r/1309183 (https://phabricator.wikimedia.org/T429150) (owner: 10Brouberol) [13:55:10] (03PS2) 10Cathal Mooney: Update BGP policy naming: k8s-ml-serve and related [homer/public] - 10https://gerrit.wikimedia.org/r/1309201 (https://phabricator.wikimedia.org/T371088) [13:55:38] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host netboxdb1003.eqiad.wmnet [13:56:25] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host cuminunpriv1001.eqiad.wmnet [13:57:19] !log installing requests security updates [13:57:20] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [13:58:27] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.4 point update - https://phabricator.wikimedia.org/T420240#12105375 (10MoritzMuehlenhoff) [13:59:24] !log blake@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage [13:59:28] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb1003.eqiad.wmnet [13:59:31] (03CR) 10Cathal Mooney: [C:03+2] Update BGP policy naming: k8s-ml-serve and related [homer/public] - 10https://gerrit.wikimedia.org/r/1309201 (https://phabricator.wikimedia.org/T371088) (owner: 10Cathal Mooney) [13:59:59] (03PS1) 10DCausse: cirrus: change the restart strategy to exponential-delay [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309210 (https://phabricator.wikimedia.org/T428863) [14:00:15] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cuminunpriv1001.eqiad.wmnet [14:00:21] (03PS2) 10Aqu: airflow: add gcs-token-search-console secret [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309204 (https://phabricator.wikimedia.org/T427457) [14:01:03] (03Merged) 10jenkins-bot: Update BGP policy naming: k8s-ml-serve and related [homer/public] - 10https://gerrit.wikimedia.org/r/1309201 (https://phabricator.wikimedia.org/T371088) (owner: 10Cathal Mooney) [14:01:46] (03Abandoned) 10Aqu: Airflow-main: Add GCP Airflow connection [deployment-charts] - 10https://gerrit.wikimedia.org/r/1307184 (https://phabricator.wikimedia.org/T427457) (owner: 10Aqu) [14:03:38] (03CR) 10Bking: [C:03+1] opensearch: match pod anti-affinity to the actual cluster name [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309108 (https://phabricator.wikimedia.org/T402408) (owner: 10Btullis) [14:04:38] !log jhancock@cumin2002 START - Cookbook sre.dns.netbox [14:06:35] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.4 point update - https://phabricator.wikimedia.org/T420240#12105418 (10MoritzMuehlenhoff) [14:07:30] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply [14:07:31] (03CR) 10Bking: [C:03+1] opensearch-cluster: render nodePool affinity through tpl [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309107 (https://phabricator.wikimedia.org/T402408) (owner: 10Btullis) [14:07:49] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply [14:09:15] !log jhancock@cumin2002 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" [14:09:21] !log jhancock@cumin2002 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002" [14:09:21] !log jhancock@cumin2002 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [14:10:32] (03CR) 10Dzahn: "I think we should default to "no special checks" where we can to not have the effect of "it's different from prod"." [puppet] - 10https://gerrit.wikimedia.org/r/1308659 (https://phabricator.wikimedia.org/T431582) (owner: 10Hashar) [14:11:28] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply [14:11:36] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply [14:13:41] !log update druid indexation job for webrequest_sampled_live - T427068 [14:13:43] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:13:44] T427068: Add X-Provenance data to webrequest_sampled_live - https://phabricator.wikimedia.org/T427068 [14:14:47] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply [14:14:53] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply [14:15:34] (03CR) 10Majavah: [C:04-1] dumps: add nfs_client for dumps-nfs.w.o (035 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1308126 (https://phabricator.wikimedia.org/T411248) (owner: 10Filippo Giunchedi) [14:15:37] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply [14:15:43] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply [14:15:48] 06SRE, 06Infrastructure-Foundations, 10netops, 13Patch-For-Review: GSHUT (and other?) community matching/actions not working on SR-Linux - https://phabricator.wikimedia.org/T430810#12105447 (10cmooney) The patch is merged, and we can see the local-pref=100 on standard routes being learnt now: ` A:ssw1-d8-e... [14:17:32] (03PS1) 10Elukey: turnilo: add new X-provenance dimensions [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309212 (https://phabricator.wikimedia.org/T427068) [14:17:51] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply [14:18:02] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply [14:19:53] !log jynus@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade [14:20:18] !log blake@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1052.eqiad.wmnet with OS trixie [14:20:40] (03CR) 10Joal: [C:03+1] "LGTM!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309212 (https://phabricator.wikimedia.org/T427068) (owner: 10Elukey) [14:23:02] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply [14:23:14] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply [14:26:26] (03CR) 10Elukey: [C:03+2] turnilo: add new X-provenance dimensions [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309212 (https://phabricator.wikimedia.org/T427068) (owner: 10Elukey) [14:26:58] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply [14:28:13] !log elukey@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync [14:28:36] !log elukey@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync [14:29:06] 06SRE, 06MediaWiki-Media-Platform-Team, 06ServiceOps new, 10ServiceOps-Services-Oids, 10Thumbor: Thumbor-k8s performance improvements - https://phabricator.wikimedia.org/T333445#12105563 (10JTweed-WMF) [14:30:04] Deploy window Test Kitchen Experiment Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260709T1430) [14:30:06] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply [14:30:32] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply [14:30:37] !log elukey@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync [14:30:45] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply [14:30:53] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply [14:31:01] !log elukey@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync [14:31:34] blake@cumin1003 renumber-node (PID 3243349) is awaiting input [14:31:50] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply [14:32:14] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply [14:32:30] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply [14:35:30] (03CR) 10Federico Ceratto: cookbooks/sre/mysql/decommission: add cookbook (031 comment) [cookbooks] - 10https://gerrit.wikimedia.org/r/1291952 (https://phabricator.wikimedia.org/T426613) (owner: 10Federico Ceratto) [14:35:47] 10SRE-swift-storage, 10Thumbor: Gradually drop all thumbnails as a one-off clean up - https://phabricator.wikimedia.org/T379942#12105629 (10JTweed-WMF) [14:35:48] 06SRE, 10SRE-swift-storage, 10Thumbor, 06Traffic: Cache thumbs in our caching infrastructure (e.g. ATS) - https://phabricator.wikimedia.org/T345334#12105630 (10JTweed-WMF) [14:37:45] (03PS4) 10Scott French: P:mediawiki::php: Refactor support for debian bookworm [puppet] - 10https://gerrit.wikimedia.org/r/1295057 (https://phabricator.wikimedia.org/T423714) [14:39:42] blake@cumin1003 renumber-node (PID 3243349) is awaiting input [14:40:13] (03CR) 10Scott French: [C:03+1] "Thanks, for reviving / rebasing this, Raine! I think this looks good, and indeed the PCC diffs remain noops for both debian versions." [puppet] - 10https://gerrit.wikimedia.org/r/1295057 (https://phabricator.wikimedia.org/T423714) (owner: 10Scott French) [14:40:50] (03CR) 10Kamila Součková: [C:03+2] P:mediawiki::php: Refactor support for debian bookworm [puppet] - 10https://gerrit.wikimedia.org/r/1295057 (https://phabricator.wikimedia.org/T423714) (owner: 10Scott French) [14:42:20] !log blake@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1052.eqiad.wmnet [14:42:20] !log blake@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1052.eqiad.wmnet [14:42:23] !log blake@cumin1003 END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1052.eqiad.wmnet [14:43:12] (03PS1) 10Clare Ming: Test Kitchen UI: Deploy v1.4.7 release to staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309221 (https://phabricator.wikimedia.org/T428984) [14:43:16] (03PS1) 10Elukey: profile::benthos: update x-provenance's fields to "0" and "1" [puppet] - 10https://gerrit.wikimedia.org/r/1309222 (https://phabricator.wikimedia.org/T427068) [14:45:07] (03CR) 10CWilliams: cookbooks/sre/mysql/decommission: add cookbook (031 comment) [cookbooks] - 10https://gerrit.wikimedia.org/r/1291952 (https://phabricator.wikimedia.org/T426613) (owner: 10Federico Ceratto) [14:45:38] !log elukey@cumin1003 START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet [14:46:41] (03CR) 10Santiago Faci: [C:03+2] Test Kitchen UI: Deploy v1.4.7 release to staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309221 (https://phabricator.wikimedia.org/T428984) (owner: 10Clare Ming) [14:47:40] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet [14:48:29] !log elukey@cumin1003 START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet [14:48:37] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet [14:48:38] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.4 point update - https://phabricator.wikimedia.org/T420240#12105727 (10MoritzMuehlenhoff) [14:49:03] (03Merged) 10jenkins-bot: Test Kitchen UI: Deploy v1.4.7 release to staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309221 (https://phabricator.wikimedia.org/T428984) (owner: 10Clare Ming) [14:49:11] !log jynus@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade [14:51:43] !log jynus@cumin1003 START - Cookbook sre.hosts.remove-downtime for 14 hosts [14:51:51] !log jynus@cumin1003 END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts [14:52:35] (03CR) 10Kosta Harlan: [C:03+1] "@Ladsgroup@gmail.com if you're able to do it today, that would be great. Otherwise early next week would be OK for us too." [puppet] - 10https://gerrit.wikimedia.org/r/1307100 (https://phabricator.wikimedia.org/T386456) (owner: 10Dreamy Jazz) [14:52:52] (03PS7) 10Clément Goubert: sre.k8s.renumber-node: Increase puppet timeout and attempts [cookbooks] - 10https://gerrit.wikimedia.org/r/1307872 (https://phabricator.wikimedia.org/T430226) (owner: 10Jasmine) [14:53:24] (03CR) 10Clément Goubert: [C:03+1] sre.k8s.renumber-node: Increase puppet timeout and attempts [cookbooks] - 10https://gerrit.wikimedia.org/r/1307872 (https://phabricator.wikimedia.org/T430226) (owner: 10Jasmine) [14:55:49] (03PS1) 10Elukey: sre.hosts.bmc-user-mgmt: add more logging [cookbooks] - 10https://gerrit.wikimedia.org/r/1309226 (https://phabricator.wikimedia.org/T426180) [14:57:43] FIRING: KeyholderUnarmed: 1 unarmed Keyholder key(s) on netmon2002:9100 - https://wikitech.wikimedia.org/wiki/Keyholder - TODO - https://alerts.wikimedia.org/?q=alertname%3DKeyholderUnarmed [14:58:24] 10SRE-swift-storage, 10Thumbor: Gradually drop all thumbnails as a one-off clean up - https://phabricator.wikimedia.org/T379942#12105819 (10JTweed-WMF) [14:58:26] 06SRE, 10SRE-swift-storage, 10Thumbor, 06Traffic: Cache thumbs in our caching infrastructure (e.g. ATS) - https://phabricator.wikimedia.org/T345334#12105821 (10JTweed-WMF) [14:58:37] !log cjming@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply [14:59:14] !log cjming@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply [15:00:05] hashar and andre: Deploy window Train log triage (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260709T1500) [15:02:26] (03CR) 10Hashar: "This was solved by restoring the backed up path to just `/srv/gerrit/git` and `/srv/gerrit/data` instead of `/srv/gerrit` which holds a lo" [puppet] - 10https://gerrit.wikimedia.org/r/1306166 (https://phabricator.wikimedia.org/T257744) (owner: 10Arnaudb) [15:02:42] (03CR) 10Joal: [C:03+1] "LGTM!" [puppet] - 10https://gerrit.wikimedia.org/r/1309222 (https://phabricator.wikimedia.org/T427068) (owner: 10Elukey) [15:04:34] (03CR) 10SBassett: Revert^2 "varnish: Add CSP report-only header value" (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1307826 (owner: 10BCornwall) [15:06:19] !log cjming@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply [15:06:21] (03CR) 10Elukey: cookbooks/sre/mysql/decommission: add cookbook (032 comments) [cookbooks] - 10https://gerrit.wikimedia.org/r/1291952 (https://phabricator.wikimedia.org/T426613) (owner: 10Federico Ceratto) [15:06:47] !log cjming@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply [15:13:28] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply [15:13:40] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply [15:14:06] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply [15:14:16] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply [15:14:37] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply [15:14:45] !log javiermonton@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply [15:18:48] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.4 point update - https://phabricator.wikimedia.org/T420240#12105931 (10MoritzMuehlenhoff) [15:19:23] FIRING: JobUnavailable: Reduced availability for job minio in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [15:24:23] RESOLVED: JobUnavailable: Reduced availability for job minio in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [15:25:59] 06SRE, 06Infrastructure-Foundations, 10netops: Renumber link addressing to lsw1-e1-eqiad and lsw1-f1-eqiad - https://phabricator.wikimedia.org/T404248#12105978 (10cmooney) 05Open→03Resolved This was resolved some time back. [15:28:46] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308757 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [15:30:16] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in module airflow [puppet] - 10https://gerrit.wikimedia.org/r/1308716 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [15:34:42] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in profile hadoop [puppet] - 10https://gerrit.wikimedia.org/r/1308761 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [15:35:07] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in profile zookeeper [puppet] - 10https://gerrit.wikimedia.org/r/1308753 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [15:42:01] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.4 point update - https://phabricator.wikimedia.org/T420240#12106035 (10MoritzMuehlenhoff) [15:42:15] (03PS1) 10Hnowlan: lists: migrate nrpe checks to prometheus [puppet] - 10https://gerrit.wikimedia.org/r/1309238 (https://phabricator.wikimedia.org/T370157) [15:42:35] (03CR) 10Elukey: [C:03+2] profile::benthos: update x-provenance's fields to "0" and "1" [puppet] - 10https://gerrit.wikimedia.org/r/1309222 (https://phabricator.wikimedia.org/T427068) (owner: 10Elukey) [15:43:13] jhathaway: ok to puppet-merge? [15:43:22] yes [15:43:26] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.4 point update - https://phabricator.wikimedia.org/T420240#12106047 (10MoritzMuehlenhoff) [15:43:45] proceeding! [15:44:16] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308760 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [15:45:06] (03CR) 10Hnowlan: [V:03+1] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/8966/co" [puppet] - 10https://gerrit.wikimedia.org/r/1309238 (https://phabricator.wikimedia.org/T370157) (owner: 10Hnowlan) [15:45:22] (03CR) 10RLazarus: "It's https://gerrit.wikimedia.org/r/1308802 -- did the link not work? That's weird." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1308788 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [15:45:58] (03CR) 10Jasmine: sre.k8s.renumber-node: Increase puppet timeout and attempts (031 comment) [cookbooks] - 10https://gerrit.wikimedia.org/r/1307872 (https://phabricator.wikimedia.org/T430226) (owner: 10Jasmine) [15:47:46] !log jynus@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backupmon1001.eqiad.wmnet with reason: restart [15:48:43] (03PS1) 10Clare Ming: Test Kitchen UI: Deploy v1.4.8 release to staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309240 (https://phabricator.wikimedia.org/T428984) [15:49:02] (03CR) 10Majavah: statistics: optional support for dumps-nfs.w.o (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1308127 (https://phabricator.wikimedia.org/T411248) (owner: 10Filippo Giunchedi) [15:49:13] !log jynus@cumin1003 START - Cookbook sre.hosts.remove-downtime for 14 hosts [15:49:23] !log jynus@cumin1003 END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts [15:50:15] (03PS1) 10Pppery: Update translations [phabricator/translations] (wmf/stable) - 10https://gerrit.wikimedia.org/r/1309241 [15:50:38] (03PS1) 10JMeybohm: kind.sh: Use istio 1.29.4 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309242 (https://phabricator.wikimedia.org/T427401) [15:51:22] (03CR) 10Santiago Faci: [C:03+2] Test Kitchen UI: Deploy v1.4.8 release to staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309240 (https://phabricator.wikimedia.org/T428984) (owner: 10Clare Ming) [15:51:27] !log restarting backupmon1001 [15:51:28] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [15:53:40] (03Merged) 10jenkins-bot: Test Kitchen UI: Deploy v1.4.8 release to staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309240 (https://phabricator.wikimedia.org/T428984) (owner: 10Clare Ming) [15:53:46] (03CR) 10JMeybohm: [C:03+2] kind.sh: Use istio 1.29.4 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309242 (https://phabricator.wikimedia.org/T427401) (owner: 10JMeybohm) [15:54:29] !log cjming@deploy2003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply [15:55:00] !log cjming@deploy2003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply [15:55:28] (03Merged) 10jenkins-bot: kind.sh: Use istio 1.29.4 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309242 (https://phabricator.wikimedia.org/T427401) (owner: 10JMeybohm) [15:57:37] (03CR) 10Ladsgroup: [C:03+1] "Hi, I get to it ASAP. I need to fix something first and then do this." [puppet] - 10https://gerrit.wikimedia.org/r/1307100 (https://phabricator.wikimedia.org/T386456) (owner: 10Dreamy Jazz) [16:00:05] jhathaway and rzl: #bothumor My software never has bugs. It just develops random features. Rise for Puppet request window. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260709T1600). [16:00:05] No Gerrit patches in the queue for this window AFAICS. [16:01:02] (03PS2) 10JHathaway: Puppet 8: Replace legacy facts in misc kafka [puppet] - 10https://gerrit.wikimedia.org/r/1308760 (https://phabricator.wikimedia.org/T372666) [16:01:11] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1308760 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [16:02:20] PROBLEM - Check if Pybal has been restarted after pybal.conf was changed on lvs2012 is CRITICAL: CRITICAL: Service pybal.service has not been restarted after /etc/pybal/pybal.conf was changed (gt 1h). https://wikitech.wikimedia.org/wiki/PyBal%23Pybal_service_has_not_been_restarted [16:05:51] (03PS1) 10Pppery: Fix broken link in scripts/ and support/ docs [phabricator/translations] (wmf/stable) - 10https://gerrit.wikimedia.org/r/1309245 (https://phabricator.wikimedia.org/T429078) [16:05:59] (03PS1) 10Hnowlan: wikidough: migrate NRPE checks to prometheus [puppet] - 10https://gerrit.wikimedia.org/r/1309246 (https://phabricator.wikimedia.org/T384438) [16:07:06] (03CR) 10Hnowlan: [V:03+1] "PCC SUCCESS (DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/8967/console" [puppet] - 10https://gerrit.wikimedia.org/r/1309246 (https://phabricator.wikimedia.org/T384438) (owner: 10Hnowlan) [16:09:08] (03CR) 10Hnowlan: [V:03+1] "PCC SUCCESS (): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/8968/console" [puppet] - 10https://gerrit.wikimedia.org/r/1309246 (https://phabricator.wikimedia.org/T384438) (owner: 10Hnowlan) [16:11:12] (03PS1) 10Pppery: Remove the commented out bit about AM/PM [phabricator/translations] (wmf/stable) - 10https://gerrit.wikimedia.org/r/1309247 (https://phabricator.wikimedia.org/T428957) [16:11:24] PROBLEM - Wikidough DoH Check -IPv4- on doh5004 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/Wikidough/Monitoring%23Wikidough_Basic_Check [16:11:24] PROBLEM - Wikidough DoH Check -IPv6- on doh5004 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/Wikidough/Monitoring%23Wikidough_Basic_Check [16:12:11] PROBLEM - Wikidough DoT Check -IPv6- on doh5004 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/Wikidough/Monitoring%23Wikidough_Basic_Check [16:12:11] PROBLEM - SSH on doh5004 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/SSH/monitoring [16:12:11] PROBLEM - Wikidough DoT Check -IPv4- on doh5004 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/Wikidough/Monitoring%23Wikidough_Basic_Check [16:14:32] (03PS2) 10Pppery: Remove the commented out bit about AM/PM [phabricator/translations] (wmf/stable) - 10https://gerrit.wikimedia.org/r/1309247 (https://phabricator.wikimedia.org/T428957) [16:15:02] RECOVERY - Wikidough DoT Check -IPv6- on doh5004 is OK: TCP OK - 0.468 second response time on 2001:df2:e500:3:103:102:166:99 port 853 https://wikitech.wikimedia.org/wiki/Wikidough/Monitoring%23Wikidough_Basic_Check [16:15:02] RECOVERY - Wikidough DoT Check -IPv4- on doh5004 is OK: TCP OK - 0.510 second response time on 103.102.166.99 port 853 https://wikitech.wikimedia.org/wiki/Wikidough/Monitoring%23Wikidough_Basic_Check [16:15:02] RECOVERY - SSH on doh5004 is OK: SSH OK - OpenSSH_9.2p1 Debian-2+deb12u9 (protocol 2.0) https://wikitech.wikimedia.org/wiki/SSH/monitoring [16:15:11] awkward timing :E [16:15:14] RECOVERY - Wikidough DoH Check -IPv6- on doh5004 is OK: HTTP OK: HTTP/1.1 200 OK - 595 bytes in 0.925 second response time https://wikitech.wikimedia.org/wiki/Wikidough/Monitoring%23Wikidough_Basic_Check [16:15:15] RECOVERY - Wikidough DoH Check -IPv4- on doh5004 is OK: HTTP OK: HTTP/1.1 200 OK - 595 bytes in 0.988 second response time https://wikitech.wikimedia.org/wiki/Wikidough/Monitoring%23Wikidough_Basic_Check [16:15:22] (03CR) 10Hnowlan: [V:03+1] "PCC SUCCESS (NOOP 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/8969/console" [puppet] - 10https://gerrit.wikimedia.org/r/1309246 (https://phabricator.wikimedia.org/T384438) (owner: 10Hnowlan) [16:15:53] (03PS3) 10Pppery: Remove the commented out bit about AM/PM [phabricator/translations] (wmf/stable) - 10https://gerrit.wikimedia.org/r/1309247 (https://phabricator.wikimedia.org/T428957) [16:17:34] (03CR) 10Jdlrobson: "No problem! Please backport this when you're ready. I just wanted to get the patch up on Gerrit!" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308671 (https://phabricator.wikimedia.org/T431514) (owner: 10Jdlrobson) [16:18:35] (03CR) 10Ladsgroup: "Thank you <3" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308671 (https://phabricator.wikimedia.org/T431514) (owner: 10Jdlrobson) [16:20:12] PROBLEM - SSH on doh5004 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/SSH/monitoring [16:20:24] PROBLEM - Wikidough DoH Check -IPv4- on doh5004 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/Wikidough/Monitoring%23Wikidough_Basic_Check [16:20:24] PROBLEM - Wikidough DoH Check -IPv6- on doh5004 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/Wikidough/Monitoring%23Wikidough_Basic_Check [16:21:10] PROBLEM - Wikidough DoT Check -IPv6- on doh5004 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/Wikidough/Monitoring%23Wikidough_Basic_Check [16:21:12] PROBLEM - Wikidough DoT Check -IPv4- on doh5004 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/Wikidough/Monitoring%23Wikidough_Basic_Check [16:21:16] hnowlan: awkward in what way? Should we be aware of some maintenance going on? [16:24:00] RECOVERY - Wikidough DoT Check -IPv6- on doh5004 is OK: TCP OK - 0.468 second response time on 2001:df2:e500:3:103:102:166:99 port 853 https://wikitech.wikimedia.org/wiki/Wikidough/Monitoring%23Wikidough_Basic_Check [16:24:04] RECOVERY - Wikidough DoT Check -IPv4- on doh5004 is OK: TCP OK - 1.888 second response time on 103.102.166.99 port 853 https://wikitech.wikimedia.org/wiki/Wikidough/Monitoring%23Wikidough_Basic_Check [16:24:05] oh, I see, the nrpe timing check changes [16:24:21] indeed, that is awkward! [16:24:23] FIRING: JobUnavailable: Reduced availability for job wikidough in ops@eqsin - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:28:42] 10ops-eqiad, 06SRE, 06DC-Ops: Degraded RAID on backup1019 - https://phabricator.wikimedia.org/T431367#12106217 (10jcrespo) Thank you a lot, @VRiley-WMF ! Feel free to replace it if it can be done in a hot way, but if let me know with no isue if you require stopping the server, or it would simplify a lot the... [16:29:06] brett: sorry for the alarm! they're still in review, nothing has changed (I HOPE) [16:29:55] hnowlan: No worries, we do actually have something wrong with that host right now! [16:35:04] (03CR) 10Dzahn: [C:03+1] "@hashar any concerns on your end? if no veto I will deploy this" [puppet] - 10https://gerrit.wikimedia.org/r/1308081 (https://phabricator.wikimedia.org/T431398) (owner: 10Arnaudb) [16:35:10] PROBLEM - Wikidough DoT Check -IPv6- on doh5004 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/Wikidough/Monitoring%23Wikidough_Basic_Check [16:35:57] jouncebot: nowandnext [16:35:57] For the next 0 hour(s) and 24 minute(s): Puppet request window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260709T1600) [16:35:57] In 0 hour(s) and 24 minute(s): Cloud Services/Technical Documentation weekly deploy (Toolhub, Developer portal, Striker) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260709T1700) [16:35:57] In 0 hour(s) and 24 minute(s): MediaWiki infrastructure (UTC late) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260709T1700) [16:36:11] okay, trying to deploy the thumbor thing right now for real [16:36:12] PROBLEM - Wikidough DoT Check -IPv4- on doh5004 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/Wikidough/Monitoring%23Wikidough_Basic_Check [16:38:09] (03CR) 10Dzahn: "it seems the right way to do this first on secondary backends - but on the other hand .. if we have no traffic on it and no way to test - " [puppet] - 10https://gerrit.wikimedia.org/r/1304808 (https://phabricator.wikimedia.org/T429749) (owner: 10Arnaudb) [16:39:40] (03PS1) 10Pppery: Export source strings again [phabricator/translations] (wmf/stable) - 10https://gerrit.wikimedia.org/r/1309254 [16:41:08] (03PS1) 10Ladsgroup: thumbor: Load files from swift async [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309255 (https://phabricator.wikimedia.org/T333445) [16:44:06] RECOVERY - Wikidough DoT Check -IPv4- on doh5004 is OK: TCP OK - 3.965 second response time on 103.102.166.99 port 853 https://wikitech.wikimedia.org/wiki/Wikidough/Monitoring%23Wikidough_Basic_Check [16:44:06] RECOVERY - Wikidough DoT Check -IPv6- on doh5004 is OK: TCP OK - 4.949 second response time on 2001:df2:e500:3:103:102:166:99 port 853 https://wikitech.wikimedia.org/wiki/Wikidough/Monitoring%23Wikidough_Basic_Check [16:44:34] PROBLEM - Bird Internet Routing Daemon on doh5004 is CRITICAL: PROCS CRITICAL: 0 processes with command name bird https://wikitech.wikimedia.org/wiki/Anycast%23Bird_daemon_not_running [16:45:57] ^I've stopped bird.service, will downtime [16:47:38] (03CR) 10Hnowlan: [C:03+1] thumbor: Load files from swift async [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309255 (https://phabricator.wikimedia.org/T333445) (owner: 10Ladsgroup) [16:47:54] (03CR) 10Ladsgroup: [C:03+2] thumbor: Load files from swift async [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309255 (https://phabricator.wikimedia.org/T333445) (owner: 10Ladsgroup) [16:47:55] !log brett@cumin2002 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on doh5004.wikimedia.org with reason: random high load, investigating [16:49:18] (03PS2) 10Pppery: Export source strings again [phabricator/translations] (wmf/stable) - 10https://gerrit.wikimedia.org/r/1309254 [16:49:23] RESOLVED: JobUnavailable: Reduced availability for job wikidough in ops@eqsin - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:49:39] !log planet1003/planet2003 - rebooting [16:49:41] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [16:50:17] (03Merged) 10jenkins-bot: thumbor: Load files from swift async [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309255 (https://phabricator.wikimedia.org/T333445) (owner: 10Ladsgroup) [16:50:18] RECOVERY - Wikidough DoH Check -IPv6- on doh5004 is OK: HTTP OK: HTTP/1.1 200 OK - 595 bytes in 3.346 second response time https://wikitech.wikimedia.org/wiki/Wikidough/Monitoring%23Wikidough_Basic_Check [16:52:56] !log ladsgroup@deploy2003 helmfile [staging] START helmfile.d/services/thumbor: apply [16:53:23] !log ladsgroup@deploy2003 helmfile [staging] DONE helmfile.d/services/thumbor: apply [16:54:23] FIRING: JobUnavailable: Reduced availability for job wikidough in ops@eqsin - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:54:24] !log start full in-place reindex of cloudelastic cluster [16:54:25] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [16:59:15] !log stewards1001/stewards2001 - reboot for maintenance [16:59:17] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [16:59:22] RESOLVED: JobUnavailable: Reduced availability for job wikidough in ops@eqsin - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [17:00:05] bd808: That opportune time for a Cloud Services/Technical Documentation weekly deploy (Toolhub, Developer portal, Striker) deploy is upon us again. Don't be afraid. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260709T1700). [17:00:05] Deploy window MediaWiki infrastructure (UTC late) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260709T1700) [17:00:56] Nothing for my window today [17:01:55] (03PS12) 10Btullis: admin_ng: enable the topolvm CSI driver on dse-k8s [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305978 (https://phabricator.wikimedia.org/T429331) [17:01:55] (03PS1) 10Btullis: admin_ng: configure topolvm storage classes and node placement [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309261 (https://phabricator.wikimedia.org/T428235) [17:03:10] !log start full in-place reindex of codfw cirrussearch cluster [17:03:11] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [17:06:10] RECOVERY - SSH on doh5004 is OK: SSH OK - OpenSSH_9.2p1 Debian-2+deb12u9 (protocol 2.0) https://wikitech.wikimedia.org/wiki/SSH/monitoring [17:08:28] !log dzahn@cumin2002 START - Cookbook sre.hosts.reboot-cluster [17:08:29] !log dzahn@cumin2002 END (FAIL) - Cookbook sre.hosts.reboot-cluster (exit_code=99) [17:09:30] !log start full in-place reindex of eqiad cirrussearch cluster [17:09:31] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [17:09:51] !log zuul[12]00[123] - rebooting for maintenance [17:09:52] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [17:13:15] 06SRE, 06Infrastructure-Foundations, 10netops: Export network fixed-limit usage stats in gnmi - https://phabricator.wikimedia.org/T395998#12106342 (10cmooney) [17:14:01] 06SRE, 06Infrastructure-Foundations, 10netops: Export network fixed-limit usage stats in gnmi - https://phabricator.wikimedia.org/T395998#12106344 (10cmooney) The first of these is tracked in T428685 for Nokia devices. Leaving this one open for now to do the work for Juniper. [17:14:08] (03PS1) 10CDobbins: put testwiki back into enforce mode [puppet] - 10https://gerrit.wikimedia.org/r/1309263 [17:14:23] FIRING: [2x] JobUnavailable: Reduced availability for job atlas_exporter in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [17:15:14] (03PS2) 10CDobbins: varnish: put testwiki back into enforce mode [puppet] - 10https://gerrit.wikimedia.org/r/1309263 [17:17:48] (03CR) 10CDobbins: "@sbassett@wikimedia.org I've created two patches to address your feedback." [puppet] - 10https://gerrit.wikimedia.org/r/1307826 (owner: 10BCornwall) [17:19:23] FIRING: CertAlmostExpired: gNMI TLS certificate for lsw1-d6-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [17:23:16] RECOVERY - Wikidough DoH Check -IPv4- on doh5004 is OK: HTTP OK: HTTP/1.1 200 OK - 595 bytes in 0.990 second response time https://wikitech.wikimedia.org/wiki/Wikidough/Monitoring%23Wikidough_Basic_Check [17:24:23] RESOLVED: JobUnavailable: Reduced availability for job wikidough in ops@eqsin - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [17:24:23] 06SRE, 06Infrastructure-Foundations, 10netops: eqiad: pod CD switch upgrade (2026) - https://phabricator.wikimedia.org/T431745 (10cmooney) 03NEW p:05Triage→03High [17:25:21] 06SRE, 06Infrastructure-Foundations, 10netops: eqiad: pod CD switch upgrade (2026) - https://phabricator.wikimedia.org/T431745#12106371 (10cmooney) [17:25:52] 06SRE, 06Infrastructure-Foundations, 10netops, 06ServiceOps new: Nokia SR-Linux DHCP Relay Bug - https://phabricator.wikimedia.org/T411054#12106375 (10cmooney) [17:25:53] 06SRE, 06Infrastructure-Foundations, 10netops: eqiad: pod CD switch upgrade (2026) - https://phabricator.wikimedia.org/T431745#12106376 (10cmooney) [17:26:40] 06SRE, 06Infrastructure-Foundations, 10netops: Eqiad: move row-wide vlan gateways to Nokia switches - https://phabricator.wikimedia.org/T416872#12106378 (10cmooney) [17:26:42] 06SRE, 06Infrastructure-Foundations, 10netops: eqiad: pod CD switch upgrade (2026) - https://phabricator.wikimedia.org/T431745#12106379 (10cmooney) [17:26:58] !log ladsgroup@deploy2003 helmfile [codfw] START helmfile.d/services/thumbor: apply [17:26:58] 06SRE, 06Infrastructure-Foundations, 10netops: eqiad: pod CD switch upgrade (2026) - https://phabricator.wikimedia.org/T431745#12106381 (10cmooney) [17:26:59] 06SRE, 06Infrastructure-Foundations, 10netops: Eqiad C/D refresh: move legacy switch uplinks to Nokias and migrate Vlan GWs - https://phabricator.wikimedia.org/T405562#12106382 (10cmooney) [17:27:00] 06SRE, 06Infrastructure-Foundations, 10netops: Nokia SR-Linux - wonky routing with IPv6 RAs and EVPN Anycast GW - https://phabricator.wikimedia.org/T420706#12106380 (10cmooney) [17:27:22] 06SRE, 06Infrastructure-Foundations, 10netops: eqiad: pod CD switch upgrade (2026) - https://phabricator.wikimedia.org/T431745#12106384 (10cmooney) [17:27:23] 06SRE, 06Infrastructure-Foundations, 10netops: Blackbox probe for TLS cert expiry failing on multiple eqiad SR-Linux nodes - https://phabricator.wikimedia.org/T429242#12106383 (10cmooney) [17:27:34] 06SRE, 06Infrastructure-Foundations, 10netops: eqiad: pod CD switch upgrade (2026) - https://phabricator.wikimedia.org/T431745#12106387 (10cmooney) [17:27:56] 06SRE, 06Infrastructure-Foundations, 10netops: SR-Linux: applying analytics-in acl to irb sub-interface blocks ARP - https://phabricator.wikimedia.org/T429499#12106388 (10cmooney) [17:27:57] 06SRE, 06Infrastructure-Foundations, 10netops: eqiad: pod CD switch upgrade (2026) - https://phabricator.wikimedia.org/T431745#12106389 (10cmooney) [17:29:14] !log ladsgroup@deploy2003 helmfile [codfw] DONE helmfile.d/services/thumbor: apply [17:35:58] !log ladsgroup@deploy2003 helmfile [eqiad] START helmfile.d/services/thumbor: apply [17:38:29] (03PS3) 10CDobbins: put testwiki back into enforce mode [puppet] - 10https://gerrit.wikimedia.org/r/1309263 [17:38:38] !log ladsgroup@deploy2003 helmfile [eqiad] DONE helmfile.d/services/thumbor: apply [17:39:22] 06SRE, 06Infrastructure-Foundations, 10netops: Create netbox script to support moving a cable from one network port to another - https://phabricator.wikimedia.org/T355869#12106433 (10cmooney) 05Open→03Declined We've done without it thus far ok. If we need it again we can revisit. [17:39:23] RESOLVED: CertAlmostExpired: gNMI TLS certificate for lsw1-d6-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [17:42:12] PROBLEM - Host contint2003 is DOWN: PING CRITICAL - Packet loss = 100% [17:43:28] RECOVERY - Bird Internet Routing Daemon on doh5004 is OK: PROCS OK: 1 process with command name bird https://wikitech.wikimedia.org/wiki/Anycast%23Bird_daemon_not_running [17:43:32] RECOVERY - Host contint2003 is UP: PING OK - Packet loss = 0%, RTA = 31.65 ms [17:45:53] !log brett@cumin2002 START - Cookbook sre.hosts.remove-downtime for doh5004.wikimedia.org [17:45:54] !log brett@cumin2002 END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for doh5004.wikimedia.org [17:45:57] 06SRE, 06Infrastructure-Foundations, 10netops: Core router upgrades 2026 #2 - https://phabricator.wikimedia.org/T431748 (10cmooney) 03NEW p:05Triage→03Medium [17:47:11] (03PS1) 10Pppery: Add a bunch more locale files [phabricator/translations] (wmf/stable) - 10https://gerrit.wikimedia.org/r/1309266 (https://phabricator.wikimedia.org/T412651) [18:09:23] FIRING: JobUnavailable: Reduced availability for job atlas_exporter in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [18:14:23] RESOLVED: JobUnavailable: Reduced availability for job atlas_exporter in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [18:15:12] !log rzl@deploy2003 helmfile [staging-codfw] START helmfile.d/admin 'apply'. [18:16:55] !log rzl@deploy2003 helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. [18:41:40] !log rzl@deploy2003 helmfile [staging-eqiad] START helmfile.d/admin 'apply'. [18:42:56] !log rzl@deploy2003 helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. [18:57:43] FIRING: KeyholderUnarmed: 1 unarmed Keyholder key(s) on netmon2002:9100 - https://wikitech.wikimedia.org/wiki/Keyholder - TODO - https://alerts.wikimedia.org/?q=alertname%3DKeyholderUnarmed [19:14:23] FIRING: JobUnavailable: Reduced availability for job atlas_exporter in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [19:19:23] RESOLVED: JobUnavailable: Reduced availability for job atlas_exporter in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [19:25:44] !log rzl@deploy2003 helmfile [codfw] START helmfile.d/admin 'apply'. [19:27:07] !log rzl@deploy2003 helmfile [codfw] DONE helmfile.d/admin 'apply'. [19:28:52] !log rzl@deploy2003 helmfile [eqiad] START helmfile.d/admin 'apply'. [19:30:03] !log rzl@deploy2003 helmfile [eqiad] DONE helmfile.d/admin 'apply'. [19:33:38] PROBLEM - Host logging-sd2007 is DOWN: PING CRITICAL - Packet loss = 100% [19:35:04] RECOVERY - Host logging-sd2007 is UP: PING OK - Packet loss = 0%, RTA = 30.48 ms [19:43:28] !log mvernon@cumin2003 START - Cookbook sre.hosts.provision for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [19:46:50] PROBLEM - Host thanos-be1009 is DOWN: PING CRITICAL - Packet loss = 100% [19:52:22] RECOVERY - Host thanos-be1009 is UP: PING OK - Packet loss = 0%, RTA = 0.31 ms [19:54:14] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [19:55:18] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host thanos-be1009.eqiad.wmnet with OS trixie [20:00:05] RoanKattouw, urbanecm, TheresNoTime, kindrobot, and cjming: May I have your attention please! UTC late backport window. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260709T2000) [20:00:05] No Gerrit patches in the queue for this window AFAICS. [20:02:47] !log ladsgroup@cumin1003 START - Cookbook sre.wikireplicas.update-views [20:11:31] PROBLEM - Host logging-sd1001 is DOWN: PING CRITICAL - Packet loss = 100% [20:11:31] PROBLEM - Host logging-hd1001 is DOWN: PING CRITICAL - Packet loss = 100% [20:12:45] !log ladsgroup@cumin1003 END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) [20:12:59] RECOVERY - Host logging-hd1001 is UP: PING OK - Packet loss = 0%, RTA = 0.36 ms [20:12:59] RECOVERY - Host logging-sd1001 is UP: PING OK - Packet loss = 0%, RTA = 0.35 ms [20:13:17] ooh, no patches, I'm going to backport so hard [20:15:24] (03PS1) 10Ladsgroup: Introduce sharing of section function [core] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1309290 (https://phabricator.wikimedia.org/T18691) [20:15:42] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage [20:15:56] (03PS1) 10Ladsgroup: Drop the share icon from the mobile site [skins/MinervaNeue] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1309291 (https://phabricator.wikimedia.org/T18691) [20:16:19] !log ladsgroup@cumin1003 START - Cookbook sre.wikireplicas.update-views [20:16:28] 06SRE, 06Data-Engineering, 06Data-Platform-SRE (2026-07-03 - 2026-07-31), 07Sustainability (Incident Followup): The webrequest_sampled_live data pipeline and its query tools have become mission-critical and require re-engineering for resilience - https://phabricator.wikimedia.org/T431112#12106896 (10bd8... [20:18:52] (03CR) 10RLazarus: [C:03+2] wikifunctions: Enable custom dnsConfig pointing at coredns-internalonly [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305798 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [20:19:20] ladsgroup@cumin1003 update-views (PID 3291191) is awaiting input [20:20:06] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage [20:20:28] PROBLEM - Host logging-hd1002 is DOWN: PING CRITICAL - Packet loss = 100% [20:20:36] PROBLEM - Host logging-sd1002 is DOWN: PING CRITICAL - Packet loss = 100% [20:21:01] (03CR) 10TrainBranchBot: [C:03+2] "Approved by ladsgroup@deploy2003 using scap backport" [core] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1309290 (https://phabricator.wikimedia.org/T18691) (owner: 10Ladsgroup) [20:21:01] (03CR) 10TrainBranchBot: [C:03+2] "Approved by ladsgroup@deploy2003 using scap backport" [skins/MinervaNeue] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1309291 (https://phabricator.wikimedia.org/T18691) (owner: 10Ladsgroup) [20:21:07] (03Merged) 10jenkins-bot: wikifunctions: Enable custom dnsConfig pointing at coredns-internalonly [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305798 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [20:21:41] !log bking@cumin2003 START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - T431658 [20:21:49] !log bking@cumin2003 END (ERROR) - Cookbook sre.elasticsearch.rolling-operation (exit_code=97) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - T431658 [20:22:01] !log bking@cumin2003 START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - T431658 [20:22:03] !log bking@cumin2003 END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - T431658 [20:22:04] RECOVERY - Host logging-sd1002 is UP: PING OK - Packet loss = 0%, RTA = 0.79 ms [20:22:30] (03CR) 10SBassett: "Ok, I'll have a look. Will this change set be abandoned then?" [puppet] - 10https://gerrit.wikimedia.org/r/1307826 (owner: 10BCornwall) [20:22:58] RECOVERY - Host logging-hd1002 is UP: PING OK - Packet loss = 0%, RTA = 0.31 ms [20:23:08] !log bking@cumin2003 START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet [20:23:09] !log rzl@deploy2003 helmfile [staging] START helmfile.d/services/wikifunctions: apply [20:23:25] !log bking@cumin2003 END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host relforge1008.eqiad.wmnet [20:23:49] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS bookworm [20:24:42] !log rzl@deploy2003 helmfile [staging] DONE helmfile.d/services/wikifunctions: apply [20:29:39] !log bking@cumin2003 START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet [20:30:39] (03Merged) 10jenkins-bot: Introduce sharing of section function [core] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1309290 (https://phabricator.wikimedia.org/T18691) (owner: 10Ladsgroup) [20:30:44] (03CR) 10CI reject: [V:04-1] Drop the share icon from the mobile site [skins/MinervaNeue] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1309291 (https://phabricator.wikimedia.org/T18691) (owner: 10Ladsgroup) [20:31:39] !log ladsgroup@cumin1003 END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) [20:31:56] !log ladsgroup@cumin1003 START - Cookbook sre.wikireplicas.update-views [20:32:42] !log bking@cumin2003 START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - T431658 [20:32:44] !log bking@cumin2003 END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - T431658 [20:33:42] !log bking@cumin2003 START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - T431658 [20:33:50] (03CR) 10Bking: [C:03+2] Fix rolling-operation datetimes [cookbooks] - 10https://gerrit.wikimedia.org/r/1305766 (https://phabricator.wikimedia.org/T426862) (owner: 10Ryan Kemper) [20:36:34] (03PS5) 10Danielyepezgarces: extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1307836 (https://phabricator.wikimedia.org/T431334) [20:36:45] FIRING: SwiftLowContainerAvailability: Swift eqiad container availability low - https://wikitech.wikimedia.org/wiki/Swift/How_To - https://grafana.wikimedia.org/d/OPgmB1Eiz/swift?panelId=8&fullscreen&orgId=1&var-DC=eqiad - https://alerts.wikimedia.org/?q=alertname%3DSwiftLowContainerAvailability [20:36:45] FIRING: SwiftLowObjectAvailability: Swift eqiad object availability low - https://wikitech.wikimedia.org/wiki/Swift/How_To - https://grafana.wikimedia.org/d/OPgmB1Eiz/swift?panelId=8&fullscreen&orgId=1&var-DC=eqiad - https://alerts.wikimedia.org/?q=alertname%3DSwiftLowObjectAvailability [20:37:51] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1009.eqiad.wmnet with OS trixie [20:37:55] PROBLEM - OpenSearch health check for shards on 9200 on cloudelastic1008 is CRITICAL: CRITICAL - elasticsearch inactive shards 261 threshold =0.15 breach: cluster_name: cloudelastic-chi-eqiad, status: red, timed_out: False, number_of_nodes: 5, number_of_data_nodes: 5, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 792, active_shards: 1306, relocating_shards: 0, initializing_shards: 0, unassigned_shards: [20:37:55] ayed_unassigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 83.34396936821953 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:37:55] PROBLEM - OpenSearch health check for shards on 9400 on cloudelastic1008 is CRITICAL: CRITICAL - elasticsearch inactive shards 275 threshold =0.15 breach: cluster_name: cloudelastic-omega-eqiad, status: yellow, timed_out: False, number_of_nodes: 5, number_of_data_nodes: 5, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 827, active_shards: 1382, relocating_shards: 0, initializing_shards: 0, unassigned_sha [20:37:55] , delayed_unassigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 83.40374170187084 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:37:55] PROBLEM - OpenSearch health check for shards on 9400 on cloudelastic1007 is CRITICAL: CRITICAL - elasticsearch inactive shards 275 threshold =0.15 breach: cluster_name: cloudelastic-omega-eqiad, status: yellow, timed_out: False, number_of_nodes: 5, number_of_data_nodes: 5, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 827, active_shards: 1382, relocating_shards: 0, initializing_shards: 0, unassigned_sha [20:37:56] , delayed_unassigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 83.40374170187084 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:37:56] PROBLEM - OpenSearch health check for shards on 9600 on cloudelastic1007 is CRITICAL: CRITICAL - elasticsearch inactive shards 272 threshold =0.15 breach: cluster_name: cloudelastic-psi-eqiad, status: yellow, timed_out: False, number_of_nodes: 5, number_of_data_nodes: 5, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 817, active_shards: 1365, relocating_shards: 0, initializing_shards: 0, unassigned_shard [20:37:57] delayed_unassigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 83.38423946243128 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:37:57] PROBLEM - OpenSearch health check for shards on 9200 on cloudelastic1007 is CRITICAL: CRITICAL - elasticsearch inactive shards 261 threshold =0.15 breach: cluster_name: cloudelastic-chi-eqiad, status: red, timed_out: False, number_of_nodes: 5, number_of_data_nodes: 5, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 792, active_shards: 1306, relocating_shards: 0, initializing_shards: 0, unassigned_shards: [20:37:58] ayed_unassigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 83.34396936821953 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:37:58] PROBLEM - OpenSearch health check for shards on 9400 on cloudelastic1010 is CRITICAL: CRITICAL - elasticsearch inactive shards 275 threshold =0.15 breach: cluster_name: cloudelastic-omega-eqiad, status: yellow, timed_out: False, number_of_nodes: 5, number_of_data_nodes: 5, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 827, active_shards: 1382, relocating_shards: 0, initializing_shards: 0, unassigned_sha [20:37:59] , delayed_unassigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 83.40374170187084 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:37:59] PROBLEM - OpenSearch health check for shards on 9600 on cloudelastic1010 is CRITICAL: CRITICAL - elasticsearch inactive shards 272 threshold =0.15 breach: cluster_name: cloudelastic-psi-eqiad, status: yellow, timed_out: False, number_of_nodes: 5, number_of_data_nodes: 5, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 817, active_shards: 1365, relocating_shards: 0, initializing_shards: 0, unassigned_shard [20:38:00] delayed_unassigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 83.38423946243128 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:38:00] PROBLEM - OpenSearch health check for shards on 9600 on cloudelastic1008 is CRITICAL: CRITICAL - elasticsearch inactive shards 272 threshold =0.15 breach: cluster_name: cloudelastic-psi-eqiad, status: yellow, timed_out: False, number_of_nodes: 5, number_of_data_nodes: 5, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 817, active_shards: 1365, relocating_shards: 0, initializing_shards: 0, unassigned_shard [20:38:01] delayed_unassigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 83.38423946243128 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:38:01] PROBLEM - OpenSearch health check for shards on 9200 on cloudelastic1010 is CRITICAL: CRITICAL - elasticsearch inactive shards 261 threshold =0.15 breach: cluster_name: cloudelastic-chi-eqiad, status: red, timed_out: False, number_of_nodes: 5, number_of_data_nodes: 5, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 792, active_shards: 1306, relocating_shards: 0, initializing_shards: 0, unassigned_shards: [20:38:02] ayed_unassigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 83.34396936821953 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:38:02] PROBLEM - OpenSearch health check for shards on 9200 on cloudelastic1009 is CRITICAL: CRITICAL - elasticsearch inactive shards 261 threshold =0.15 breach: cluster_name: cloudelastic-chi-eqiad, status: red, timed_out: False, number_of_nodes: 5, number_of_data_nodes: 5, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 792, active_shards: 1306, relocating_shards: 0, initializing_shards: 0, unassigned_shards: [20:38:03] ayed_unassigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 83.34396936821953 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:38:03] PROBLEM - OpenSearch health check for shards on 9400 on cloudelastic1009 is CRITICAL: CRITICAL - elasticsearch inactive shards 275 threshold =0.15 breach: cluster_name: cloudelastic-omega-eqiad, status: yellow, timed_out: False, number_of_nodes: 5, number_of_data_nodes: 5, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 827, active_shards: 1382, relocating_shards: 0, initializing_shards: 0, unassigned_sha [20:38:04] , delayed_unassigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 83.40374170187084 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:38:04] PROBLEM - OpenSearch health check for shards on 9600 on cloudelastic1009 is CRITICAL: CRITICAL - elasticsearch inactive shards 272 threshold =0.15 breach: cluster_name: cloudelastic-psi-eqiad, status: yellow, timed_out: False, number_of_nodes: 5, number_of_data_nodes: 5, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 817, active_shards: 1365, relocating_shards: 0, initializing_shards: 0, unassigned_shard [20:38:04] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12106964 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host thanos-be1009.eqiad.wmnet with OS trixie compl... [20:38:05] delayed_unassigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 83.38423946243128 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:38:05] PROBLEM - OpenSearch health check for shards on 9200 on cloudelastic1011 is CRITICAL: CRITICAL - elasticsearch inactive shards 261 threshold =0.15 breach: cluster_name: cloudelastic-chi-eqiad, status: red, timed_out: False, number_of_nodes: 5, number_of_data_nodes: 5, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 792, active_shards: 1306, relocating_shards: 0, initializing_shards: 0, unassigned_shards: [20:38:06] ayed_unassigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 83.34396936821953 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:38:06] PROBLEM - OpenSearch health check for shards on 9600 on cloudelastic1011 is CRITICAL: CRITICAL - elasticsearch inactive shards 272 threshold =0.15 breach: cluster_name: cloudelastic-psi-eqiad, status: yellow, timed_out: False, number_of_nodes: 5, number_of_data_nodes: 5, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 817, active_shards: 1365, relocating_shards: 0, initializing_shards: 0, unassigned_shard [20:38:07] delayed_unassigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 83.38423946243128 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:38:07] PROBLEM - OpenSearch health check for shards on 9400 on cloudelastic1011 is CRITICAL: CRITICAL - elasticsearch inactive shards 275 threshold =0.15 breach: cluster_name: cloudelastic-omega-eqiad, status: yellow, timed_out: False, number_of_nodes: 5, number_of_data_nodes: 5, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 827, active_shards: 1382, relocating_shards: 0, initializing_shards: 0, unassigned_sha [20:38:08] , delayed_unassigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 83.40374170187084 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:39:55] RECOVERY - OpenSearch health check for shards on 9600 on cloudelastic1007 is OK: OK - elasticsearch status cloudelastic-psi-eqiad: cluster_name: cloudelastic-psi-eqiad, status: green, timed_out: False, number_of_nodes: 6, number_of_data_nodes: 6, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 817, active_shards: 1638, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, delayed_unassigned_ [20:39:55] 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:39:55] RECOVERY - OpenSearch health check for shards on 9400 on cloudelastic1007 is OK: OK - elasticsearch status cloudelastic-omega-eqiad: cluster_name: cloudelastic-omega-eqiad, status: green, timed_out: False, number_of_nodes: 6, number_of_data_nodes: 6, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 827, active_shards: 1658, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, delayed_unassig [20:39:55] ds: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:39:55] RECOVERY - OpenSearch health check for shards on 9200 on cloudelastic1007 is OK: OK - elasticsearch status cloudelastic-chi-eqiad: cluster_name: cloudelastic-chi-eqiad, status: yellow, timed_out: False, number_of_nodes: 6, number_of_data_nodes: 6, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 798, active_shards: 1557, relocating_shards: 0, initializing_shards: 6, unassigned_shards: 5, delayed_unassigned [20:39:56] 0, number_of_pending_tasks: 1, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 99.2984693877551 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:39:56] RECOVERY - OpenSearch health check for shards on 9400 on cloudelastic1008 is OK: OK - elasticsearch status cloudelastic-omega-eqiad: cluster_name: cloudelastic-omega-eqiad, status: green, timed_out: False, number_of_nodes: 6, number_of_data_nodes: 6, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 827, active_shards: 1658, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, delayed_unassig [20:39:57] ds: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:39:57] RECOVERY - OpenSearch health check for shards on 9200 on cloudelastic1008 is OK: OK - elasticsearch status cloudelastic-chi-eqiad: cluster_name: cloudelastic-chi-eqiad, status: yellow, timed_out: False, number_of_nodes: 6, number_of_data_nodes: 6, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 798, active_shards: 1557, relocating_shards: 0, initializing_shards: 6, unassigned_shards: 5, delayed_unassigned [20:39:58] 0, number_of_pending_tasks: 1, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 99.2984693877551 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:39:58] RECOVERY - OpenSearch health check for shards on 9400 on cloudelastic1010 is OK: OK - elasticsearch status cloudelastic-omega-eqiad: cluster_name: cloudelastic-omega-eqiad, status: green, timed_out: False, number_of_nodes: 6, number_of_data_nodes: 6, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 827, active_shards: 1658, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, delayed_unassig [20:39:59] ds: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:39:59] RECOVERY - OpenSearch health check for shards on 9600 on cloudelastic1010 is OK: OK - elasticsearch status cloudelastic-psi-eqiad: cluster_name: cloudelastic-psi-eqiad, status: green, timed_out: False, number_of_nodes: 6, number_of_data_nodes: 6, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 817, active_shards: 1638, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, delayed_unassigned_ [20:40:00] 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:40:00] RECOVERY - OpenSearch health check for shards on 9600 on cloudelastic1008 is OK: OK - elasticsearch status cloudelastic-psi-eqiad: cluster_name: cloudelastic-psi-eqiad, status: green, timed_out: False, number_of_nodes: 6, number_of_data_nodes: 6, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 817, active_shards: 1638, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, delayed_unassigned_ [20:40:01] 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:40:01] RECOVERY - OpenSearch health check for shards on 9200 on cloudelastic1010 is OK: OK - elasticsearch status cloudelastic-chi-eqiad: cluster_name: cloudelastic-chi-eqiad, status: yellow, timed_out: False, number_of_nodes: 6, number_of_data_nodes: 6, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 798, active_shards: 1565, relocating_shards: 0, initializing_shards: 3, unassigned_shards: 0, delayed_unassigned [20:40:02] 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 99.80867346938776 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:40:02] RECOVERY - OpenSearch health check for shards on 9200 on cloudelastic1009 is OK: OK - elasticsearch status cloudelastic-chi-eqiad: cluster_name: cloudelastic-chi-eqiad, status: yellow, timed_out: False, number_of_nodes: 6, number_of_data_nodes: 6, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 798, active_shards: 1565, relocating_shards: 0, initializing_shards: 3, unassigned_shards: 0, delayed_unassigned [20:40:03] 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 99.80867346938776 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:40:03] RECOVERY - OpenSearch health check for shards on 9400 on cloudelastic1009 is OK: OK - elasticsearch status cloudelastic-omega-eqiad: cluster_name: cloudelastic-omega-eqiad, status: green, timed_out: False, number_of_nodes: 6, number_of_data_nodes: 6, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 827, active_shards: 1658, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, delayed_unassig [20:40:04] ds: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:40:04] RECOVERY - OpenSearch health check for shards on 9200 on cloudelastic1011 is OK: OK - elasticsearch status cloudelastic-chi-eqiad: cluster_name: cloudelastic-chi-eqiad, status: green, timed_out: False, number_of_nodes: 6, number_of_data_nodes: 6, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 798, active_shards: 1568, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, delayed_unassigned_ [20:40:05] 0, number_of_pending_tasks: 1, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:40:05] RECOVERY - OpenSearch health check for shards on 9600 on cloudelastic1009 is OK: OK - elasticsearch status cloudelastic-psi-eqiad: cluster_name: cloudelastic-psi-eqiad, status: green, timed_out: False, number_of_nodes: 6, number_of_data_nodes: 6, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 817, active_shards: 1638, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, delayed_unassigned_ [20:40:06] 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:40:06] RECOVERY - OpenSearch health check for shards on 9600 on cloudelastic1011 is OK: OK - elasticsearch status cloudelastic-psi-eqiad: cluster_name: cloudelastic-psi-eqiad, status: green, timed_out: False, number_of_nodes: 6, number_of_data_nodes: 6, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 817, active_shards: 1638, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, delayed_unassigned_ [20:40:07] 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:40:07] RECOVERY - OpenSearch health check for shards on 9400 on cloudelastic1011 is OK: OK - elasticsearch status cloudelastic-omega-eqiad: cluster_name: cloudelastic-omega-eqiad, status: green, timed_out: False, number_of_nodes: 6, number_of_data_nodes: 6, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 827, active_shards: 1658, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, delayed_unassig [20:40:08] ds: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:40:09] !log bking@cumin2003 END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - T431658 [20:40:17] !log bking@cumin2003 END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1008.eqiad.wmnet [20:41:27] !log ladsgroup@cumin1003 END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) [20:41:36] !log ladsgroup@cumin1003 START - Cookbook sre.wikireplicas.update-views [20:42:41] (03CR) 10Ladsgroup: [C:03+2] Drop the share icon from the mobile site [skins/MinervaNeue] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1309291 (https://phabricator.wikimedia.org/T18691) (owner: 10Ladsgroup) [20:42:53] PROBLEM - Host logging-hd1003 is DOWN: PING CRITICAL - Packet loss = 100% [20:42:55] !log bking@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage [20:43:33] PROBLEM - Host logging-sd1003 is DOWN: PING CRITICAL - Packet loss = 100% [20:45:05] RECOVERY - Host logging-sd1003 is UP: PING OK - Packet loss = 0%, RTA = 0.38 ms [20:45:05] RECOVERY - Host logging-hd1003 is UP: PING OK - Packet loss = 0%, RTA = 0.29 ms [20:45:11] (03CR) 10SBassett: varnish: update CSP report-only header (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1308772 (owner: 10CDobbins) [20:47:34] !log bking@cumin2003 START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 [20:47:59] (03Merged) 10jenkins-bot: Drop the share icon from the mobile site [skins/MinervaNeue] (wmf/1.47.0-wmf.10) - 10https://gerrit.wikimedia.org/r/1309291 (https://phabricator.wikimedia.org/T18691) (owner: 10Ladsgroup) [20:50:02] (03PS4) 10RLazarus: admin_ng: Grant network access to kube-dns-internalonly [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305799 (https://phabricator.wikimedia.org/T427864) [20:50:02] (03PS1) 10RLazarus: admin_ng: Restrict wikifunctions network access to DNS [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309293 (https://phabricator.wikimedia.org/T427864) [20:50:02] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage [20:50:54] PROBLEM - OpenSearch health check for shards on 9400 on cloudelastic1007 is CRITICAL: CRITICAL - elasticsearch inactive shards 276 threshold =0.15 breach: cluster_name: cloudelastic-omega-eqiad, status: yellow, timed_out: False, number_of_nodes: 5, number_of_data_nodes: 5, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 827, active_shards: 1381, relocating_shards: 0, initializing_shards: 0, unassigned_sha [20:50:54] , delayed_unassigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 83.34339167169583 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:50:54] PROBLEM - OpenSearch health check for shards on 9600 on cloudelastic1007 is CRITICAL: CRITICAL - elasticsearch inactive shards 272 threshold =0.15 breach: cluster_name: cloudelastic-psi-eqiad, status: yellow, timed_out: False, number_of_nodes: 5, number_of_data_nodes: 5, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 817, active_shards: 1365, relocating_shards: 0, initializing_shards: 0, unassigned_shard [20:50:54] delayed_unassigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 83.38423946243128 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:50:54] PROBLEM - OpenSearch health check for shards on 9200 on cloudelastic1007 is CRITICAL: CRITICAL - elasticsearch inactive shards 293 threshold =0.15 breach: cluster_name: cloudelastic-chi-eqiad, status: red, timed_out: False, number_of_nodes: 5, number_of_data_nodes: 5, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 826, active_shards: 1340, relocating_shards: 0, initializing_shards: 0, unassigned_shards: [20:50:55] ayed_unassigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 82.05756276791182 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:50:56] PROBLEM - OpenSearch health check for shards on 9400 on cloudelastic1010 is CRITICAL: CRITICAL - elasticsearch inactive shards 276 threshold =0.15 breach: cluster_name: cloudelastic-omega-eqiad, status: yellow, timed_out: False, number_of_nodes: 5, number_of_data_nodes: 5, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 827, active_shards: 1381, relocating_shards: 0, initializing_shards: 0, unassigned_sha [20:50:56] , delayed_unassigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 83.34339167169583 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:50:56] PROBLEM - OpenSearch health check for shards on 9600 on cloudelastic1010 is CRITICAL: CRITICAL - elasticsearch inactive shards 272 threshold =0.15 breach: cluster_name: cloudelastic-psi-eqiad, status: yellow, timed_out: False, number_of_nodes: 5, number_of_data_nodes: 5, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 817, active_shards: 1365, relocating_shards: 0, initializing_shards: 0, unassigned_shard [20:50:57] delayed_unassigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 83.38423946243128 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:50:58] PROBLEM - OpenSearch health check for shards on 9200 on cloudelastic1010 is CRITICAL: CRITICAL - elasticsearch inactive shards 293 threshold =0.15 breach: cluster_name: cloudelastic-chi-eqiad, status: red, timed_out: False, number_of_nodes: 5, number_of_data_nodes: 5, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 826, active_shards: 1340, relocating_shards: 0, initializing_shards: 0, unassigned_shards: [20:50:58] ayed_unassigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 82.05756276791182 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:50:58] PROBLEM - OpenSearch health check for shards on 9200 on cloudelastic1009 is CRITICAL: CRITICAL - elasticsearch inactive shards 293 threshold =0.15 breach: cluster_name: cloudelastic-chi-eqiad, status: red, timed_out: False, number_of_nodes: 5, number_of_data_nodes: 5, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 826, active_shards: 1340, relocating_shards: 0, initializing_shards: 0, unassigned_shards: [20:50:59] ayed_unassigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 82.05756276791182 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:50:59] PROBLEM - OpenSearch health check for shards on 9400 on cloudelastic1009 is CRITICAL: CRITICAL - elasticsearch inactive shards 276 threshold =0.15 breach: cluster_name: cloudelastic-omega-eqiad, status: yellow, timed_out: False, number_of_nodes: 5, number_of_data_nodes: 5, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 827, active_shards: 1381, relocating_shards: 0, initializing_shards: 0, unassigned_sha [20:51:00] , delayed_unassigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 83.34339167169583 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:51:00] PROBLEM - OpenSearch health check for shards on 9200 on cloudelastic1011 is CRITICAL: CRITICAL - elasticsearch inactive shards 293 threshold =0.15 breach: cluster_name: cloudelastic-chi-eqiad, status: red, timed_out: False, number_of_nodes: 5, number_of_data_nodes: 5, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 826, active_shards: 1340, relocating_shards: 0, initializing_shards: 0, unassigned_shards: [20:51:01] ayed_unassigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 82.05756276791182 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:51:01] PROBLEM - OpenSearch health check for shards on 9400 on cloudelastic1011 is CRITICAL: CRITICAL - elasticsearch inactive shards 276 threshold =0.15 breach: cluster_name: cloudelastic-omega-eqiad, status: yellow, timed_out: False, number_of_nodes: 5, number_of_data_nodes: 5, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 827, active_shards: 1381, relocating_shards: 0, initializing_shards: 0, unassigned_sha [20:51:02] , delayed_unassigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 83.34339167169583 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:51:02] PROBLEM - OpenSearch health check for shards on 9600 on cloudelastic1011 is CRITICAL: CRITICAL - elasticsearch inactive shards 272 threshold =0.15 breach: cluster_name: cloudelastic-psi-eqiad, status: yellow, timed_out: False, number_of_nodes: 5, number_of_data_nodes: 5, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 817, active_shards: 1365, relocating_shards: 0, initializing_shards: 0, unassigned_shard [20:51:03] delayed_unassigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 83.38423946243128 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:51:03] PROBLEM - OpenSearch health check for shards on 9600 on cloudelastic1009 is CRITICAL: CRITICAL - elasticsearch inactive shards 272 threshold =0.15 breach: cluster_name: cloudelastic-psi-eqiad, status: yellow, timed_out: False, number_of_nodes: 5, number_of_data_nodes: 5, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 817, active_shards: 1365, relocating_shards: 0, initializing_shards: 0, unassigned_shard [20:51:04] delayed_unassigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 83.38423946243128 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:51:06] !log ladsgroup@cumin1003 END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99) [20:51:49] (03CR) 10SBassett: varnish: put testwiki back into enforce mode (033 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1309263 (owner: 10CDobbins) [20:51:54] RESOLVED: SwiftLowContainerAvailability: Swift eqiad container availability low - https://wikitech.wikimedia.org/wiki/Swift/How_To - https://grafana.wikimedia.org/d/OPgmB1Eiz/swift?panelId=8&fullscreen&orgId=1&var-DC=eqiad - https://alerts.wikimedia.org/?q=alertname%3DSwiftLowContainerAvailability [20:51:59] RESOLVED: SwiftLowObjectAvailability: Swift eqiad object availability low - https://wikitech.wikimedia.org/wiki/Swift/How_To - https://grafana.wikimedia.org/d/OPgmB1Eiz/swift?panelId=8&fullscreen&orgId=1&var-DC=eqiad - https://alerts.wikimedia.org/?q=alertname%3DSwiftLowObjectAvailability [20:52:54] (03CR) 10SBassett: varnish: put testwiki back into enforce mode (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1309263 (owner: 10CDobbins) [20:53:13] !log ladsgroup@deploy2003 Started scap sync-world: Backport for [[gerrit:1309290|Introduce sharing of section function (T18691)]], [[gerrit:1309291|Drop the share icon from the mobile site (T18691)]] [20:53:17] T18691: RFC: Section header "share" link - https://phabricator.wikimedia.org/T18691 [20:53:17] !log ladsgroup@cumin1003 START - Cookbook sre.wikireplicas.update-views [20:53:54] RECOVERY - OpenSearch health check for shards on 9600 on cloudelastic1007 is OK: OK - elasticsearch status cloudelastic-psi-eqiad: cluster_name: cloudelastic-psi-eqiad, status: yellow, timed_out: False, number_of_nodes: 6, number_of_data_nodes: 6, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 817, active_shards: 1608, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 30, delayed_unassigne [20:53:54] : 0, number_of_pending_tasks: 1, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 98.16849816849816 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:53:54] RECOVERY - OpenSearch health check for shards on 9400 on cloudelastic1007 is OK: OK - elasticsearch status cloudelastic-omega-eqiad: cluster_name: cloudelastic-omega-eqiad, status: yellow, timed_out: False, number_of_nodes: 6, number_of_data_nodes: 6, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 827, active_shards: 1615, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 43, delayed_unass [20:53:54] ards: 0, number_of_pending_tasks: 1, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 97.40651387213511 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:53:54] RECOVERY - OpenSearch health check for shards on 9200 on cloudelastic1007 is OK: OK - elasticsearch status cloudelastic-chi-eqiad: cluster_name: cloudelastic-chi-eqiad, status: yellow, timed_out: False, number_of_nodes: 6, number_of_data_nodes: 6, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 831, active_shards: 1480, relocating_shards: 0, initializing_shards: 4, unassigned_shards: 118, delayed_unassign [20:53:55] s: 0, number_of_pending_tasks: 2, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 66, active_shards_percent_as_number: 92.38451935081149 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:53:56] RECOVERY - OpenSearch health check for shards on 9600 on cloudelastic1010 is OK: OK - elasticsearch status cloudelastic-psi-eqiad: cluster_name: cloudelastic-psi-eqiad, status: yellow, timed_out: False, number_of_nodes: 6, number_of_data_nodes: 6, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 817, active_shards: 1620, relocating_shards: 0, initializing_shards: 4, unassigned_shards: 14, delayed_unassigne [20:53:56] : 0, number_of_pending_tasks: 2, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 38, active_shards_percent_as_number: 98.9010989010989 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:53:56] RECOVERY - OpenSearch health check for shards on 9400 on cloudelastic1010 is OK: OK - elasticsearch status cloudelastic-omega-eqiad: cluster_name: cloudelastic-omega-eqiad, status: yellow, timed_out: False, number_of_nodes: 6, number_of_data_nodes: 6, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 827, active_shards: 1631, relocating_shards: 0, initializing_shards: 7, unassigned_shards: 20, delayed_unass [20:53:57] ards: 0, number_of_pending_tasks: 6, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 65, active_shards_percent_as_number: 98.37153196622437 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:53:58] RECOVERY - OpenSearch health check for shards on 9200 on cloudelastic1009 is OK: OK - elasticsearch status cloudelastic-chi-eqiad: cluster_name: cloudelastic-chi-eqiad, status: yellow, timed_out: False, number_of_nodes: 6, number_of_data_nodes: 6, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 831, active_shards: 1494, relocating_shards: 0, initializing_shards: 7, unassigned_shards: 101, delayed_unassign [20:53:58] s: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 93.25842696629213 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:53:58] RECOVERY - OpenSearch health check for shards on 9200 on cloudelastic1010 is OK: OK - elasticsearch status cloudelastic-chi-eqiad: cluster_name: cloudelastic-chi-eqiad, status: yellow, timed_out: False, number_of_nodes: 6, number_of_data_nodes: 6, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 831, active_shards: 1495, relocating_shards: 0, initializing_shards: 6, unassigned_shards: 101, delayed_unassign [20:53:59] s: 0, number_of_pending_tasks: 1, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 93.32084893882646 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:53:59] RECOVERY - OpenSearch health check for shards on 9400 on cloudelastic1009 is OK: OK - elasticsearch status cloudelastic-omega-eqiad: cluster_name: cloudelastic-omega-eqiad, status: green, timed_out: False, number_of_nodes: 6, number_of_data_nodes: 6, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 827, active_shards: 1658, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, delayed_unassig [20:54:00] ds: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:54:00] RECOVERY - OpenSearch health check for shards on 9200 on cloudelastic1011 is OK: OK - elasticsearch status cloudelastic-chi-eqiad: cluster_name: cloudelastic-chi-eqiad, status: yellow, timed_out: False, number_of_nodes: 6, number_of_data_nodes: 6, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 831, active_shards: 1496, relocating_shards: 0, initializing_shards: 6, unassigned_shards: 100, delayed_unassign [20:54:01] s: 0, number_of_pending_tasks: 1, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 93.3832709113608 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:54:01] RECOVERY - OpenSearch health check for shards on 9600 on cloudelastic1011 is OK: OK - elasticsearch status cloudelastic-psi-eqiad: cluster_name: cloudelastic-psi-eqiad, status: green, timed_out: False, number_of_nodes: 6, number_of_data_nodes: 6, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 817, active_shards: 1638, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, delayed_unassigned_ [20:54:02] 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:54:02] RECOVERY - OpenSearch health check for shards on 9600 on cloudelastic1009 is OK: OK - elasticsearch status cloudelastic-psi-eqiad: cluster_name: cloudelastic-psi-eqiad, status: green, timed_out: False, number_of_nodes: 6, number_of_data_nodes: 6, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 817, active_shards: 1638, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, delayed_unassigned_ [20:54:03] 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:54:03] RECOVERY - OpenSearch health check for shards on 9400 on cloudelastic1011 is OK: OK - elasticsearch status cloudelastic-omega-eqiad: cluster_name: cloudelastic-omega-eqiad, status: green, timed_out: False, number_of_nodes: 6, number_of_data_nodes: 6, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 827, active_shards: 1658, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, delayed_unassig [20:54:04] ds: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:54:43] !log bking@cumin2003 END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003 [20:55:37] inflatador: if those icinga alerts are expected when you run those reboots, do you think you can downtime em? :) [21:00:05] Deploy window Readers deployment window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260709T2100) [21:01:08] is anyone deploying? I would like to deploy to private settings [21:01:50] rzl oops, there is a problem with this cookbook I didn't notice [21:02:17] Will set a downtime manually [21:02:22] thanks <3 [21:03:05] np, thanks for the call-out. I need to check pybal now too ;) [21:07:28] !log bking@cumin2003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 6 hosts with reason: reboots [21:07:38] PROBLEM - Host logging-sd1004 is DOWN: PING CRITICAL - Packet loss = 100% [21:08:05] !log ladsgroup@cumin1003 END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0) [21:08:42] preparing to run scap [21:09:06] RECOVERY - Host logging-sd1004 is UP: PING OK - Packet loss = 0%, RTA = 4.01 ms [21:09:23] FIRING: JobUnavailable: Reduced availability for job atlas_exporter in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [21:10:48] oh someone is doing a backport [21:11:03] !log ladsgroup@deploy2003 ladsgroup: Backport for [[gerrit:1309290|Introduce sharing of section function (T18691)]], [[gerrit:1309291|Drop the share icon from the mobile site (T18691)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [21:11:06] T18691: RFC: Section header "share" link - https://phabricator.wikimedia.org/T18691 [21:11:25] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [21:13:31] maryum, Amir1: highlighting you for each other :) [21:13:47] rzl: thank you! [21:13:49] ah thanks [21:13:55] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2002.codfw.wmnet with OS bookworm [21:13:59] sorry. My deploy took way longer for various reasons [21:13:59] looks like you're wrapping up soon Amir1 [21:14:04] yup yup [21:14:23] RESOLVED: JobUnavailable: Reduced availability for job atlas_exporter in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [21:15:10] !log ladsgroup@deploy2003 ladsgroup: Continuing with deployment [21:16:19] !log bking@cumin2003 START - Cookbook sre.wdqs.data-transfer (T430879, restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards [21:16:22] T430879: Migrate WCQS to Bookworm or later - https://phabricator.wikimedia.org/T430879 [21:16:28] !log bking@cumin2003 END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) (T430879, restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards [21:20:17] !log bking@deploy2003 Started deploy [wdqs/wdqs@e8fb00c] (wcqs): T430879 [21:20:23] !log bking@deploy2003 Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): T430879 (duration: 00m 22s) [21:22:00] !log bking@cumin2003 START - Cookbook sre.wdqs.data-transfer (T430879, restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards [21:22:03] T430879: Migrate WCQS to Bookworm or later - https://phabricator.wikimedia.org/T430879 [21:23:20] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie [21:23:42] !log bking@cumin2003 START - Cookbook sre.hosts.move-vlan for host wcqs1002 [21:23:42] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1002 [21:27:09] > Running for 34m 29s [21:27:27] !log ladsgroup@deploy2003 Finished scap sync-world: Backport for [[gerrit:1309290|Introduce sharing of section function (T18691)]], [[gerrit:1309291|Drop the share icon from the mobile site (T18691)]] (duration: 34m 14s) [21:27:32] T18691: RFC: Section header "share" link - https://phabricator.wikimedia.org/T18691 [21:27:35] https://media3.giphy.com/media/v1.Y2lkPTc5MGI3NjExbHJ1cTBlamg1MHdzMHNpd3d3bHBudmMxbGZvejlhemRnYXVqMmQyYSZlcD12MV9pbnRlcm5hbF9naWZfYnlfaWQmY3Q9Zw/OpPyw0U5IGZDog5K4U/giphy.gif [21:27:42] okay, finally it's done [21:27:54] (03CR) 10Scott French: [C:03+1] "As discussed out of band, narrowing this to _just_ an additive step that enables access to kube-dns-internalonly seems like the right path" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305799 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [21:27:59] maryum: feel free to go ahead. I have a couple other changes to push but they can wai [21:28:00] (03PS3) 10Bking: cumin: Fix/simplify cirrussearch aliases [puppet] - 10https://gerrit.wikimedia.org/r/1309288 (https://phabricator.wikimedia.org/T431484) [21:28:01] *wait [21:28:29] Amir1 thank you! [21:28:34] I need to take a break too [21:28:52] just doing one scap run and should be done very shortly [21:29:31] (03CR) 10RLazarus: [C:03+2] admin_ng: Grant network access to kube-dns-internalonly [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305799 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [21:34:23] FIRING: JobUnavailable: Reduced availability for job jmx_wcqs_blazegraph in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [21:38:00] (03Merged) 10jenkins-bot: admin_ng: Grant network access to kube-dns-internalonly [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305799 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [21:38:12] !log bking@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage [21:42:06] !log Deploy fix for T431684 [21:42:07] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [21:42:11] scap run finished Amir1 [21:43:05] !log rzl@deploy2003 helmfile [staging-codfw] START helmfile.d/admin 'apply'. [21:43:47] !log rzl@deploy2003 helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. [21:44:22] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage [21:45:46] !log rzl@deploy2003 helmfile [staging-eqiad] START helmfile.d/admin 'apply'. [21:46:50] (03PS4) 10Bking: Cirrussearch: Replace temp dc-specific migration hieradata w/common hieradata [puppet] - 10https://gerrit.wikimedia.org/r/1309270 (https://phabricator.wikimedia.org/T431311) [21:46:59] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1309270 (https://phabricator.wikimedia.org/T431311) (owner: 10Bking) [21:46:59] !log rzl@deploy2003 helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. [21:53:06] !log rzl@deploy2003 helmfile [codfw] START helmfile.d/admin 'apply'. [21:53:10] !log jasmine@cumin2002 START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1165.eqiad.wmnet [21:53:14] !log jasmine@cumin2002 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1165.eqiad.wmnet [21:53:21] PROBLEM - Host logging-sd1006 is DOWN: PING CRITICAL - Packet loss = 100% [21:53:38] !log rzl@deploy2003 helmfile [codfw] DONE helmfile.d/admin 'apply'. [21:53:59] RECOVERY - Host logging-sd1006 is UP: PING OK - Packet loss = 0%, RTA = 0.30 ms [21:54:23] !log jasmine@cumin2002 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1165.eqiad.wmnet [21:54:46] !log rzl@deploy2003 helmfile [eqiad] START helmfile.d/admin 'apply'. [21:54:48] !log jasmine@cumin2002 START - Cookbook sre.hosts.reimage for host wikikube-worker1165.eqiad.wmnet with OS trixie [21:55:09] !log jasmine@cumin2002 START - Cookbook sre.hosts.move-vlan for host wikikube-worker1165 [21:57:34] !log jasmine@cumin2002 START - Cookbook sre.dns.netbox [22:02:02] !log rzl@deploy2003 helmfile [eqiad] DONE helmfile.d/admin 'apply'. [22:02:13] !log jasmine@cumin2002 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" [22:02:18] !log jasmine@cumin2002 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002" [22:02:19] !log jasmine@cumin2002 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [22:02:19] !log jasmine@cumin2002 START - Cookbook sre.dns.wipe-cache wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [22:02:23] !log jasmine@cumin2002 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [22:02:23] !log jasmine@cumin2002 START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1165 [22:02:48] !log jasmine@cumin2002 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1165 [22:02:48] !log jasmine@cumin2002 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1165 [22:04:17] !log rzl@deploy2003 helmfile [codfw] START helmfile.d/services/wikifunctions: apply [22:04:53] !log rzl@deploy2003 helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply [22:10:45] FIRING: CirrusStreamingUpdaterUnknownErrors: CirrusSearch consumer-search@codfw is failing write requests because of unknown errors - https://wikitech.wikimedia.org/wiki/Search#Streaming_Updater - https://grafana.wikimedia.org/d/jKqki4MSk/cirrus-streaming-updater - https://alerts.wikimedia.org/?q=alertname%3DCirrusStreamingUpdaterUnknownErrors [22:12:16] !log rzl@deploy2003 helmfile [eqiad] START helmfile.d/services/wikifunctions: apply [22:13:01] !log rzl@deploy2003 helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply [22:15:13] (03PS1) 10Danielyepezgarces: Enable campaignEvents on cowikimedia [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309302 (https://phabricator.wikimedia.org/T431765) [22:16:59] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, July 13 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309302 (https://phabricator.wikimedia.org/T431765) (owner: 10Danielyepezgarces) [22:17:05] (03PS5) 10Ryan Kemper: Cirrussearch: Replace temp dc-specific migration hieradata w/common hieradata [puppet] - 10https://gerrit.wikimedia.org/r/1309270 (https://phabricator.wikimedia.org/T431311) (owner: 10Bking) [22:17:10] (03PS4) 10Ryan Kemper: cumin: Fix/simplify cirrussearch aliases [puppet] - 10https://gerrit.wikimedia.org/r/1309288 (https://phabricator.wikimedia.org/T431484) (owner: 10Bking) [22:17:18] !log jasmine@cumin2002 START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage [22:18:43] (03CR) 10Ryan Kemper: "check puppet" [puppet] - 10https://gerrit.wikimedia.org/r/1309288 (https://phabricator.wikimedia.org/T431484) (owner: 10Bking) [22:18:58] (03CR) 10Ryan Kemper: "check puppet" [puppet] - 10https://gerrit.wikimedia.org/r/1309270 (https://phabricator.wikimedia.org/T431311) (owner: 10Bking) [22:20:45] RESOLVED: CirrusStreamingUpdaterUnknownErrors: CirrusSearch consumer-search@codfw is failing write requests because of unknown errors - https://wikitech.wikimedia.org/wiki/Search#Streaming_Updater - https://grafana.wikimedia.org/d/jKqki4MSk/cirrus-streaming-updater - https://alerts.wikimedia.org/?q=alertname%3DCirrusStreamingUpdaterUnknownErrors [22:22:03] (03CR) 10Ryan Kemper: [C:03+1] Cirrussearch: Replace temp dc-specific migration hieradata w/common hieradata [puppet] - 10https://gerrit.wikimedia.org/r/1309270 (https://phabricator.wikimedia.org/T431311) (owner: 10Bking) [22:24:37] 06SRE, 06MediaWiki-Media-Platform-Team, 06ServiceOps new, 10ServiceOps-Services-Oids, 10Thumbor: Thumbor-k8s performance improvements - https://phabricator.wikimedia.org/T333445#12107169 (10Ladsgroup) Cormac suggested we close this ticket since it has been a while and a lot has changed since then. I agree. [22:24:45] (03CR) 10Ryan Kemper: [C:03+1] cumin: Fix/simplify cirrussearch aliases [puppet] - 10https://gerrit.wikimedia.org/r/1309288 (https://phabricator.wikimedia.org/T431484) (owner: 10Bking) [22:25:07] !log jasmine@cumin2002 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage [22:31:43] (03PS3) 10Scott French: Consolidate and modernize test dependencies and entrypoint [software/etcd-mirror] - 10https://gerrit.wikimedia.org/r/1309303 (https://phabricator.wikimedia.org/T418915) [22:36:58] !log bking@cumin2003 END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) (T430879, restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards [22:37:02] T430879: Migrate WCQS to Bookworm or later - https://phabricator.wikimedia.org/T430879 [22:37:24] !log arlolra@deploy2003 helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply [22:37:54] !log arlolra@deploy2003 helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply [22:37:55] !log arlolra@deploy2003 helmfile [codfw] START helmfile.d/services/mw-parsoid: apply [22:38:28] !log arlolra@deploy2003 helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply [22:39:23] RESOLVED: JobUnavailable: Reduced availability for job jmx_wcqs_blazegraph in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [22:43:12] (03CR) 10RLazarus: "I split this out from Icffd4702 -- see https://phabricator.wikimedia.org/T427864#12107156." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309293 (https://phabricator.wikimedia.org/T427864) (owner: 10RLazarus) [22:45:12] !log jasmine@cumin2002 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1165.eqiad.wmnet with OS trixie [22:48:12] jasmine@cumin2002 renumber-node (PID 1017065) is awaiting input [22:51:45] FIRING: CirrusStreamingUpdaterUnknownErrors: CirrusSearch consumer-search@codfw is failing write requests because of unknown errors - https://wikitech.wikimedia.org/wiki/Search#Streaming_Updater - https://grafana.wikimedia.org/d/jKqki4MSk/cirrus-streaming-updater - https://alerts.wikimedia.org/?q=alertname%3DCirrusStreamingUpdaterUnknownErrors [22:56:56] !log jasmine@cumin2002 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1165.eqiad.wmnet [22:56:58] !log jasmine@cumin2002 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1165.eqiad.wmnet [22:57:00] !log jasmine@cumin2002 END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1165.eqiad.wmnet [22:57:43] FIRING: KeyholderUnarmed: 1 unarmed Keyholder key(s) on netmon2002:9100 - https://wikitech.wikimedia.org/wiki/Keyholder - TODO - https://alerts.wikimedia.org/?q=alertname%3DKeyholderUnarmed [22:59:02] (03PS1) 10Ladsgroup: beta: Enable section share on Persian Wikipedia [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309307 [23:01:39] (03CR) 10Ladsgroup: [C:03+2] beta: Enable section share on Persian Wikipedia [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309307 (owner: 10Ladsgroup) [23:01:45] RESOLVED: CirrusStreamingUpdaterUnknownErrors: CirrusSearch consumer-search@codfw is failing write requests because of unknown errors - https://wikitech.wikimedia.org/wiki/Search#Streaming_Updater - https://grafana.wikimedia.org/d/jKqki4MSk/cirrus-streaming-updater - https://alerts.wikimedia.org/?q=alertname%3DCirrusStreamingUpdaterUnknownErrors [23:02:34] (03Merged) 10jenkins-bot: beta: Enable section share on Persian Wikipedia [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309307 (owner: 10Ladsgroup) [23:14:15] FIRING: CirrusStreamingUpdaterUnknownErrors: CirrusSearch consumer-search@codfw is failing write requests because of unknown errors - https://wikitech.wikimedia.org/wiki/Search#Streaming_Updater - https://grafana.wikimedia.org/d/jKqki4MSk/cirrus-streaming-updater - https://alerts.wikimedia.org/?q=alertname%3DCirrusStreamingUpdaterUnknownErrors [23:18:01] (03PS2) 10Ladsgroup: Enable section share on Persian Wikipedia [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308671 (https://phabricator.wikimedia.org/T431514) (owner: 10Jdlrobson) [23:18:18] jouncebot: nowandnext [23:18:19] No deployments scheduled for the next 6 hour(s) and 41 minute(s) [23:18:19] In 6 hour(s) and 41 minute(s): MediaWiki infrastructure (UTC early) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260710T0600) [23:18:34] (03CR) 10TrainBranchBot: [C:03+2] "Approved by ladsgroup@deploy2003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308671 (https://phabricator.wikimedia.org/T431514) (owner: 10Jdlrobson) [23:20:18] (03Merged) 10jenkins-bot: Enable section share on Persian Wikipedia [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1308671 (https://phabricator.wikimedia.org/T431514) (owner: 10Jdlrobson) [23:20:28] !log ladsgroup@deploy2003 Started scap sync-world: Backport for [[gerrit:1308671|Enable section share on Persian Wikipedia (T431514)]] [23:20:31] T431514: Run an A/B test on section share link for logged out users - https://phabricator.wikimedia.org/T431514 [23:22:19] !log ladsgroup@deploy2003 ladsgroup, jdlrobson: Backport for [[gerrit:1308671|Enable section share on Persian Wikipedia (T431514)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [23:24:15] RESOLVED: CirrusStreamingUpdaterUnknownErrors: CirrusSearch consumer-search@codfw is failing write requests because of unknown errors - https://wikitech.wikimedia.org/wiki/Search#Streaming_Updater - https://grafana.wikimedia.org/d/jKqki4MSk/cirrus-streaming-updater - https://alerts.wikimedia.org/?q=alertname%3DCirrusStreamingUpdaterUnknownErrors [23:29:35] !log ladsgroup@deploy2003 ladsgroup, jdlrobson: Continuing with deployment [23:30:05] (03Abandoned) 10TrainBranchBot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1308790 (owner: 10TrainBranchBot) [23:33:53] !log ladsgroup@deploy2003 Finished scap sync-world: Backport for [[gerrit:1308671|Enable section share on Persian Wikipedia (T431514)]] (duration: 13m 26s) [23:33:57] T431514: Run an A/B test on section share link for logged out users - https://phabricator.wikimedia.org/T431514 [23:42:32] (03PS1) 10TrainBranchBot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1309309 [23:42:32] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1309309 (owner: 10TrainBranchBot) [23:50:18] (03Merged) 10jenkins-bot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1309309 (owner: 10TrainBranchBot)