[00:00:15] yes i'll kick off a test run - need to turn on the instrument first [00:01:10] cool -- happy to start that for you, or happy to leave it to you [00:01:50] (and, of course! thanks for your flexibility with the timing) [00:02:18] would you mind starting? i'm unclear what the exact cmd should be [00:02:28] sure, let me know when you're ready [00:02:40] ready [00:03:05] (or if you'd like the practice, it's the steps at https://wikitech.wikimedia.org/wiki/Mw-cron_jobs#Manually_running_a_CronJob -- in our case instead of mediawiki-main-serviceops-version it's testkitchen-constructiveedits [00:03:07] ) [00:05:15] going ahead [00:06:12] hmm, failed [00:06:18] `aawiki Fatal error: Cannot declare class WikimediaEvents\Maintenance\InstrumentConstructiveEdits, because the name is already in use in /srv/mediawiki/php-1.47.0-wmf.13/extensions/WikimediaEvents/maintenance/InstrumentConstructiveEdits.php on line 22` [00:06:29] no constructive edits exist :( [00:06:52] oh - probably we need some events first? [00:07:38] FIRING: [2x] CertAlmostExpired: gNMI TLS certificate for lsw1-d6-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [00:10:54] rzl: thanks for running - i think we need some constructive edits first [00:11:43] I'm a little surprised to get that PHP error in that situation, but you're the expert in that script :) [00:11:51] cjming: rzl: this seems kinda cryptic, but I wonder if this is what happens if the script file is missing the [00:11:51] $maintClass = InstrumentConstructiveEdits::class; [00:11:51] require_once RUN_MAINTENANCE_IF_MAIN; [00:12:03] yeah that's exactly what I was just looking at [00:12:04] ... which seems to be the case here [00:14:06] swfrench-wmf: where does that need to be added? to InstrumentConstructiveEdits.php? [00:14:18] yeah, take a look at UpdatePeriodicMetrics.php by comparison [00:14:29] (in the same directory) [00:14:59] (03CR) 10Jforrester: [C:03+2] "…" [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1320288 (owner: 10TrainBranchBot) [00:15:52] Yeah, the boilerplate is there for reasons :) [00:17:11] (03PS2) 10Cwhite: logstash: enable reversing security plugin enable [puppet] - 10https://gerrit.wikimedia.org/r/1320291 (https://phabricator.wikimedia.org/T350516) [00:17:11] very curious about how exactly we end up with the duplicate class declaration error without it (if that's indeed *the* issue, rather than just *an* issue), but that will need added for this to work as expected [00:18:02] (03PS3) 10Cwhite: logstash: enable reversing security plugin enable [puppet] - 10https://gerrit.wikimedia.org/r/1320291 (https://phabricator.wikimedia.org/T350516) [00:19:10] shoot - ok i will add that and maybe hopefully deploy/backport it rn [00:19:17] (03PS4) 10Cwhite: logstash: enable reversing security plugin enable [puppet] - 10https://gerrit.wikimedia.org/r/1320291 (https://phabricator.wikimedia.org/T350516) [00:19:23] thanks for the help and fix [00:20:01] (03Merged) 10jenkins-bot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1320288 (owner: 10TrainBranchBot) [00:21:22] (03PS5) 10Effie Mouzeli: api-gateway:: remove API Gateway specific ratelimiting logic #3 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319563 (https://phabricator.wikimedia.org/T428625) (owner: 10Aaron Schulz) [00:21:43] cjming: want us to leave it as-is in the meantime, or disable it? might open some nuisance alert tickets if we leave it, but won't harm anything unless that tricks someone into investigating something we already figured out; if we disable it it'll be another puppet patch to re-enable it, but that's easy [00:21:54] I'd lean toward switching it off, but up to you [00:22:13] is it ok if i merge the fix real quick and backport it? [00:22:53] oh like right now! no objection from me, I just might not be at keys for too much longer, so if you decide you want it disabled after all, feel free to leave me a note and I'll take care of that tomorrow [00:23:00] FIRING: HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-rollback - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs-next - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [00:23:23] rzl: thank you so much - i'm going to be bold and backport it now [00:24:21] Are we down? https://usercontent.irccloud-cdn.com/file/RQEFUHDw/image.png [00:24:27] https://en.wikipedia.org/wiki/Agreement_on_Climate_Change,_Trade_and_Sustainability [00:24:51] https://usercontent.irccloud-cdn.com/file/u9Z3uDbB/image.png [00:25:03] Jake_Park: nothing widespread, might just be your ISP! but if you can grab the infomration at https://wikitech-static.wikimedia.org/wiki/Reporting_a_connectivity_issue we can take a look [00:26:52] jouncebot: nowandnext [00:26:53] No deployments scheduled for the next 1 hour(s) and 33 minute(s) [00:26:53] In 1 hour(s) and 33 minute(s): Automatic branching of MediaWiki, extensions, skins, and vendor – see Heterogeneous deployment/Train deploys (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260804T0200) [00:27:24] rzl: Will do, thanks! [00:27:50] if it's ok, i'm going to self-merge https://gerrit.wikimedia.org/r/c/mediawiki/extensions/WikimediaEvents/+/1320295 and backport it in the next few minutes [00:28:01] Jake_Park: thank you! sorry for the trouble [00:31:06] It's back (i knew this would happen lol) [00:34:08] (03PS1) 10Clare Ming: Fix InstrumentConstructiveEdits script [extensions/WikimediaEvents] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1320296 (https://phabricator.wikimedia.org/T431493) [00:34:44] (03CR) 10TrainBranchBot: [C:03+2] "Approved by cjming@deploy1003 using scap backport" [extensions/WikimediaEvents] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1320296 (https://phabricator.wikimedia.org/T431493) (owner: 10Clare Ming) [00:38:35] (03Merged) 10jenkins-bot: Fix InstrumentConstructiveEdits script [extensions/WikimediaEvents] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1320296 (https://phabricator.wikimedia.org/T431493) (owner: 10Clare Ming) [00:39:30] (03PS7) 10KineticPelagic: rest: Add test server option to REST Sandbox for Wikipedia projects [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1313827 (https://phabricator.wikimedia.org/T408816) [00:39:33] !log cjming@deploy1003 Started scap sync-world: Backport for [[gerrit:1320296|Fix InstrumentConstructiveEdits script (T431493)]] [00:39:38] T431493: [DE 1.1.9] An instrument that produces edit_survived events - https://phabricator.wikimedia.org/T431493 [00:41:18] !log cjming@deploy1003 cjming: Backport for [[gerrit:1320296|Fix InstrumentConstructiveEdits script (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [00:41:46] !log cjming@deploy1003 cjming: Continuing with deployment [00:42:10] PROBLEM - Blazegraph Port for wdqs-blazegraph on wdqs2013 is CRITICAL: connect to address 127.0.0.1 and port 9999: Connection refused https://wikitech.wikimedia.org/wiki/Wikidata_query_service/Runbook [00:43:04] RECOVERY - Blazegraph Port for wdqs-blazegraph on wdqs2013 is OK: TCP OK - 0.000 second response time on 127.0.0.1 port 9999 https://wikitech.wikimedia.org/wiki/Wikidata_query_service/Runbook [00:45:53] !log cjming@deploy1003 Finished scap sync-world: Backport for [[gerrit:1320296|Fix InstrumentConstructiveEdits script (T431493)]] (duration: 06m 20s) [00:45:58] T431493: [DE 1.1.9] An instrument that produces edit_survived events - https://phabricator.wikimedia.org/T431493 [00:48:03] FIRING: [2x] ProbeDown: Service titan1002:443 has failed probes (http_thanos_wikimedia_org_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#titan1002:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [00:53:03] RESOLVED: [2x] ProbeDown: Service titan1002:443 has failed probes (http_thanos_wikimedia_org_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#titan1002:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [00:58:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 20.1% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [01:08:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.97% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [01:11:02] (03PS1) 10TrainBranchBot: Branch commit for wmf/1.47.0-wmf.14 [core] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1320310 (https://phabricator.wikimedia.org/T430833) [01:11:05] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/1.47.0-wmf.14 [core] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1320310 (https://phabricator.wikimedia.org/T430833) (owner: 10TrainBranchBot) [01:12:11] (03PS1) 10TrainBranchBot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1320311 [01:12:11] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1320311 (owner: 10TrainBranchBot) [01:14:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.9% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [01:19:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.97% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [01:21:08] (03Merged) 10jenkins-bot: Branch commit for wmf/1.47.0-wmf.14 [core] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1320310 (https://phabricator.wikimedia.org/T430833) (owner: 10TrainBranchBot) [01:26:44] (03Merged) 10jenkins-bot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1320311 (owner: 10TrainBranchBot) [02:00:05] Deploy window Automatic branching of MediaWiki, extensions, skins, and vendor – see Heterogeneous deployment/Train deploys (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260804T0200) [02:00:56] !log mwpresync@deploy1003 Started scap build-images: Publishing wmf/next image [02:03:01] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:07:28] !log mwpresync@deploy1003 Finished scap build-images: Publishing wmf/next image (duration: 06m 32s) [02:08:00] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:13:57] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:18:00] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:31:42] (03CR) 10Andrea Denisse: [C:03+1] "LGTM, thanks!!" [puppet] - 10https://gerrit.wikimedia.org/r/1320291 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [02:56:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [03:00:05] Deploy window Automatic deployment of MediaWiki, extensions, skins, and vendor to testwikis only – see Heterogeneous deployment/Train deploys (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260804T0300) [03:01:57] (03PS1) 10TrainBranchBot: testwikis to 1.47.0-wmf.14 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320351 (https://phabricator.wikimedia.org/T430833) [03:02:00] (03CR) 10TrainBranchBot: [C:03+2] "Initiated by mwpresync@deploy1003" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320351 (https://phabricator.wikimedia.org/T430833) (owner: 10TrainBranchBot) [03:05:20] (03Merged) 10jenkins-bot: testwikis to 1.47.0-wmf.14 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320351 (https://phabricator.wikimedia.org/T430833) (owner: 10TrainBranchBot) [03:05:37] !log mwpresync@deploy1003 Started scap sync-world: testwikis to 1.47.0-wmf.14 refs T430833 [03:05:41] T430833: 1.47.0-wmf.14 deployment blockers - https://phabricator.wikimedia.org/T430833 [03:17:09] (03PS1) 10Clare Ming: Test Kitchen UI: Deploy v1.5.1 release to staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320361 (https://phabricator.wikimedia.org/T433265) [03:19:50] (03CR) 10Clare Ming: [C:03+2] Test Kitchen UI: Deploy v1.5.1 release to staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320361 (https://phabricator.wikimedia.org/T433265) (owner: 10Clare Ming) [03:20:02] PROBLEM - Improperly owned -0:0- files in /srv/mediawiki-staging on deploy2003 is CRITICAL: Improperly owned (0:0) files in /srv/mediawiki-staging https://wikitech.wikimedia.org/wiki/Monitoring/bad_directory_owner [03:22:00] (03Merged) 10jenkins-bot: Test Kitchen UI: Deploy v1.5.1 release to staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320361 (https://phabricator.wikimedia.org/T433265) (owner: 10Clare Ming) [03:22:40] !log cjming@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply [03:23:04] !log cjming@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply [03:30:02] RECOVERY - Improperly owned -0:0- files in /srv/mediawiki-staging on deploy2003 is OK: Files ownership is ok. https://wikitech.wikimedia.org/wiki/Monitoring/bad_directory_owner [03:38:34] !log mwpresync@deploy1003 Finished scap sync-world: testwikis to 1.47.0-wmf.14 refs T430833 (duration: 32m 57s) [03:38:38] T430833: 1.47.0-wmf.14 deployment blockers - https://phabricator.wikimedia.org/T430833 [04:00:05] Deploy window Automatic removal of all obsolete MediaWiki versions from the deployment and bare metal servers (except the most-recent obsolete version) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260804T0400) [04:02:35] !log mwpresync@deploy1003 Pruned MediaWiki: 1.47.0-wmf.11 (duration: 02m 29s) [04:07:38] FIRING: [2x] CertAlmostExpired: gNMI TLS certificate for lsw1-d6-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [04:23:01] FIRING: HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-rollback - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs-next - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [05:29:12] (03PS1) 10PipelineBot: mobileapps: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320429 [05:45:54] 06SRE, 10LDAP-Access-Requests: Grant Access to NDA for jaleman-vdr-wmf - https://phabricator.wikimedia.org/T433417#12181675 (10JArguello-WMF) hi @KFrancis thanks for your help, the full name is Jose Alemán, he is a vendor working with Crescendo. Please let me know what other information you need to proceed. Th... [05:55:43] 06SRE, 10Ganeti, 06Infrastructure-Foundations: SSH host key verification failures in Ganeti intra node SSH calls after Bullseye update - https://phabricator.wikimedia.org/T309724#12181678 (10ayounsi) @bking can you try https://wikitech.wikimedia.org/wiki/Ganeti#If_network_or_SSH_between_ganeti_nodes_is_unava... [06:00:05] Deploy window MediaWiki infrastructure (UTC early) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260804T0600) [06:00:05] marostegui, Amir1, and federico3: OwO what's this, a deployment window?? Primary database switchover. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260804T0600). nyaa~ [06:18:01] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [06:32:02] (03PS1) 10Jcrespo: dbbackups: Replace db1150 usages for newer db1265 [puppet] - 10https://gerrit.wikimedia.org/r/1320627 (https://phabricator.wikimedia.org/T433825) [06:32:14] (03CR) 10Jcrespo: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1320627 (https://phabricator.wikimedia.org/T433825) (owner: 10Jcrespo) [06:36:56] (03PS1) 10Jcrespo: dbbackups: Replace db1171 usages for newer db1285 [puppet] - 10https://gerrit.wikimedia.org/r/1320628 (https://phabricator.wikimedia.org/T433825) [06:38:02] (03CR) 10Jcrespo: [C:03+1] "https://puppet-compiler.wmflabs.org/output/1320627/7444/" [puppet] - 10https://gerrit.wikimedia.org/r/1320627 (https://phabricator.wikimedia.org/T433825) (owner: 10Jcrespo) [06:38:09] (03CR) 10Jcrespo: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1320628 (https://phabricator.wikimedia.org/T433825) (owner: 10Jcrespo) [06:40:58] (03PS2) 10Jcrespo: dbbackups: Replace db1150 usages for newer db1265 [puppet] - 10https://gerrit.wikimedia.org/r/1320627 (https://phabricator.wikimedia.org/T433825) [06:41:02] (03CR) 10Jcrespo: [C:03+1] "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1320627 (https://phabricator.wikimedia.org/T433825) (owner: 10Jcrespo) [06:42:35] (03CR) 10Jcrespo: [C:03+1] "https://puppet-compiler.wmflabs.org/output/1320628/7445/" [puppet] - 10https://gerrit.wikimedia.org/r/1320628 (https://phabricator.wikimedia.org/T433825) (owner: 10Jcrespo) [06:48:08] (03CR) 10Slyngshede: [C:03+2] IDM: Update IDM to version 0.1.18 [dns] - 10https://gerrit.wikimedia.org/r/1320124 (owner: 10Slyngshede) [06:48:17] !log slyngshede@dns1004 START - running authdns-update [06:50:05] !log slyngshede@dns1004 END - running authdns-update [06:56:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [06:58:45] (03CR) 10Jcrespo: [C:03+2] dbbackups: Replace db1150 usages for newer db1265 [puppet] - 10https://gerrit.wikimedia.org/r/1320627 (https://phabricator.wikimedia.org/T433825) (owner: 10Jcrespo) [07:00:05] Amir1, urbanecm, and awight: #bothumor My software never has bugs. It just develops random features. Rise for UTC morning backport window. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260804T0700). [07:00:05] thedj: A patch you scheduled for UTC morning backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [07:01:14] (03PS2) 10Jcrespo: dbbackups: Replace db1171 usages for newer db1285 [puppet] - 10https://gerrit.wikimedia.org/r/1320628 (https://phabricator.wikimedia.org/T433825) [07:01:16] (03CR) 10Jcrespo: [C:03+1] "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1320628 (https://phabricator.wikimedia.org/T433825) (owner: 10Jcrespo) [07:11:43] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-be1093.eqiad.wmnet with OS trixie [07:11:51] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12181743 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-be1093.eqiad.wmnet with OS trixie [07:12:01] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host ms-be2092.codfw.wmnet with OS trixie [07:12:10] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12181745 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host ms-be2092.codfw.wmnet with OS trixie [07:15:14] Amir1 or urbanecm about ? [07:26:13] (03PS1) 10MVernon: swift: move ms-be10{69,70,71} to new-style storage, remove from rings [puppet] - 10https://gerrit.wikimedia.org/r/1320658 (https://phabricator.wikimedia.org/T429630) [07:29:37] !log running extra backups to test db1265 T433825 [07:29:41] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [07:29:42] T433825: decommission db1150.eqiad.wmnet - https://phabricator.wikimedia.org/T433825 [07:30:50] (03CR) 10Jcrespo: [C:03+2] dbbackups: Replace db1171 usages for newer db1285 [puppet] - 10https://gerrit.wikimedia.org/r/1320628 (https://phabricator.wikimedia.org/T433825) (owner: 10Jcrespo) [07:32:57] !log mvernon@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage [07:33:43] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage [07:38:58] (03PS2) 10MVernon: swift: move ms-be10{69,70,71} to new-style storage, remove from rings [puppet] - 10https://gerrit.wikimedia.org/r/1320658 (https://phabricator.wikimedia.org/T429630) [07:39:34] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage [07:40:20] ok. shifted to the afternoon [07:43:34] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage [07:45:41] (03CR) 10Jcrespo: [C:03+1] swift: move ms-be10{69,70,71} to new-style storage, remove from rings [puppet] - 10https://gerrit.wikimedia.org/r/1320658 (https://phabricator.wikimedia.org/T429630) (owner: 10MVernon) [07:48:05] (03CR) 10MVernon: [C:03+2] swift: move ms-be10{69,70,71} to new-style storage, remove from rings [puppet] - 10https://gerrit.wikimedia.org/r/1320658 (https://phabricator.wikimedia.org/T429630) (owner: 10MVernon) [07:51:17] !log cwilliams@cumin1003 START - Cookbook sre.hosts.reboot-single for host db1266.eqiad.wmnet [07:53:23] (03CR) 10Slyngshede: [C:03+2] varnish: Split out media functions from upload [puppet] - 10https://gerrit.wikimedia.org/r/1311110 (https://phabricator.wikimedia.org/T427465) (owner: 10BCornwall) [07:55:59] !log running extra backups to test db1285 T433826 [07:56:03] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [07:56:03] T433826: decommission db1171.eqiad.wmnet - https://phabricator.wikimedia.org/T433826 [07:56:53] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1266.eqiad.wmnet [07:58:05] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1093.eqiad.wmnet with OS trixie [07:58:21] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12181815 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-be1093.eqiad.wmnet with OS trixie completed... [07:59:02] !log cwilliams@cumin1003 START - Cookbook sre.hosts.reboot-single for host db1267.eqiad.wmnet [08:00:05] jnuche and jeena: #bothumor Q:How do functions break up? A:They stop calling each other. Rise for MediaWiki train - Utc-0+Utc-7 Version deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260804T0800). [08:00:11] morning, I'll be rolling out the train in the next few minutes [08:00:17] (03PS1) 10Jcrespo: mariadb: Set db1150 as insetup for decommissioning [puppet] - 10https://gerrit.wikimedia.org/r/1320682 (https://phabricator.wikimedia.org/T433825) [08:03:20] (03PS1) 10Effie Mouzeli: site.pp: retire parsoidtest and testreduce [puppet] - 10https://gerrit.wikimedia.org/r/1320683 (https://phabricator.wikimedia.org/T421484) [08:03:42] (03PS1) 10Jcrespo: mariadb: Set db1171 as insetup for decommissioning [puppet] - 10https://gerrit.wikimedia.org/r/1320684 (https://phabricator.wikimedia.org/T433826) [08:04:01] (03CR) 10Jcrespo: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1320682 (https://phabricator.wikimedia.org/T433825) (owner: 10Jcrespo) [08:04:07] (03CR) 10Jcrespo: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1320684 (https://phabricator.wikimedia.org/T433826) (owner: 10Jcrespo) [08:04:16] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1267.eqiad.wmnet [08:04:22] (03CR) 10Effie Mouzeli: profile::memcached::instance: migrate to PKI (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1320123 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [08:04:30] !log cwilliams@cumin1003 START - Cookbook sre.hosts.reboot-single for host db1268.eqiad.wmnet [08:04:57] (03PS1) 10TrainBranchBot: group0 to 1.47.0-wmf.14 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320685 (https://phabricator.wikimedia.org/T430833) [08:05:00] (03CR) 10TrainBranchBot: [C:03+2] "Initiated by jnuche@deploy1003" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320685 (https://phabricator.wikimedia.org/T430833) (owner: 10TrainBranchBot) [08:06:07] (03PS13) 10Effie Mouzeli: profile::memcached::instance: migrate to PKI [puppet] - 10https://gerrit.wikimedia.org/r/1320123 (https://phabricator.wikimedia.org/T353511) [08:06:22] (03Merged) 10jenkins-bot: group0 to 1.47.0-wmf.14 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320685 (https://phabricator.wikimedia.org/T430833) (owner: 10TrainBranchBot) [08:06:26] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2092.codfw.wmnet with OS trixie [08:06:38] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12181839 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-be2092.codfw.wmnet with OS trixie completed... [08:07:38] FIRING: [2x] CertAlmostExpired: gNMI TLS certificate for lsw1-d6-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [08:09:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 20.14% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [08:09:44] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1268.eqiad.wmnet [08:09:57] !log cwilliams@cumin1003 START - Cookbook sre.hosts.reboot-single for host db1269.eqiad.wmnet [08:14:43] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host ms-be2093.codfw.wmnet with OS trixie [08:14:58] FIRING: ProbeDown: Service text-https:443 has failed probes (http_text-https_ip6) - https://wikitech.wikimedia.org/wiki/Runbook#text-https:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [08:15:00] (03CR) 10Marostegui: [C:03+1] mariadb: Set db1171 as insetup for decommissioning [puppet] - 10https://gerrit.wikimedia.org/r/1320684 (https://phabricator.wikimedia.org/T433826) (owner: 10Jcrespo) [08:15:01] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12181855 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host ms-be2093.codfw.wmnet with OS trixie [08:15:11] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1269.eqiad.wmnet [08:15:24] (03CR) 10Marostegui: [C:03+1] "it may appear on som other backup.cnf files?" [puppet] - 10https://gerrit.wikimedia.org/r/1320684 (https://phabricator.wikimedia.org/T433826) (owner: 10Jcrespo) [08:15:24] !log cwilliams@cumin1003 START - Cookbook sre.hosts.reboot-single for host db1270.eqiad.wmnet [08:15:33] (03PS1) 10Jgiannelos: mobileapps: Pin staging to latest master [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320742 [08:18:15] FIRING: MediaWikiLatencyExceeded: p75 latency high: eqiad mw-web releases routed via main (k8s) 930.6ms - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook#Average_latency_exceeded - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=55&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiLatencyExceeded [08:18:45] FIRING: WidespreadPuppetFailure: Puppet has failed in magru - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [08:19:05] Oh, the Puppet in magru is me [08:19:58] RESOLVED: ProbeDown: Service text-https:443 has failed probes (http_text-https_ip6) - https://wikitech.wikimedia.org/wiki/Runbook#text-https:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [08:20:39] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1270.eqiad.wmnet [08:20:53] !log cwilliams@cumin1003 START - Cookbook sre.hosts.reboot-single for host db1271.eqiad.wmnet [08:21:45] !log jnuche@deploy1003 rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.14 refs T430833 [08:21:50] T430833: 1.47.0-wmf.14 deployment blockers - https://phabricator.wikimedia.org/T430833 [08:23:01] FIRING: HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-rollback - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs-next - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [08:23:15] RESOLVED: MediaWikiLatencyExceeded: p75 latency high: eqiad mw-web releases routed via main (k8s) 930.6ms - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook#Average_latency_exceeded - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=55&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiLatencyExceeded [08:23:24] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-be1094.eqiad.wmnet with OS trixie [08:23:41] 06SRE, 10SRE-swift-storage, 07Essential-Work, 13Patch-For-Review: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12181873 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-be1094.eqiad.wmnet with OS trixie [08:24:57] (03PS1) 10Slyngshede: Revert "varnish: Split out media functions from upload" [puppet] - 10https://gerrit.wikimedia.org/r/1320749 [08:25:05] (03PS1) 10Cathal Mooney: L3 Switch Policy: core-out for unicast EBGP not allowing Calico /122s [homer/public] - 10https://gerrit.wikimedia.org/r/1320750 (https://phabricator.wikimedia.org/T433881) [08:25:11] (03CR) 10Jgiannelos: [C:03+2] mobileapps: Pin staging to latest master [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320742 (owner: 10Jgiannelos) [08:26:07] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1271.eqiad.wmnet [08:26:18] (03CR) 10Slyngshede: [C:03+2] Revert "varnish: Split out media functions from upload" [puppet] - 10https://gerrit.wikimedia.org/r/1320749 (owner: 10Slyngshede) [08:27:26] (03Merged) 10jenkins-bot: mobileapps: Pin staging to latest master [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320742 (owner: 10Jgiannelos) [08:28:50] !log dcausse@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply [08:29:04] !log dcausse@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply [08:29:18] (03CR) 10Cathal Mooney: [C:03+2] L3 Switch Policy: core-out for unicast EBGP not allowing Calico /122s [homer/public] - 10https://gerrit.wikimedia.org/r/1320750 (https://phabricator.wikimedia.org/T433881) (owner: 10Cathal Mooney) [08:30:57] (03Merged) 10jenkins-bot: L3 Switch Policy: core-out for unicast EBGP not allowing Calico /122s [homer/public] - 10https://gerrit.wikimedia.org/r/1320750 (https://phabricator.wikimedia.org/T433881) (owner: 10Cathal Mooney) [08:32:52] (03PS1) 10Marostegui: clouddb1033: Remove note [puppet] - 10https://gerrit.wikimedia.org/r/1320753 [08:33:35] (03CR) 10Marostegui: [C:03+2] clouddb1033: Remove note [puppet] - 10https://gerrit.wikimedia.org/r/1320753 (owner: 10Marostegui) [08:34:40] !log jgiannelos@deploy1003 helmfile [staging] START helmfile.d/services/mobileapps: apply [08:35:00] !log jgiannelos@deploy1003 helmfile [staging] DONE helmfile.d/services/mobileapps: apply [08:35:38] 06SRE, 06Infrastructure-Foundations, 10netops: Consider removing BFD on datacentre IBGP peerings - https://phabricator.wikimedia.org/T433675#12181927 (10ayounsi) Sounds good to me! [08:36:13] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage [08:38:32] !log dcausse@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply [08:38:41] !log dcausse@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply [08:39:13] dcausse@deploy1003: Failed to log message to wiki. Somebody should check the error logs. [08:42:37] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage [08:43:42] (03PS1) 10JavierMonton: stream: pageview.trending.relative.v1 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320757 (https://phabricator.wikimedia.org/T432204) [08:44:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.17% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [08:44:15] FIRING: MediaWikiLatencyExceeded: p75 latency high: eqiad mw-web releases routed via main (k8s) 882ms - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook#Average_latency_exceeded - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=55&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiLatencyExceeded [08:44:20] !log mvernon@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage [08:45:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.59% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [08:45:18] !log cwilliams@cumin1003 START - Cookbook sre.hosts.reboot-single for host db1272.eqiad.wmnet [08:48:20] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage [08:49:15] RESOLVED: MediaWikiLatencyExceeded: p75 latency high: eqiad mw-web releases routed via main (k8s) 882ms - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook#Average_latency_exceeded - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=55&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiLatencyExceeded [08:49:48] !log dcausse@deploy1003 helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply [08:49:58] !log dcausse@deploy1003 helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply [08:50:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.31% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [08:50:32] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1272.eqiad.wmnet [08:50:39] FIRING: KubernetesAPILatency: High Kubernetes API latency (LIST events) on k8s-dse@eqiad - https://wikitech.wikimedia.org/wiki/Kubernetes - https://grafana.wikimedia.org/d/ddNd-sLnk/kubernetes-api-details?var-site=eqiad&var-cluster=k8s-dse&var-latency_percentile=0.95&var-verb=LIST - https://alerts.wikimedia.org/?q=alertname%3DKubernetesAPILatency [08:50:46] !log cwilliams@cumin1003 START - Cookbook sre.hosts.reboot-single for host db1273.eqiad.wmnet [08:52:26] (03PS1) 10Slyngshede: Revert^2 "varnish: Split out media functions from upload" [puppet] - 10https://gerrit.wikimedia.org/r/1320759 [08:52:32] 10ops-codfw, 06SRE, 06DC-Ops, 06ServiceOps: Rerack mc-gp2006 and mc2046 - https://phabricator.wikimedia.org/T432716#12182010 (10cmooney) >>! In T432716#12180337, @Jhancock.wm wrote: > it's magical to me. Updated the server location. I was gonna give the reimage a try but i didn't know which distro was on i... [08:53:29] (03CR) 10Joal: [C:03+1] stream: pageview.trending.relative.v1 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320757 (https://phabricator.wikimedia.org/T432204) (owner: 10JavierMonton) [08:55:42] (03PS1) 10Effie Mouzeli: changeprop-jobqueue: removing node affinity [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320761 (https://phabricator.wikimedia.org/T433881) [08:56:00] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1273.eqiad.wmnet [08:56:13] !log cwilliams@cumin1003 START - Cookbook sre.hosts.reboot-single for host db1274.eqiad.wmnet [08:56:15] RESOLVED: WidespreadPuppetFailure: Puppet has failed in magru - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [08:56:36] (03PS1) 10AikoChou: changeprop: Reduce revertrisk-wikidata concurrency 5 -> 2 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320762 (https://phabricator.wikimedia.org/T420883) [08:57:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.69% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [08:57:47] (03CR) 10Effie Mouzeli: [C:03+2] profile::memcached::instance: migrate to PKI [puppet] - 10https://gerrit.wikimedia.org/r/1320123 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [08:59:31] (03CR) 10Blake: [C:03+1] changeprop-jobqueue: removing node affinity [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320761 (https://phabricator.wikimedia.org/T433881) (owner: 10Effie Mouzeli) [09:00:01] (03CR) 10Effie Mouzeli: [C:03+2] changeprop-jobqueue: removing node affinity [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320761 (https://phabricator.wikimedia.org/T433881) (owner: 10Effie Mouzeli) [09:00:39] (03CR) 10DCausse: [C:03+2] opensearch-semantic-search-test: bump to opensearch 3.7.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320153 (https://phabricator.wikimedia.org/T433697) (owner: 10DCausse) [09:01:20] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1274.eqiad.wmnet [09:01:34] !log cwilliams@cumin1003 START - Cookbook sre.hosts.reboot-single for host db1275.eqiad.wmnet [09:02:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.69% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [09:02:17] (03Merged) 10jenkins-bot: changeprop-jobqueue: removing node affinity [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320761 (https://phabricator.wikimedia.org/T433881) (owner: 10Effie Mouzeli) [09:03:11] (03Merged) 10jenkins-bot: opensearch-semantic-search-test: bump to opensearch 3.7.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320153 (https://phabricator.wikimedia.org/T433697) (owner: 10DCausse) [09:03:18] RECOVERY - Host mc2046 is UP: PING OK - Packet loss = 0%, RTA = 31.70 ms [09:04:18] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2093.codfw.wmnet with OS trixie [09:04:30] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12182048 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-be2093.codfw.wmnet with OS trixie completed: - ms-be2093 (**PASS*... [09:06:36] (03CR) 10Jcrespo: [C:04-1] "Waiting for backup testing." [puppet] - 10https://gerrit.wikimedia.org/r/1320682 (https://phabricator.wikimedia.org/T433825) (owner: 10Jcrespo) [09:06:36] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1275.eqiad.wmnet [09:06:49] !log cwilliams@cumin1003 START - Cookbook sre.hosts.reboot-single for host db1276.eqiad.wmnet [09:06:58] (03CR) 10Jcrespo: [C:04-1] "Waiting for backup testing" [puppet] - 10https://gerrit.wikimedia.org/r/1320684 (https://phabricator.wikimedia.org/T433826) (owner: 10Jcrespo) [09:08:25] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1094.eqiad.wmnet with OS trixie [09:08:36] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12182074 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-be1094.eqiad.wmnet with OS trixie completed: - ms-be1094 (**PASS*... [09:10:58] (03CR) 10Jcrespo: [C:04-1] "It was removed @" [puppet] - 10https://gerrit.wikimedia.org/r/1320684 (https://phabricator.wikimedia.org/T433826) (owner: 10Jcrespo) [09:11:43] (03PS3) 10Blake: httpbb: Add a --request-header argument. [software/httpbb] - 10https://gerrit.wikimedia.org/r/1318682 (https://phabricator.wikimedia.org/T428972) [09:11:43] (03CR) 10Blake: httpbb: Add a --request-header argument. (032 comments) [software/httpbb] - 10https://gerrit.wikimedia.org/r/1318682 (https://phabricator.wikimedia.org/T428972) (owner: 10Blake) [09:11:54] jouncebot: now [09:11:54] For the next 0 hour(s) and 48 minute(s): MediaWiki train - Utc-0+Utc-7 Version (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260804T0800) [09:12:03] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1276.eqiad.wmnet [09:12:15] !log dcausse@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply [09:12:16] !log cwilliams@cumin1003 START - Cookbook sre.hosts.reboot-single for host db1277.eqiad.wmnet [09:12:29] !log dcausse@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply [09:12:58] !log jiji@deploy1003 helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply [09:13:31] !log jiji@deploy1003 helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply [09:13:32] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host ms-be2094.codfw.wmnet with OS trixie [09:13:44] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12182090 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host ms-be2094.codfw.wmnet with OS trixie [09:13:54] (03PS3) 10Mpostoronca: WikimediaAntiAbuse: Document required load order after Echo [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319086 (https://phabricator.wikimedia.org/T432452) [09:15:49] (03PS3) 10DCausse: opensearch-semantic-search: bump to opensearch 3.7.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319845 (https://phabricator.wikimedia.org/T433697) [09:15:49] (03PS1) 10DCausse: opensearch-semantic-search-test: config values must be strings [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320765 [09:17:06] (03CR) 10DCausse: [C:03+2] opensearch-semantic-search-test: config values must be strings [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320765 (owner: 10DCausse) [09:17:30] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1277.eqiad.wmnet [09:19:22] (03Merged) 10jenkins-bot: opensearch-semantic-search-test: config values must be strings [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320765 (owner: 10DCausse) [09:20:02] !log brouberol@cumin1003 START - Cookbook sre.ganeti.reboot-vm for VM archiva1002.wikimedia.org [09:20:39] RESOLVED: KubernetesAPILatency: High Kubernetes API latency (LIST events) on k8s-dse@eqiad - https://wikitech.wikimedia.org/wiki/Kubernetes - https://grafana.wikimedia.org/d/ddNd-sLnk/kubernetes-api-details?var-site=eqiad&var-cluster=k8s-dse&var-latency_percentile=0.95&var-verb=LIST - https://alerts.wikimedia.org/?q=alertname%3DKubernetesAPILatency [09:23:02] (03PS6) 10Effie Mouzeli: hieradata: switch mc1059 to PKI [puppet] - 10https://gerrit.wikimedia.org/r/1320132 (https://phabricator.wikimedia.org/T353511) [09:23:58] !log brouberol@cumin1003 END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM archiva1002.wikimedia.org [09:24:03] (03PS3) 10Slyngshede: varnish: Split out media functions from upload [puppet] - 10https://gerrit.wikimedia.org/r/1320759 (https://phabricator.wikimedia.org/T427465) [09:24:08] (03CR) 10Slyngshede: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1320759 (https://phabricator.wikimedia.org/T427465) (owner: 10Slyngshede) [09:24:27] !log dcausse@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply [09:24:36] !log dcausse@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply [09:27:10] !log mvernon@cumin1003 START - Cookbook sre.swift.convert-disks for host ms-be1069 [09:28:32] (03PS1) 10Jcrespo: backup: Set backup1003 & backup2003 as insetup for decom. [puppet] - 10https://gerrit.wikimedia.org/r/1320768 (https://phabricator.wikimedia.org/T420506) [09:30:26] (03PS2) 10Jcrespo: backup: Set backup1003 & backup2003 as insetup for decom. [puppet] - 10https://gerrit.wikimedia.org/r/1320768 (https://phabricator.wikimedia.org/T420506) [09:30:45] (03CR) 10Jcrespo: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1320768 (https://phabricator.wikimedia.org/T420506) (owner: 10Jcrespo) [09:33:09] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie [09:33:16] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12182252 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-be1095.eqiad.wmnet with OS trixie [09:33:55] !log mvernon@cumin1003 START - Cookbook sre.swift.convert-disks for host ms-be1070 [09:34:14] (03CR) 10Effie Mouzeli: [C:03+2] hieradata: switch mc1059 to PKI [puppet] - 10https://gerrit.wikimedia.org/r/1320132 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [09:34:20] !log mvernon@cumin1003 START - Cookbook sre.swift.convert-disks for host ms-be1071 [09:34:33] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage [09:35:17] (03CR) 10Gkyziridis: [C:03+1] "Thnx! LGTM" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320762 (https://phabricator.wikimedia.org/T420883) (owner: 10AikoChou) [09:35:40] 10ops-eqiad, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 10netops: eqiad row A/B switch refresh prep - https://phabricator.wikimedia.org/T418012#12182272 (10cmooney) [09:40:06] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage [09:44:08] !log dcausse@deploy1003 helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply [09:44:23] !log dcausse@deploy1003 helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply [09:53:27] (03CR) 10AikoChou: [C:03+2] changeprop: Reduce revertrisk-wikidata concurrency 5 -> 2 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320762 (https://phabricator.wikimedia.org/T420883) (owner: 10AikoChou) [09:53:46] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply [09:53:47] !log mvernon@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage [09:54:14] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply [09:54:28] (03CR) 10Effie Mouzeli: [C:03+2] site.pp: retire parsoidtest and testreduce [puppet] - 10https://gerrit.wikimedia.org/r/1320683 (https://phabricator.wikimedia.org/T421484) (owner: 10Effie Mouzeli) [09:55:26] (03PS1) 10Mszwarc: UIC: Add user name to server-side instrumentation events [extensions/CheckUser] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1320781 (https://phabricator.wikimedia.org/T433816) [09:55:40] (03PS1) 10Mszwarc: UIC: Add user name to server-side instrumentation events [extensions/CheckUser] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1320782 (https://phabricator.wikimedia.org/T433816) [09:55:46] (03Merged) 10jenkins-bot: changeprop: Reduce revertrisk-wikidata concurrency 5 -> 2 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320762 (https://phabricator.wikimedia.org/T420883) (owner: 10AikoChou) [09:58:20] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage [10:00:04] Deploy window MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260804T1000) [10:00:52] (03CR) 10DCausse: "tested an upgrade of the test cluster (moving from 3.3 to 3.7) and it went well" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319845 (https://phabricator.wikimedia.org/T433697) (owner: 10DCausse) [10:01:22] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2094.codfw.wmnet with OS trixie [10:01:29] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12182352 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-be2094.codfw.wmnet with OS trixie completed: - ms-be2094 (**PASS*... [10:08:44] !log jayme@cumin1003 START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1139.eqiad.wmnet [10:08:48] !log jayme@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1139.eqiad.wmnet [10:09:23] !log jayme@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1139.eqiad.wmnet [10:10:05] !log jayme@cumin1003 START - Cookbook sre.hosts.reimage for host wikikube-worker1139.eqiad.wmnet with OS trixie [10:10:23] !log jayme@cumin1003 START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1140.eqiad.wmnet [10:10:26] !log jayme@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1140.eqiad.wmnet [10:10:33] !log jayme@cumin1003 START - Cookbook sre.hosts.move-vlan for host wikikube-worker1139 [10:10:46] !log jayme@cumin1003 START - Cookbook sre.dns.netbox [10:11:01] !log jayme@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1140.eqiad.wmnet [10:11:12] !log jayme@cumin1003 START - Cookbook sre.hosts.reimage for host wikikube-worker1140.eqiad.wmnet with OS trixie [10:11:40] !log jayme@cumin1003 START - Cookbook sre.hosts.move-vlan for host wikikube-worker1140 [10:12:37] !log mvernon@cumin1003 END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1069 [10:13:04] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-be1069.eqiad.wmnet with OS trixie [10:13:19] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12182408 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-be1069.eqiad.wmnet with OS trixie [10:14:09] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie [10:14:19] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12182419 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host ms-be2095.codfw.wmnet with OS trixie [10:15:02] !log jayme@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" [10:15:46] !log jayme@cumin1003 START - Cookbook sre.dns.netbox [10:15:47] !log jayme@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003" [10:15:47] !log jayme@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [10:15:47] !log jayme@cumin1003 START - Cookbook sre.dns.wipe-cache wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [10:15:50] !log jayme@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [10:15:51] !log jayme@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1139 [10:17:08] !log mvernon@cumin1003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be1095.eqiad.wmnet with OS trixie [10:17:12] !log jayme@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1139 [10:17:12] !log jayme@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1139 [10:17:14] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12182424 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-be1095.eqiad.wmnet with OS trixie completed: - ms-be1095 (**FAIL*... [10:17:16] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12182425 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-be1095.eqiad.wmnet with OS trixie executed with errors: - ms-be10... [10:17:39] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie [10:17:46] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12182426 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-be1095.eqiad.wmnet with OS trixie [10:18:01] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [10:20:26] !log mvernon@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage [10:21:09] !log mvernon@cumin1003 END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1070 [10:21:34] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-be1070.eqiad.wmnet with OS trixie [10:21:35] !log mvernon@cumin1003 END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1071 [10:21:42] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12182434 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-be1070.eqiad.wmnet with OS trixie [10:21:45] !log jayme@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" [10:21:50] !log jayme@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003" [10:21:50] !log jayme@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [10:21:50] !log jayme@cumin1003 START - Cookbook sre.dns.wipe-cache wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [10:21:51] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-be1071.eqiad.wmnet with OS trixie [10:21:53] !log jayme@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [10:21:54] !log jayme@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1140 [10:21:59] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12182435 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-be1071.eqiad.wmnet with OS trixie [10:23:11] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage [10:23:54] !log jayme@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1140 [10:23:54] !log jayme@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1140 [10:27:43] (03PS1) 10Ayounsi: Add OSPF for codfw HE transport [homer/public] - 10https://gerrit.wikimedia.org/r/1320789 (https://phabricator.wikimedia.org/T429653) [10:30:12] !log mvernon@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage [10:31:31] FIRING: [2x] ProbeDown: Service gerrit2003:443 has failed probes (http_gerrit_tls_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#gerrit2003:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [10:32:05] (03PS1) 10Ayounsi: Add OSPF for codfw HE transport [homer/public] - 10https://gerrit.wikimedia.org/r/1320807 (https://phabricator.wikimedia.org/T429653) [10:33:07] !log jayme@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage [10:33:10] 06SRE, 06ServiceOps, 10ServiceOps-Upgrades-Hardware, 07Kubernetes, 13Patch-For-Review: wikikube-ctrl2006 implementation tracking - https://phabricator.wikimedia.org/T406596#12182487 (10JMeybohm) Pinging here since this appears stale for almost a month. Is there anything we can do to get this sorted? [10:33:19] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage [10:33:23] (03Abandoned) 10JavierMonton: topic: webrequest-pageview [deployment-charts] - 10https://gerrit.wikimedia.org/r/1310018 (https://phabricator.wikimedia.org/T430675) (owner: 10JavierMonton) [10:33:50] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1095.eqiad.wmnet with OS trixie [10:33:57] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12182504 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-be1095.eqiad.wmnet with OS trixie completed: - ms-be1095 (**PASS*... [10:35:05] (03CR) 10Cathal Mooney: [C:03+1] "lgtm as long as you're good with the alert in eqsin :)" [homer/public] - 10https://gerrit.wikimedia.org/r/1320789 (https://phabricator.wikimedia.org/T429653) (owner: 10Ayounsi) [10:35:23] (03CR) 10Ayounsi: [C:03+2] Add OSPF for codfw HE transport [homer/public] - 10https://gerrit.wikimedia.org/r/1320789 (https://phabricator.wikimedia.org/T429653) (owner: 10Ayounsi) [10:35:26] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage [10:36:31] RESOLVED: [2x] ProbeDown: Service gerrit2003:443 has failed probes (http_gerrit_tls_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#gerrit2003:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [10:37:02] 10ops-codfw, 06DC-Ops, 10Prod-Kubernetes, 06ServiceOps, 10ServiceOps-Upgrades-Hardware: move wikikube-worker servers in same rack - https://phabricator.wikimedia.org/T431585#12182518 (10JMeybohm) [10:37:16] (03Merged) 10jenkins-bot: Add OSPF for codfw HE transport [homer/public] - 10https://gerrit.wikimedia.org/r/1320789 (https://phabricator.wikimedia.org/T429653) (owner: 10Ayounsi) [10:37:55] !log jayme@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage [10:38:04] !log jayme@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage [10:38:49] !log mvernon@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage [10:39:02] !log mvernon@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage [10:41:33] !log jayme@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage [10:41:55] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie [10:42:03] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12182576 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-be1096.eqiad.wmnet with OS trixie [10:42:33] 06SRE, 06App Experience, 06Content-Platform-Team, 06ServiceOps, and 2 others: Investigate Code 414 error when selecting zh-classical (lzh) language from article toolbar - https://phabricator.wikimedia.org/T425545#12182590 (10Phreelance) @Raine above you mentioned a hotfix for `zh-classical` – could the sam... [10:42:41] PROBLEM - OSPF status on cr1-codfw is CRITICAL: OSPFv2: 5/9 UP : OSPFv3: 5/9 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [10:43:02] (03PS1) 10Effie Mouzeli: memcached: implement refresh_certs [puppet] - 10https://gerrit.wikimedia.org/r/1320809 (https://phabricator.wikimedia.org/T353511) [10:43:37] (03CR) 10CI reject: [V:04-1] memcached: implement refresh_certs [puppet] - 10https://gerrit.wikimedia.org/r/1320809 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [10:43:57] FIRING: [2x] HelmReleaseBadStatus: Helm release wdqs/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [10:44:49] !log mvernon@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage [10:45:13] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage [10:46:27] (03PS2) 10Effie Mouzeli: memcached: implement refresh_certs [puppet] - 10https://gerrit.wikimedia.org/r/1320809 (https://phabricator.wikimedia.org/T353511) [10:47:03] (03CR) 10CI reject: [V:04-1] memcached: implement refresh_certs [puppet] - 10https://gerrit.wikimedia.org/r/1320809 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [10:48:24] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage [10:49:47] PROBLEM - OSPF status on cr2-eqsin is CRITICAL: OSPFv2: 4/5 UP : OSPFv3: 4/5 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [10:50:19] (03PS3) 10Effie Mouzeli: memcached: implement refresh_certs [puppet] - 10https://gerrit.wikimedia.org/r/1320809 (https://phabricator.wikimedia.org/T353511) [10:52:42] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage [10:53:02] (03PS1) 10Ayounsi: HE drmrs-codfw OSPF: fix typo [homer/public] - 10https://gerrit.wikimedia.org/r/1320814 [10:53:05] (03CR) 10David Caro: [C:03+1] "❤️" [puppet] - 10https://gerrit.wikimedia.org/r/1318720 (owner: 10Andrew Bogott) [10:53:14] (03Abandoned) 10Ayounsi: Add OSPF for codfw HE transport [homer/public] - 10https://gerrit.wikimedia.org/r/1320807 (https://phabricator.wikimedia.org/T429653) (owner: 10Ayounsi) [10:53:17] FIRING: [2x] HelmReleaseBadStatus: Helm release wdqs/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [10:54:46] (03PS4) 10Effie Mouzeli: memcached: implement refresh_certs [puppet] - 10https://gerrit.wikimedia.org/r/1320809 (https://phabricator.wikimedia.org/T353511) [10:55:19] (03CR) 10Ayounsi: [C:03+2] HE drmrs-codfw OSPF: fix typo [homer/public] - 10https://gerrit.wikimedia.org/r/1320814 (owner: 10Ayounsi) [10:55:35] (03CR) 10Cathal Mooney: [C:03+1] HE drmrs-codfw OSPF: fix typo [homer/public] - 10https://gerrit.wikimedia.org/r/1320814 (owner: 10Ayounsi) [10:55:44] (03PS1) 10Majavah: hieradata: Remove data for sso Cloud VPS project [puppet] - 10https://gerrit.wikimedia.org/r/1320815 [10:55:45] (03PS1) 10Majavah: P:idp::standalone: Remove profile [puppet] - 10https://gerrit.wikimedia.org/r/1320816 [10:55:45] (03PS1) 10Majavah: P:idp: Remove support for memcached [puppet] - 10https://gerrit.wikimedia.org/r/1320817 (https://phabricator.wikimedia.org/T273950) [10:56:27] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage [10:56:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [10:56:43] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1069.eqiad.wmnet with OS trixie [10:56:50] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12182697 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-be1069.eqiad.wmnet with OS trixie completed: - ms-be1069 (**PASS*... [10:57:14] (03Merged) 10jenkins-bot: HE drmrs-codfw OSPF: fix typo [homer/public] - 10https://gerrit.wikimedia.org/r/1320814 (owner: 10Ayounsi) [10:57:43] (03CR) 10Majavah: [V:03+1] "PCC SUCCESS (CORE_DIFF 3 NOOP 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/" [puppet] - 10https://gerrit.wikimedia.org/r/1320817 (https://phabricator.wikimedia.org/T273950) (owner: 10Majavah) [11:00:07] !log jayme@cumin1003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1139.eqiad.wmnet with OS trixie [11:00:08] !log jayme@cumin1003 END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1139.eqiad.wmnet [11:01:29] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie [11:01:36] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12182716 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-be1096.eqiad.wmnet with OS trixie completed: - ms-be1096 (**PASS*... [11:01:46] (03PS1) 10Slyngshede: Geo-map: August update for Meta [dns] - 10https://gerrit.wikimedia.org/r/1320821 [11:02:13] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie [11:02:20] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12182723 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-be1097.eqiad.wmnet with OS trixie [11:02:27] !log jayme@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1140.eqiad.wmnet with OS trixie [11:03:26] (03PS1) 10JavierMonton: namespaces: webrequest-pageview [puppet] - 10https://gerrit.wikimedia.org/r/1320823 (https://phabricator.wikimedia.org/T433962) [11:04:20] (03PS2) 10Slyngshede: Geo-map: August update for Meta [dns] - 10https://gerrit.wikimedia.org/r/1320821 [11:04:37] !log mvernon@cumin1003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be1097.eqiad.wmnet with OS trixie [11:04:42] (03PS2) 10Majavah: P:idp: Remove support for memcached [puppet] - 10https://gerrit.wikimedia.org/r/1320817 (https://phabricator.wikimedia.org/T273950) [11:04:42] (03PS1) 10Majavah: hieradata: Migrate cloudidp-dev to Redis ticket storage [puppet] - 10https://gerrit.wikimedia.org/r/1320824 (https://phabricator.wikimedia.org/T433963) [11:04:46] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12182757 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-be1097.eqiad.wmnet with OS trixie executed with errors: - ms-be10... [11:05:01] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie [11:05:08] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12182760 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-be1097.eqiad.wmnet with OS trixie [11:06:55] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1071.eqiad.wmnet with OS trixie [11:07:09] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12182761 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-be1071.eqiad.wmnet with OS trixie completed: - ms-be1071 (**PASS*... [11:08:21] (03CR) 10Majavah: [V:03+1] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9109/co" [puppet] - 10https://gerrit.wikimedia.org/r/1320824 (https://phabricator.wikimedia.org/T433963) (owner: 10Majavah) [11:09:38] (03PS3) 10Majavah: P:idp: Remove support for memcached [puppet] - 10https://gerrit.wikimedia.org/r/1320817 (https://phabricator.wikimedia.org/T273950) [11:09:59] !log cwilliams@cumin1003 START - Cookbook sre.hosts.reboot-single for host db1278.eqiad.wmnet [11:13:12] (03PS1) 10JavierMonton: stream: webrequest-pageview [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320828 (https://phabricator.wikimedia.org/T433962) [11:13:17] !log jayme@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1140.eqiad.wmnet [11:13:18] !log jayme@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1140.eqiad.wmnet [11:13:18] (03PS1) 10Cathal Mooney: eqsin nokia switches: add to rancid config backup [puppet] - 10https://gerrit.wikimedia.org/r/1320829 (https://phabricator.wikimedia.org/T418439) [11:13:19] !log jayme@cumin1003 END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1140.eqiad.wmnet [11:13:55] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1070.eqiad.wmnet with OS trixie [11:14:03] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12182782 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-be1070.eqiad.wmnet with OS trixie completed: - ms-be1070 (**PASS*... [11:14:10] (03CR) 10Majavah: [V:03+1] "PCC SUCCESS (CORE_DIFF 3 NOOP 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/" [puppet] - 10https://gerrit.wikimedia.org/r/1320817 (https://phabricator.wikimedia.org/T273950) (owner: 10Majavah) [11:14:13] !log jayme@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1139.eqiad.wmnet [11:14:15] !log jayme@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1139.eqiad.wmnet [11:15:15] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, August 04 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployc" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320757 (https://phabricator.wikimedia.org/T432204) (owner: 10JavierMonton) [11:15:15] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1278.eqiad.wmnet [11:15:28] !log cwilliams@cumin1003 START - Cookbook sre.hosts.reboot-single for host db1279.eqiad.wmnet [11:15:47] (03CR) 10Ayounsi: [C:03+1] eqsin nokia switches: add to rancid config backup [puppet] - 10https://gerrit.wikimedia.org/r/1320829 (https://phabricator.wikimedia.org/T418439) (owner: 10Cathal Mooney) [11:16:25] !log mvernon@cumin2003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2095.codfw.wmnet with OS trixie [11:16:35] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12182792 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-be2095.codfw.wmnet with OS trixie completed: - ms-be2095 (**FAIL*... [11:16:36] !log jayme@cumin1003 START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1141.eqiad.wmnet [11:16:37] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12182793 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-be2095.codfw.wmnet with OS trixie executed with errors: - ms-be20... [11:16:40] !log jayme@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1141.eqiad.wmnet [11:17:16] !log jayme@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1141.eqiad.wmnet [11:18:37] (03CR) 10Cathal Mooney: [C:03+2] eqsin nokia switches: add to rancid config backup [puppet] - 10https://gerrit.wikimedia.org/r/1320829 (https://phabricator.wikimedia.org/T418439) (owner: 10Cathal Mooney) [11:20:17] jayme@cumin1003 renumber-node (PID 438480) is awaiting input [11:20:19] !log cmooney@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin with reason: upgrade new Nokia swtiches in eqsin to SR Linux v26 [11:20:45] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1279.eqiad.wmnet [11:20:48] (03PS5) 10Effie Mouzeli: memcached: implement refresh_certs [puppet] - 10https://gerrit.wikimedia.org/r/1320809 (https://phabricator.wikimedia.org/T353511) [11:20:58] !log cwilliams@cumin1003 START - Cookbook sre.hosts.reboot-single for host db1280.eqiad.wmnet [11:21:31] !log jayme@cumin1003 START - Cookbook sre.hosts.reimage for host wikikube-worker1141.eqiad.wmnet with OS trixie [11:22:00] !log jayme@cumin1003 START - Cookbook sre.hosts.move-vlan for host wikikube-worker1141 [11:22:02] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, August 04 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployc" [extensions/CheckUser] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1320781 (https://phabricator.wikimedia.org/T433816) (owner: 10Mszwarc) [11:22:13] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, August 04 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployc" [extensions/CheckUser] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1320782 (https://phabricator.wikimedia.org/T433816) (owner: 10Mszwarc) [11:22:26] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, August 04 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployc" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319086 (https://phabricator.wikimedia.org/T432452) (owner: 10Mpostoronca) [11:23:06] (03PS2) 10Slyngshede: Tofurkey: switch to a secrets file [labs/private] - 10https://gerrit.wikimedia.org/r/1320172 [11:24:13] !log mvernon@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage [11:25:04] jayme@cumin1003 renumber-node (PID 438480) is awaiting input [11:25:13] !log jayme@cumin1003 START - Cookbook sre.dns.netbox [11:26:33] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1280.eqiad.wmnet [11:26:46] !log cwilliams@cumin1003 START - Cookbook sre.hosts.reboot-single for host db1281.eqiad.wmnet [11:26:55] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie [11:27:03] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12182819 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host ms-be2095.codfw.wmnet with OS trixie [11:27:43] (03CR) 10Klausman: [C:03+1] namespaces: webrequest-pageview (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1320823 (https://phabricator.wikimedia.org/T433962) (owner: 10JavierMonton) [11:28:31] FIRING: ProbeDown: Service gerrit2003:443 has failed probes (http_gerrit_tls_ip6) - https://wikitech.wikimedia.org/wiki/Runbook#gerrit2003:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [11:29:28] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage [11:29:45] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage [11:31:16] jayme@cumin1003 renumber-node (PID 438480) is awaiting input [11:32:10] !log cmooney@cumin1003 START - Cookbook sre.network.tls for network device lswtest-d8-eqiad [11:32:17] !log cmooney@cumin1003 END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad [11:32:17] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1281.eqiad.wmnet [11:32:30] !log cwilliams@cumin1003 START - Cookbook sre.hosts.reboot-single for host db1282.eqiad.wmnet [11:32:32] !log jayme@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" [11:32:36] !log jayme@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003" [11:32:36] !log jayme@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [11:32:36] !log jayme@cumin1003 START - Cookbook sre.dns.wipe-cache wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [11:32:40] !log jayme@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [11:32:40] !log jayme@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1141 [11:32:57] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage [11:33:02] !log jayme@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1141 [11:33:03] !log jayme@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1141 [11:33:31] RESOLVED: ProbeDown: Service gerrit2003:443 has failed probes (http_gerrit_tls_ip6) - https://wikitech.wikimedia.org/wiki/Runbook#gerrit2003:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [11:34:35] (03CR) 10Effie Mouzeli: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1320809 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [11:37:04] !log cmooney@cumin1003 START - Cookbook sre.network.tls for network device lswtest-d8-eqiad [11:37:05] !log cmooney@cumin1003 END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad [11:37:23] FIRING: [4x] CertAlmostExpired: gNMI TLS certificate for asw1-603-eqsin.mgmt.eqsin.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [11:37:42] 06SRE, 06App Experience, 06Content-Platform-Team, 06ServiceOps, and 2 others: Investigate Code 414 error when selecting zh-classical (lzh) language from article toolbar - https://phabricator.wikimedia.org/T425545#12182832 (10Raine) >>! In T425545#12182590, @Phreelance wrote: > @Raine above you mentioned a... [11:37:42] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1282.eqiad.wmnet [11:37:56] !log cwilliams@cumin1003 START - Cookbook sre.hosts.reboot-single for host db1283.eqiad.wmnet [11:38:25] (03PS3) 10Slyngshede: Tofurkey: switch to a secrets file [labs/private] - 10https://gerrit.wikimedia.org/r/1320172 [11:39:04] !log cmooney@cumin1003 START - Cookbook sre.network.tls for network device lswtest-d8-eqiad [11:39:04] !log cmooney@cumin1003 END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad [11:40:28] PROBLEM - Postfix SMTP on crm2001 is CRITICAL: CRITICAL - Certificate crm2001.codfw.wmnet expires in 15 day(s) (Thu 20 Aug 2026 11:40:00 AM GMT +0000). https://wikitech.wikimedia.org/wiki/Mail%23Troubleshooting [11:42:23] FIRING: [4x] CertAlmostExpired: gNMI TLS certificate for asw1-603-eqsin.mgmt.eqsin.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [11:42:37] (03CR) 10Slyngshede: [V:03+2 C:03+2] Tofurkey: switch to a secrets file [labs/private] - 10https://gerrit.wikimedia.org/r/1320172 (owner: 10Slyngshede) [11:42:43] !log klausman@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows [11:43:12] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1283.eqiad.wmnet [11:43:26] !log cwilliams@cumin1003 START - Cookbook sre.hosts.reboot-single for host db1284.eqiad.wmnet [11:43:47] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2095.codfw.wmnet with OS trixie [11:43:58] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12182837 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-be2095.codfw.wmnet with OS trixie completed: - ms-be2095 (**PASS*... [11:48:43] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1284.eqiad.wmnet [11:48:57] !log cwilliams@cumin1003 START - Cookbook sre.hosts.reboot-single for host db1286.eqiad.wmnet [11:49:15] !log jayme@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage [11:51:44] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host ms-be2096.codfw.wmnet with OS trixie [11:51:59] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12182853 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host ms-be2096.codfw.wmnet with OS trixie [11:53:42] !log jayme@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage [11:54:14] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1286.eqiad.wmnet [11:54:28] !log cwilliams@cumin1003 START - Cookbook sre.hosts.reboot-single for host db1287.eqiad.wmnet [11:54:47] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1097.eqiad.wmnet with OS trixie [11:55:01] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12182865 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-be1097.eqiad.wmnet with OS trixie completed: - ms-be1097 (**PASS*... [11:56:29] 10ops-codfw, 06SRE, 06DC-Ops, 06ServiceOps: Rerack mc-gp2006 and mc2046 - https://phabricator.wikimedia.org/T432716#12182867 (10Clement_Goubert) >>! In T432716#12180337, @Jhancock.wm wrote: > it's magical to me. Updated the server location. I was gonna give the reimage a try but i didn't know which distro... [11:57:23] FIRING: [5x] CertAlmostExpired: gNMI TLS certificate for asw1-603-eqsin.mgmt.eqsin.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [11:59:40] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1287.eqiad.wmnet [11:59:54] !log cwilliams@cumin1003 START - Cookbook sre.hosts.reboot-single for host db1288.eqiad.wmnet [12:00:05] Deploy window Mobileapps/RESTBase/Wikifeeds (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260804T1200) [12:02:19] (03CR) 10JMeybohm: "First of all you can check the CI output for the "actual diff" of each kube-state-metrics deployment. This should give you a better idea o" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319069 (https://phabricator.wikimedia.org/T427405) (owner: 10Blake) [12:02:23] FIRING: [4x] CertAlmostExpired: gNMI TLS certificate for asw1-604-eqsin.mgmt.eqsin.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [12:05:11] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1288.eqiad.wmnet [12:05:25] !log cwilliams@cumin1003 START - Cookbook sre.hosts.reboot-single for host db1289.eqiad.wmnet [12:10:38] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1289.eqiad.wmnet [12:10:49] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage [12:12:56] (03CR) 10David Caro: webservice-runner: default to 8000 only if PORT and TOOL_WEB_PORT has no value (031 comment) [docker-images/toollabs-images] - 10https://gerrit.wikimedia.org/r/1310212 (https://phabricator.wikimedia.org/T432078) (owner: 10Raymond Ndibe) [12:14:07] !log jayme@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1141.eqiad.wmnet with OS trixie [12:14:46] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage [12:21:26] (03PS1) 10Cathal Mooney: eqsin: bump sr-linux verion to v26 after upgrade [homer/public] - 10https://gerrit.wikimedia.org/r/1320891 (https://phabricator.wikimedia.org/T418439) [12:22:44] !log mvernon@cumin1003 START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie [12:22:54] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12182909 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin1003 for host ms-be1096.eqiad.wmnet with OS trixie [12:27:38] (03PS1) 10Marostegui: mariadb: Move db1164 to m1 [puppet] - 10https://gerrit.wikimedia.org/r/1320898 (https://phabricator.wikimedia.org/T433464) [12:28:24] (03PS34) 10Slyngshede: P:tofurkey Add tofurkey [puppet] - 10https://gerrit.wikimedia.org/r/1260730 (https://phabricator.wikimedia.org/T355446) [12:29:00] (03CR) 10CI reject: [V:04-1] P:tofurkey Add tofurkey [puppet] - 10https://gerrit.wikimedia.org/r/1260730 (https://phabricator.wikimedia.org/T355446) (owner: 10Slyngshede) [12:31:03] (03PS35) 10Slyngshede: P:tofurkey Add tofurkey [puppet] - 10https://gerrit.wikimedia.org/r/1260730 (https://phabricator.wikimedia.org/T355446) [12:33:03] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2096.codfw.wmnet with OS trixie [12:33:09] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12182930 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-be2096.codfw.wmnet with OS trixie completed: - ms-be2096 (**PASS*... [12:33:16] (03CR) 10CI reject: [V:04-1] P:tofurkey Add tofurkey [puppet] - 10https://gerrit.wikimedia.org/r/1260730 (https://phabricator.wikimedia.org/T355446) (owner: 10Slyngshede) [12:38:15] !log marostegui@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1164,1217].eqiad.wmnet with reason: cloning [12:39:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.76% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [12:41:47] (03CR) 10Marostegui: [C:03+2] mariadb: Move db1164 to m1 [puppet] - 10https://gerrit.wikimedia.org/r/1320898 (https://phabricator.wikimedia.org/T433464) (owner: 10Marostegui) [12:42:13] !log mvernon@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage [12:42:54] !log jynus@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1150.eqiad.wmnet with reason: decom [12:43:25] !log jynus@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1171.eqiad.wmnet with reason: decom [12:43:48] jayme@cumin1003 renumber-node (PID 438480) is awaiting input [12:44:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.48% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [12:44:25] (03CR) 10Jcrespo: [C:03+2] mariadb: Set db1150 as insetup for decommissioning [puppet] - 10https://gerrit.wikimedia.org/r/1320682 (https://phabricator.wikimedia.org/T433825) (owner: 10Jcrespo) [12:44:42] (03CR) 10Jcrespo: [C:03+2] mariadb: Set db1171 as insetup for decommissioning [puppet] - 10https://gerrit.wikimedia.org/r/1320684 (https://phabricator.wikimedia.org/T433826) (owner: 10Jcrespo) [12:45:45] (03PS3) 10Jcrespo: backup: Set backup1003 & backup2003 as insetup for decom. [puppet] - 10https://gerrit.wikimedia.org/r/1320768 (https://phabricator.wikimedia.org/T420506) [12:46:43] (03PS3) 10Ebernhardson: cirrus: Enable building redirect documents [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320230 (https://phabricator.wikimedia.org/T204089) [12:49:03] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage [12:49:34] (03PS2) 10JavierMonton: namespaces: webrequest-pageview [puppet] - 10https://gerrit.wikimedia.org/r/1320823 (https://phabricator.wikimedia.org/T433962) [12:50:13] (03CR) 10CI reject: [V:04-1] namespaces: webrequest-pageview [puppet] - 10https://gerrit.wikimedia.org/r/1320823 (https://phabricator.wikimedia.org/T433962) (owner: 10JavierMonton) [12:55:39] jayme@cumin1003 renumber-node (PID 438480) is awaiting input [12:56:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.79% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [12:56:18] (03PS36) 10Slyngshede: P:tofurkey Add tofurkey [puppet] - 10https://gerrit.wikimedia.org/r/1260730 (https://phabricator.wikimedia.org/T355446) [12:57:52] (03CR) 10Clément Goubert: [C:03+1] "No-op in prod except for the api-gateway metrics, ship it." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319122 (https://phabricator.wikimedia.org/T428625) (owner: 10Aaron Schulz) [12:58:01] RESOLVED: HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-rollback - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs-next - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [12:58:05] (03CR) 10DCausse: [C:03+1] cirrus: Enable building redirect documents [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320230 (https://phabricator.wikimedia.org/T204089) (owner: 10Ebernhardson) [12:58:40] I'll be here for my deployments in 15-20 mins [12:58:59] (03PS1) 10Bartosz Wójtowicz: ml-services: declare ephemeral-storage for all ml-staging isvcs [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320919 (https://phabricator.wikimedia.org/T431089) [12:59:05] (I mean I'll be back then, not that the deployments have to happen at that time) [12:59:33] i'm here [13:00:05] Lucas_WMDE, urbanecm, and TheresNoTime: #bothumor Q:How do functions break up? A:They stop calling each other. Rise for UTC afternoon backport window deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260804T1300). [13:00:05] thedj, javiermonton, and Msz2001: A patch you scheduled for UTC afternoon backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [13:00:42] you should be fine to deploy, just be advised mw-web is under a bit of strain right now, I'm looking into it [13:01:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.79% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [13:01:25] RESOLVED: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:02:39] (03CR) 10Scott French: [C:03+2] wmnet: Reduce _etcd._tcp.conftool (R/W) SRV TTL to 10s [dns] - 10https://gerrit.wikimedia.org/r/1319194 (https://phabricator.wikimedia.org/T433554) (owner: 10Scott French) [13:03:04] !log swfrench@dns1004 START - running authdns-update [13:03:13] (03CR) 10Jcrespo: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1320768 (https://phabricator.wikimedia.org/T420506) (owner: 10Jcrespo) [13:04:36] \o [13:05:05] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12183010 (10MatthewVernon) [13:05:13] !log swfrench@dns1004 END - running authdns-update [13:06:02] (03PS4) 10Jcrespo: backup: Set backup1003 & backup2003 as insetup for decom. [puppet] - 10https://gerrit.wikimedia.org/r/1320768 (https://phabricator.wikimedia.org/T420506) [13:06:02] (03PS1) 10Jcrespo: mariadb: Remove all references on puppet to db1150 & db1171 [puppet] - 10https://gerrit.wikimedia.org/r/1320923 (https://phabricator.wikimedia.org/T433825) [13:06:17] (03PS1) 10Majavah: hieradata: codfw1dev: Update NTP server [puppet] - 10https://gerrit.wikimedia.org/r/1320924 (https://phabricator.wikimedia.org/T401810) [13:06:29] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie [13:06:31] (03PS1) 10Arnaudb: gerrit: concatenation typo on dns wipe-cache [cookbooks] - 10https://gerrit.wikimedia.org/r/1320921 [13:06:31] (03CR) 10Arnaudb: "for sanity check, this was overlooked previously so I prefer to ask for another set of eyes :-)" [cookbooks] - 10https://gerrit.wikimedia.org/r/1320921 (owner: 10Arnaudb) [13:06:35] 06SRE, 10SRE-swift-storage, 07Essential-Work: Migrate production swift clusters to trixie - https://phabricator.wikimedia.org/T429630#12183018 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin1003 for host ms-be1096.eqiad.wmnet with OS trixie completed: - ms-be1096 (**PASS*... [13:06:58] (03CR) 10Jcrespo: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1320768 (https://phabricator.wikimedia.org/T420506) (owner: 10Jcrespo) [13:07:45] (03CR) 10Majavah: [C:03+2] hieradata: codfw1dev: Update NTP server [puppet] - 10https://gerrit.wikimedia.org/r/1320924 (https://phabricator.wikimedia.org/T401810) (owner: 10Majavah) [13:10:28] (03CR) 10Jcrespo: [C:04-1] "Wait to ensure it doesn't need to be undone." [puppet] - 10https://gerrit.wikimedia.org/r/1320923 (https://phabricator.wikimedia.org/T433825) (owner: 10Jcrespo) [13:12:00] (03CR) 10Jcrespo: [C:03+2] backup: Set backup1003 & backup2003 as insetup for decom. [puppet] - 10https://gerrit.wikimedia.org/r/1320768 (https://phabricator.wikimedia.org/T420506) (owner: 10Jcrespo) [13:14:45] (03CR) 10Ottomata: [C:03+2] mw-page-html-feature-counts-change-enrich: produce to kafka main-eqiad [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319116 (https://phabricator.wikimedia.org/T433507) (owner: 10Ottomata) [13:14:45] I'm back [13:14:58] Is everybody still waiting for deployer? [13:15:02] (03CR) 10Andrew Bogott: [C:03+2] scan-and-shrink.py: expand some variable names [puppet] - 10https://gerrit.wikimedia.org/r/1318720 (owner: 10Andrew Bogott) [13:15:08] yep [13:15:12] I can deploy, then [13:16:34] (03PS3) 10Cathal Mooney: sre.network.tls: add SR-Linux v26 support [cookbooks] - 10https://gerrit.wikimedia.org/r/1319814 (https://phabricator.wikimedia.org/T433105) (owner: 10Ayounsi) [13:16:47] (03CR) 10TrainBranchBot: [C:03+2] "Approved by mszwarc@deploy1003 using scap backport" [extensions/Kartographer] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1320171 (https://phabricator.wikimedia.org/T433703) (owner: 10Jforrester) [13:16:48] (03CR) 10TrainBranchBot: [C:03+2] "Approved by mszwarc@deploy1003 using scap backport" [extensions/TimedMediaHandler] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1320159 (https://phabricator.wikimedia.org/T433666) (owner: 10Jforrester) [13:16:49] (03PS1) 10Majavah: hieradata: codfw1dev: Update codfw1dev ENC server [puppet] - 10https://gerrit.wikimedia.org/r/1320927 (https://phabricator.wikimedia.org/T401810) [13:17:12] (03Merged) 10jenkins-bot: mw-page-html-feature-counts-change-enrich: produce to kafka main-eqiad [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319116 (https://phabricator.wikimedia.org/T433507) (owner: 10Ottomata) [13:18:39] (03PS1) 10Jcrespo: bacula: Remove last references to backup1003 & backup2003 [puppet] - 10https://gerrit.wikimedia.org/r/1320928 (https://phabricator.wikimedia.org/T420506) [13:18:54] (03PS2) 10Jcrespo: bacula: Remove last references to backup1003 & backup2003 [puppet] - 10https://gerrit.wikimedia.org/r/1320928 (https://phabricator.wikimedia.org/T420506) [13:19:29] (03CR) 10Jcrespo: [C:04-1] "Blocked on actual decom script running." [puppet] - 10https://gerrit.wikimedia.org/r/1320928 (https://phabricator.wikimedia.org/T420506) (owner: 10Jcrespo) [13:20:11] (03CR) 10Majavah: [C:03+2] hieradata: codfw1dev: Update codfw1dev ENC server [puppet] - 10https://gerrit.wikimedia.org/r/1320927 (https://phabricator.wikimedia.org/T401810) (owner: 10Majavah) [13:20:24] (03PS3) 10JavierMonton: namespaces: webrequest-pageview [puppet] - 10https://gerrit.wikimedia.org/r/1320823 (https://phabricator.wikimedia.org/T433962) [13:20:29] (03PS2) 10Scott French: wmnet: Switch _etcd._tcp.conftool (R/W) SRV hosts to codfw [dns] - 10https://gerrit.wikimedia.org/r/1319195 (https://phabricator.wikimedia.org/T433554) [13:20:29] (03PS2) 10Scott French: wmnet: Restore _etcd._tcp.conftool (R/W) SRV TTL to 5M [dns] - 10https://gerrit.wikimedia.org/r/1319196 (https://phabricator.wikimedia.org/T433554) [13:22:17] (03PS4) 10Cathal Mooney: sre.network.tls: add SR-Linux v26 support [cookbooks] - 10https://gerrit.wikimedia.org/r/1319814 (https://phabricator.wikimedia.org/T433105) (owner: 10Ayounsi) [13:22:23] FIRING: [3x] CertAlmostExpired: gNMI TLS certificate for lsw1-d6-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [13:22:51] !log otto@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply [13:22:55] !log otto@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply [13:25:45] (03CR) 10Slyngshede: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1260730 (https://phabricator.wikimedia.org/T355446) (owner: 10Slyngshede) [13:26:30] (03CR) 10Ottomata: [C:03+2] eventstreams: expose page_html_feature_counts_change.v1 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319117 (https://phabricator.wikimedia.org/T433507) (owner: 10Ottomata) [13:27:49] FIRING: HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs-next - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [13:28:26] !log jayme@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1141.eqiad.wmnet [13:28:27] !log jayme@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1141.eqiad.wmnet [13:28:29] !log jayme@cumin1003 END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1141.eqiad.wmnet [13:28:38] !log jayme@cumin1003 START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet [13:28:41] !log jayme@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet [13:28:42] (03CR) 10Scott French: [C:03+2] hieradata: etcd read-only in eqiad [puppet] - 10https://gerrit.wikimedia.org/r/1319191 (https://phabricator.wikimedia.org/T433554) (owner: 10Scott French) [13:28:59] (03Merged) 10jenkins-bot: eventstreams: expose page_html_feature_counts_change.v1 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319117 (https://phabricator.wikimedia.org/T433507) (owner: 10Ottomata) [13:29:11] jouncebot: nowandnext [13:29:11] For the next 0 hour(s) and 30 minute(s): UTC afternoon backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260804T1300) [13:29:11] In 0 hour(s) and 30 minute(s): Test Kitchen UI Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260804T1400) [13:29:49] !log jayme@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet [13:31:10] !log otto@deploy1003 helmfile [staging] START helmfile.d/services/eventstreams: apply [13:31:38] !log otto@deploy1003 helmfile [staging] DONE helmfile.d/services/eventstreams: apply [13:31:52] !log otto@deploy1003 helmfile [codfw] START helmfile.d/services/eventstreams: apply [13:32:40] !log otto@deploy1003 helmfile [codfw] DONE helmfile.d/services/eventstreams: apply [13:32:42] (03PS5) 10Cathal Mooney: sre.network.tls: add SR-Linux v26 support [cookbooks] - 10https://gerrit.wikimedia.org/r/1319814 (https://phabricator.wikimedia.org/T433105) (owner: 10Ayounsi) [13:32:49] jayme@cumin1003 renumber-node (PID 461261) is awaiting input [13:32:52] !log otto@deploy1003 helmfile [eqiad] START helmfile.d/services/eventstreams: apply [13:33:03] !log jayme@cumin1003 END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1154.eqiad.wmnet [13:33:24] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [13:33:37] !log otto@deploy1003 helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply [13:33:48] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [13:34:39] (03CR) 10Scott French: [C:03+2] hieradata: switch etcd replication from codfw to eqiad [puppet] - 10https://gerrit.wikimedia.org/r/1319192 (https://phabricator.wikimedia.org/T433554) (owner: 10Scott French) [13:34:53] (03PS3) 10Scott French: hieradata: switch etcd replication from codfw to eqiad [puppet] - 10https://gerrit.wikimedia.org/r/1319192 (https://phabricator.wikimedia.org/T433554) [13:35:53] (03CR) 10Scott French: [C:03+2] hieradata: switch etcd replication from codfw to eqiad [puppet] - 10https://gerrit.wikimedia.org/r/1319192 (https://phabricator.wikimedia.org/T433554) (owner: 10Scott French) [13:36:25] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [13:36:47] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [13:36:50] (03Merged) 10jenkins-bot: Thumbnail: Exclude map figures from multimediaviewer [extensions/Kartographer] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1320171 (https://phabricator.wikimedia.org/T433703) (owner: 10Jforrester) [13:37:39] (03PS1) 10PipelineBot: wikifeeds: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320934 [13:38:00] (03Merged) 10jenkins-bot: TimedText: Include SRT-only sources when listing playback tracks [extensions/TimedMediaHandler] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1320159 (https://phabricator.wikimedia.org/T433666) (owner: 10Jforrester) [13:38:48] !log mszwarc@deploy1003 Started scap sync-world: Backport for [[gerrit:1320171|Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159|TimedText: Include SRT-only sources when listing playback tracks (T433666)]] [13:38:56] T433703: Interactive map fails to load - https://phabricator.wikimedia.org/T433703 [13:38:56] T427709: Replace tright and tleft with floatright and floatleft in - https://phabricator.wikimedia.org/T427709 [13:38:58] T433666: Captions are not displayed on Commons videos (.webm files) although exist .srt files - https://phabricator.wikimedia.org/T433666 [13:41:06] !log mszwarc@deploy1003 mszwarc, jforrester: Backport for [[gerrit:1320171|Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159|TimedText: Include SRT-only sources when listing playback tracks (T433666)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [13:41:53] thedj: Please verify [13:42:00] checking [13:42:06] (03PS37) 10Slyngshede: P:tofurkey Add tofurkey [puppet] - 10https://gerrit.wikimedia.org/r/1260730 (https://phabricator.wikimedia.org/T355446) [13:43:32] (03CR) 10Scott French: [C:03+2] wmnet: Switch _etcd._tcp.conftool (R/W) SRV hosts to codfw [dns] - 10https://gerrit.wikimedia.org/r/1319195 (https://phabricator.wikimedia.org/T433554) (owner: 10Scott French) [13:43:49] !log swfrench@dns1004 START - running authdns-update [13:43:55] confirmed ok [13:44:01] !log mszwarc@deploy1003 mszwarc, jforrester: Continuing with deployment [13:44:08] JavierMonton: Can I deploy your config patch together with mine patches, to speed things up? [13:45:47] !log swfrench@dns1004 END - running authdns-update [13:45:59] (03CR) 10Scott French: [C:03+2] hieradata: etcd read-write in codfw [puppet] - 10https://gerrit.wikimedia.org/r/1319193 (https://phabricator.wikimedia.org/T433554) (owner: 10Scott French) [13:46:51] (03CR) 10Slyngshede: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1260730 (https://phabricator.wikimedia.org/T355446) (owner: 10Slyngshede) [13:46:51] yes, please! [13:47:13] (03CR) 10Mszwarc: [C:03+2] "Ahead of deployment" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320757 (https://phabricator.wikimedia.org/T432204) (owner: 10JavierMonton) [13:47:22] (03CR) 10Mszwarc: [C:03+2] "Ahead of deployment" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319086 (https://phabricator.wikimedia.org/T432452) (owner: 10Mpostoronca) [13:47:48] (03CR) 10Mszwarc: [C:03+2] "Ahead of deployment" [extensions/CheckUser] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1320781 (https://phabricator.wikimedia.org/T433816) (owner: 10Mszwarc) [13:47:59] (03CR) 10Mszwarc: [C:03+2] "Ahead of deployment" [extensions/CheckUser] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1320782 (https://phabricator.wikimedia.org/T433816) (owner: 10Mszwarc) [13:48:01] FIRING: [2x] JobUnavailable: Reduced availability for job etcdmirror in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [13:48:07] !log mszwarc@deploy1003 Finished scap sync-world: Backport for [[gerrit:1320171|Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159|TimedText: Include SRT-only sources when listing playback tracks (T433666)]] (duration: 09m 19s) [13:48:14] T433703: Interactive map fails to load - https://phabricator.wikimedia.org/T433703 [13:48:15] T427709: Replace tright and tleft with floatright and floatleft in - https://phabricator.wikimedia.org/T427709 [13:48:15] T433666: Captions are not displayed on Commons videos (.webm files) although exist .srt files - https://phabricator.wikimedia.org/T433666 [13:48:20] (03Merged) 10jenkins-bot: stream: pageview.trending.relative.v1 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320757 (https://phabricator.wikimedia.org/T432204) (owner: 10JavierMonton) [13:48:25] (03Merged) 10jenkins-bot: WikimediaAntiAbuse: Document required load order after Echo [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319086 (https://phabricator.wikimedia.org/T432452) (owner: 10Mpostoronca) [13:48:45] (03CR) 10TrainBranchBot: [C:03+2] "Approved by mszwarc@deploy1003 using scap backport" [extensions/CheckUser] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1320781 (https://phabricator.wikimedia.org/T433816) (owner: 10Mszwarc) [13:48:45] (03CR) 10TrainBranchBot: [C:03+2] "Approved by mszwarc@deploy1003 using scap backport" [extensions/CheckUser] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1320782 (https://phabricator.wikimedia.org/T433816) (owner: 10Mszwarc) [13:49:23] !log swfrench@cumin2002 conftool action : set/pooled=no; selector: name=wikikube-worker2330.codfw.wmnet [13:49:35] \o/ [13:49:49] !log swfrench@cumin2002 conftool action : set/pooled=yes; selector: name=wikikube-worker2330.codfw.wmnet [13:49:52] (03Merged) 10jenkins-bot: UIC: Add user name to server-side instrumentation events [extensions/CheckUser] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1320781 (https://phabricator.wikimedia.org/T433816) (owner: 10Mszwarc) [13:52:58] (03PS1) 10Ottomata: eventstreams - fetch all stream configs [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320938 (https://phabricator.wikimedia.org/T433507) [13:56:28] (03PS38) 10Slyngshede: P:tofurkey Add tofurkey [puppet] - 10https://gerrit.wikimedia.org/r/1260730 (https://phabricator.wikimedia.org/T355446) [13:58:47] (03PS1) 10Effie Mouzeli: memcached: reload_certs on certificate renewal [puppet] - 10https://gerrit.wikimedia.org/r/1320940 (https://phabricator.wikimedia.org/T353511) [13:59:02] (03CR) 10CI reject: [V:04-1] P:tofurkey Add tofurkey [puppet] - 10https://gerrit.wikimedia.org/r/1260730 (https://phabricator.wikimedia.org/T355446) (owner: 10Slyngshede) [13:59:28] The CI is slower than I expected, this deployment will overflow into the Test Kitchen UI window [13:59:43] (03PS1) 10Ottomata: EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320941 (https://phabricator.wikimedia.org/T433507) [14:00:05] Deploy window Test Kitchen UI Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260804T1400) [14:00:06] !log bking@cumin2003 conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw [14:00:07] !log bking@cumin2003 conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw [14:00:07] !log bking@cumin2003 conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw [14:00:17] (03PS1) 10Brouberol: airflow-dumps: allow more tasks to be executed in parallel [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320942 [14:00:46] (03PS39) 10Slyngshede: P:tofurkey Add tofurkey [puppet] - 10https://gerrit.wikimedia.org/r/1260730 (https://phabricator.wikimedia.org/T355446) [14:01:18] (03PS1) 10AOkoth: miscweb: add pod_annotiations to deployment tpl [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320943 (https://phabricator.wikimedia.org/T433588) [14:01:42] (03CR) 10CI reject: [V:04-1] P:tofurkey Add tofurkey [puppet] - 10https://gerrit.wikimedia.org/r/1260730 (https://phabricator.wikimedia.org/T355446) (owner: 10Slyngshede) [14:02:05] (03PS3) 10Blake: kube-state-metrics: update to upstream 7.3.0. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319069 (https://phabricator.wikimedia.org/T427405) [14:02:56] (03CR) 10JavierMonton: [C:03+1] eventstreams - fetch all stream configs [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320938 (https://phabricator.wikimedia.org/T433507) (owner: 10Ottomata) [14:03:15] (03CR) 10AKhatun: [C:03+1] eventstreams - fetch all stream configs [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320938 (https://phabricator.wikimedia.org/T433507) (owner: 10Ottomata) [14:03:18] (03CR) 10JavierMonton: [C:03+1] EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320941 (https://phabricator.wikimedia.org/T433507) (owner: 10Ottomata) [14:03:21] !log bking@cumin2003 START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - T324335 [14:03:26] T324335: Index the OpenSearch server application logs in the central observability cluster, using ECS format - https://phabricator.wikimedia.org/T324335 [14:04:01] (03CR) 10Klausman: [C:03+1] airflow-dumps: allow more tasks to be executed in parallel [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320942 (owner: 10Brouberol) [14:04:18] (03PS2) 10Krinkle: MathML+MathJax rollout to phase 3 (Wikibooks) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311058 (https://phabricator.wikimedia.org/T271001) [14:04:19] (03PS2) 10Krinkle: MathML+MathJax rollout to phase 4 (Wikisource) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311059 (https://phabricator.wikimedia.org/T271001) [14:04:19] (03PS2) 10Krinkle: MathML+MathJax rollout to phase 5 (Wikipedia small/medium/canary) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311060 (https://phabricator.wikimedia.org/T271001) [14:04:19] (03PS2) 10Krinkle: MathML+MathJax rollout to phase 6 (remaining large Wikipedias) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311061 (https://phabricator.wikimedia.org/T271001) [14:04:28] (03CR) 10CI reject: [V:04-1] MathML+MathJax rollout to phase 3 (Wikibooks) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311058 (https://phabricator.wikimedia.org/T271001) (owner: 10Krinkle) [14:04:33] (03CR) 10CI reject: [V:04-1] MathML+MathJax rollout to phase 4 (Wikisource) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311059 (https://phabricator.wikimedia.org/T271001) (owner: 10Krinkle) [14:04:35] (03CR) 10CI reject: [V:04-1] MathML+MathJax rollout to phase 5 (Wikipedia small/medium/canary) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311060 (https://phabricator.wikimedia.org/T271001) (owner: 10Krinkle) [14:04:41] (03Merged) 10jenkins-bot: UIC: Add user name to server-side instrumentation events [extensions/CheckUser] (wmf/1.47.0-wmf.13) - 10https://gerrit.wikimedia.org/r/1320782 (https://phabricator.wikimedia.org/T433816) (owner: 10Mszwarc) [14:04:44] (03CR) 10CI reject: [V:04-1] MathML+MathJax rollout to phase 6 (remaining large Wikipedias) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311061 (https://phabricator.wikimedia.org/T271001) (owner: 10Krinkle) [14:05:02] !log mszwarc@deploy1003 Started scap sync-world: Backport for [[gerrit:1320757|stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086|WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781|UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782|UIC: Add user name to server-side instrumentation events (T433816)]] [14:05:10] T432204: Review `pageview` and `pageview.trending.relative` Kafka topic sizes. - https://phabricator.wikimedia.org/T432204 [14:05:10] T432452: Echo notifications when new items are flagged - https://phabricator.wikimedia.org/T432452 [14:05:11] T433816: UserInfoCard: Include target user name in server-side instrumentation - https://phabricator.wikimedia.org/T433816 [14:05:29] phew... All the patched got merged only now... Let's hope it'll be quicker now [14:05:30] (03CR) 10Effie Mouzeli: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1320940 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [14:06:16] (03PS1) 10Trueg: WDQSv2: Qlever index rebuild container [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320946 (https://phabricator.wikimedia.org/T432627) [14:06:28] (03PS3) 10Krinkle: MathML+MathJax rollout to phase 3 (Wikibooks) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311058 (https://phabricator.wikimedia.org/T271001) [14:06:49] (03PS3) 10Krinkle: MathML+MathJax rollout to phase 4 (Wikisource) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311059 (https://phabricator.wikimedia.org/T271001) [14:06:49] (03PS3) 10Krinkle: MathML+MathJax rollout to phase 5 (Wikipedia small/medium/canary) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311060 (https://phabricator.wikimedia.org/T271001) [14:06:49] (03PS3) 10Krinkle: MathML+MathJax rollout to phase 6 (remaining large Wikipedias) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311061 (https://phabricator.wikimedia.org/T271001) [14:07:02] (03PS6) 10Effie Mouzeli: memcached: implement refresh_certs [puppet] - 10https://gerrit.wikimedia.org/r/1320809 (https://phabricator.wikimedia.org/T353511) [14:07:05] !log mszwarc@deploy1003 javiermonton, mszwarc, mpostoronca: Backport for [[gerrit:1320757|stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086|WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781|UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782|UIC: Add user name to server-side instrumentation events (T433816)]] synced to the testser [14:07:05] vers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [14:08:08] JavierMonton: Please verify [14:08:28] (03CR) 10Ottomata: [C:03+2] eventstreams - fetch all stream configs [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320938 (https://phabricator.wikimedia.org/T433507) (owner: 10Ottomata) [14:08:46] done! it's just a config and it appears now in the debug server [14:08:58] !log mszwarc@deploy1003 javiermonton, mszwarc, mpostoronca: Continuing with deployment [14:09:45] (03PS2) 10Effie Mouzeli: memcached: reload_certs on certificate renewal [puppet] - 10https://gerrit.wikimedia.org/r/1320940 (https://phabricator.wikimedia.org/T353511) [14:09:48] (03CR) 10CI reject: [V:04-1] kube-state-metrics: update to upstream 7.3.0. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319069 (https://phabricator.wikimedia.org/T427405) (owner: 10Blake) [14:10:19] (03CR) 10Effie Mouzeli: [C:03+1] api-gateway: further simplify chart by assuming REST Gateway mode #2 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319122 (https://phabricator.wikimedia.org/T428625) (owner: 10Aaron Schulz) [14:10:55] (03Merged) 10jenkins-bot: eventstreams - fetch all stream configs [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320938 (https://phabricator.wikimedia.org/T433507) (owner: 10Ottomata) [14:11:10] (03PS4) 10Blake: kube-state-metrics: update to upstream 7.3.0. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319069 (https://phabricator.wikimedia.org/T427405) [14:11:15] (03CR) 10Scott French: [C:03+2] wmnet: Restore _etcd._tcp.conftool (R/W) SRV TTL to 5M [dns] - 10https://gerrit.wikimedia.org/r/1319196 (https://phabricator.wikimedia.org/T433554) (owner: 10Scott French) [14:11:15] (03CR) 10Effie Mouzeli: [C:03+1] api-gateway:: remove API Gateway specific ratelimiting logic #3 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319563 (https://phabricator.wikimedia.org/T428625) (owner: 10Aaron Schulz) [14:11:25] !log swfrench@dns1004 START - running authdns-update [14:11:56] (03PS7) 10Effie Mouzeli: memcached: implement refresh_certs [puppet] - 10https://gerrit.wikimedia.org/r/1320809 (https://phabricator.wikimedia.org/T353511) [14:12:03] (03CR) 10Effie Mouzeli: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1320809 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [14:12:53] (03PS3) 10Effie Mouzeli: memcached: reload_certs on certificate renewal [puppet] - 10https://gerrit.wikimedia.org/r/1320940 (https://phabricator.wikimedia.org/T353511) [14:12:58] (03CR) 10Effie Mouzeli: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1320940 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [14:12:58] (03PS1) 10Majavah: hieradata: codfw1dev: Update dumps-mounting host [puppet] - 10https://gerrit.wikimedia.org/r/1320949 (https://phabricator.wikimedia.org/T401810) [14:13:00] !log mszwarc@deploy1003 Finished scap sync-world: Backport for [[gerrit:1320757|stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086|WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781|UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782|UIC: Add user name to server-side instrumentation events (T433816)]] (duration: 07m 58s) [14:13:06] !log otto@deploy1003 helmfile [staging] START helmfile.d/services/eventstreams: apply [14:13:08] T432204: Review `pageview` and `pageview.trending.relative` Kafka topic sizes. - https://phabricator.wikimedia.org/T432204 [14:13:08] T432452: Echo notifications when new items are flagged - https://phabricator.wikimedia.org/T432452 [14:13:09] T433816: UserInfoCard: Include target user name in server-side instrumentation - https://phabricator.wikimedia.org/T433816 [14:13:26] !log Finished deployments for UTC afternoon backport window [14:13:29] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:13:29] !log swfrench@dns1004 END - running authdns-update [14:13:30] !log otto@deploy1003 helmfile [staging] DONE helmfile.d/services/eventstreams: apply [14:13:38] (03PS1) 10MVernon: swift: add two new codfw backends [puppet] - 10https://gerrit.wikimedia.org/r/1320950 (https://phabricator.wikimedia.org/T424892) [14:13:55] 10ops-codfw, 06SRE, 06DC-Ops, 13Patch-For-Review: Q4:rack/setup/install sretest2010 Config J 1P test host - https://phabricator.wikimedia.org/T394357#12183512 (10Jhancock.wm) @MatthewVernon option 3 swap is complete [14:14:19] (03CR) 10BBlack: [V:03+1] "LGTM, Thank you!" [puppet] - 10https://gerrit.wikimedia.org/r/1260730 (https://phabricator.wikimedia.org/T355446) (owner: 10Slyngshede) [14:14:45] !log otto@deploy1003 helmfile [codfw] START helmfile.d/services/eventstreams: apply [14:15:40] !log otto@deploy1003 helmfile [codfw] DONE helmfile.d/services/eventstreams: apply [14:16:16] (03CR) 10Majavah: [C:03+2] hieradata: codfw1dev: Update dumps-mounting host [puppet] - 10https://gerrit.wikimedia.org/r/1320949 (https://phabricator.wikimedia.org/T401810) (owner: 10Majavah) [14:16:21] !log otto@deploy1003 helmfile [eqiad] START helmfile.d/services/eventstreams: apply [14:16:54] (03CR) 10Dzahn: [C:03+1] gerrit: concatenation typo on dns wipe-cache [cookbooks] - 10https://gerrit.wikimedia.org/r/1320921 (owner: 10Arnaudb) [14:17:09] !log otto@deploy1003 helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply [14:18:35] and thank you Msz2001 [14:19:24] (03CR) 10CI reject: [V:04-1] kube-state-metrics: update to upstream 7.3.0. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319069 (https://phabricator.wikimedia.org/T427405) (owner: 10Blake) [14:21:52] (03PS1) 10MVernon: swift: restore ms-be10[69-71] to rings, drain last 3 old-style nodes [puppet] - 10https://gerrit.wikimedia.org/r/1320955 (https://phabricator.wikimedia.org/T429630) [14:23:48] (03CR) 10Ottomata: [C:03+2] EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320941 (https://phabricator.wikimedia.org/T433507) (owner: 10Ottomata) [14:24:29] (03CR) 10Ayounsi: [C:03+1] eqsin: bump sr-linux verion to v26 after upgrade [homer/public] - 10https://gerrit.wikimedia.org/r/1320891 (https://phabricator.wikimedia.org/T418439) (owner: 10Cathal Mooney) [14:25:19] (03Merged) 10jenkins-bot: EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320941 (https://phabricator.wikimedia.org/T433507) (owner: 10Ottomata) [14:26:08] (03CR) 10Aaron Schulz: [C:03+2] api-gateway: further simplify chart by assuming REST Gateway mode #2 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319122 (https://phabricator.wikimedia.org/T428625) (owner: 10Aaron Schulz) [14:26:09] !log jayme@cumin1003 START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet [14:26:14] !log jayme@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet [14:26:15] !log otto@deploy1003 Started scap sync-world: Backport for [[gerrit:1320941|EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] [14:26:16] !log jayme@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet [14:26:19] T433507: Expose mediawiki.page_html_feature_counts_change.v1 in public EventStreams API - https://phabricator.wikimedia.org/T433507 [14:26:26] FYI am deploying a mw config change and rolling restarting eventgate-main... [14:26:29] !log jayme@cumin1003 START - Cookbook sre.hosts.reimage for host wikikube-worker1154.eqiad.wmnet with OS trixie [14:26:57] !log jayme@cumin1003 START - Cookbook sre.hosts.move-vlan for host wikikube-worker1154 [14:28:13] !log otto@deploy1003 otto: Backport for [[gerrit:1320941|EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [14:28:43] (03Merged) 10jenkins-bot: api-gateway: further simplify chart by assuming REST Gateway mode #2 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319122 (https://phabricator.wikimedia.org/T428625) (owner: 10Aaron Schulz) [14:30:00] jayme@cumin1003 renumber-node (PID 467549) is awaiting input [14:30:05] Deploy window Test Kitchen Experiment Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260804T1430) [14:30:06] !log jayme@cumin1003 START - Cookbook sre.dns.netbox [14:30:55] !log otto@deploy1003 otto: Continuing with deployment [14:34:16] !log jayme@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" [14:34:21] !log jayme@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003" [14:34:21] !log jayme@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [14:34:21] !log jayme@cumin1003 START - Cookbook sre.dns.wipe-cache wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [14:34:25] !log jayme@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [14:34:26] !log jayme@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1154 [14:34:54] !log otto@deploy1003 Finished scap sync-world: Backport for [[gerrit:1320941|EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] (duration: 08m 39s) [14:34:59] T433507: Expose mediawiki.page_html_feature_counts_change.v1 in public EventStreams API - https://phabricator.wikimedia.org/T433507 [14:35:49] (03CR) 10Brouberol: [C:03+2] airflow-dumps: allow more tasks to be executed in parallel [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320942 (owner: 10Brouberol) [14:36:09] 10ops-codfw, 06SRE, 06DC-Ops, 13Patch-For-Review: Q4:rack/setup/install sretest2010 Config J 1P test host - https://phabricator.wikimedia.org/T394357#12183696 (10MatthewVernon) Thanks. That swap went OK - the two drives just re-appeared in each other's places and could be remounted. Before the swap I did:... [14:36:10] (03PS1) 10CDanis: airflow: production: add cloudflare_radar connection [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320964 (https://phabricator.wikimedia.org/T433982) [14:36:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 20.79% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [14:36:27] !log jayme@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1154 [14:36:27] !log jayme@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1154 [14:36:43] 06SRE, 06Data-Platform-SRE, 06Product-Analytics, 13Patch-For-Review: Add cloudflare_radar API connection to Airflow analytics main instance - https://phabricator.wikimedia.org/T433982#12183699 (10KCVelaga_WMF) [14:36:54] !log otto@deploy1003 helmfile [staging] START helmfile.d/services/eventgate-main: sync [14:37:01] !log otto@deploy1003 helmfile [staging] DONE helmfile.d/services/eventgate-main: sync [14:37:17] !log roll restart eventgate-main to pick up stream config change - T433507 [14:37:20] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:37:36] !log otto@deploy1003 helmfile [codfw] START helmfile.d/services/eventgate-main: sync [14:38:02] !log otto@deploy1003 helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync [14:38:25] (03PS1) 10Bking: cirrussearch: Fix blackbox check [puppet] - 10https://gerrit.wikimedia.org/r/1320967 (https://phabricator.wikimedia.org/T431304) [14:38:29] !log otto@deploy1003 helmfile [eqiad] START helmfile.d/services/eventgate-main: sync [14:38:52] !log otto@deploy1003 helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync [14:39:00] (03CR) 10Brouberol: [C:03+1] "Looks great! FWIW you can also use `extra_dejson` as a dict, instead of `extra` that expects a json-encoded string." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320964 (https://phabricator.wikimedia.org/T433982) (owner: 10CDanis) [14:39:01] (03PS2) 10Blake: admin_ng: pin kube-state-metrics chart to current version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320945 (https://phabricator.wikimedia.org/T427405) [14:39:03] (03CR) 10CI reject: [V:04-1] cirrussearch: Fix blackbox check [puppet] - 10https://gerrit.wikimedia.org/r/1320967 (https://phabricator.wikimedia.org/T431304) (owner: 10Bking) [14:40:40] PROBLEM - Host wikikube-worker1271 is DOWN: PING CRITICAL - Packet loss = 80%, RTA = 8161.97 ms [14:40:58] RECOVERY - Host wikikube-worker1271 is UP: PING OK - Packet loss = 0%, RTA = 0.24 ms [14:41:04] (03PS2) 10CDanis: airflow: production: add cloudflare_radar connection [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320964 (https://phabricator.wikimedia.org/T433982) [14:41:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.28% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [14:42:40] (03PS2) 10Bking: cirrussearch: Fix blackbox check [puppet] - 10https://gerrit.wikimedia.org/r/1320967 (https://phabricator.wikimedia.org/T431304) [14:42:57] (03CR) 10KCVelaga: "Should the connection be added to analytics-main, as that is where the DAG will live, instead of platform-eng? Not sure if the connections" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320964 (https://phabricator.wikimedia.org/T433982) (owner: 10CDanis) [14:43:43] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1320967 (https://phabricator.wikimedia.org/T431304) (owner: 10Bking) [14:43:44] (03PS5) 10Blake: kube-state-metrics: update to upstream 7.3.0. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319069 (https://phabricator.wikimedia.org/T427405) [14:45:25] (03PS3) 10CDanis: airflow: production: add cloudflare_radar connection [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320964 (https://phabricator.wikimedia.org/T433982) [14:45:37] (03PS1) 10Brouberol: Fix airflow config structure [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320976 [14:45:45] (03CR) 10CDanis: "Very good catch, fixed." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320964 (https://phabricator.wikimedia.org/T433982) (owner: 10CDanis) [14:45:46] (03PS2) 10Brouberol: Fix airflow config structure [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320976 [14:45:47] (03CR) 10CI reject: [V:04-1] Fix airflow config structure [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320976 (owner: 10Brouberol) [14:46:39] (03PS4) 10CDanis: airflow: production: add cloudflare_radar connection [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320964 (https://phabricator.wikimedia.org/T433982) [14:47:05] (03PS3) 10Brouberol: Fix airflow config structure [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320976 [14:48:07] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply [14:48:38] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply [14:49:18] 10ops-eqsin, 06SRE, 06DC-Ops, 06Traffic: cp5022 is unreachable - https://phabricator.wikimedia.org/T414411#12183764 (10RobH) pre-onsite checks: host mgmt is online and i can connect to it, but I'm having issues on its idrac interface. I want to push new ilom firmware to it before handoff, but having issue... [14:51:19] (03CR) 10CWilliams: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1320950 (https://phabricator.wikimedia.org/T424892) (owner: 10MVernon) [14:51:47] !log jayme@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage [14:52:37] (03CR) 10Bking: [C:03+2] "self-merging, as alerts are currently broken and PCC is coming back cleanly." [puppet] - 10https://gerrit.wikimedia.org/r/1320967 (https://phabricator.wikimedia.org/T431304) (owner: 10Bking) [14:52:53] (03CR) 10CWilliams: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1320955 (https://phabricator.wikimedia.org/T429630) (owner: 10MVernon) [14:54:18] (03CR) 10Brouberol: [C:03+2] Fix airflow config structure [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320976 (owner: 10Brouberol) [14:54:31] (03CR) 10Scott French: [C:03+1] httpbb: Add a --request-header argument. [software/httpbb] - 10https://gerrit.wikimedia.org/r/1318682 (https://phabricator.wikimedia.org/T428972) (owner: 10Blake) [14:55:25] !log jayme@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage [14:56:38] (03PS33) 10Federico Ceratto: mysql: update replication source [cookbooks] - 10https://gerrit.wikimedia.org/r/1238368 (https://phabricator.wikimedia.org/T373436) [14:57:26] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2096 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [14:57:26] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2095 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [14:57:26] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2091 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [14:57:26] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2089 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [14:57:26] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2080 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [14:57:26] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2082 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [14:57:26] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2103 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [14:57:27] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2108 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [14:57:27] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2062 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [14:57:28] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2094 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [14:57:28] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2079 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [14:57:29] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2078 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [14:57:29] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2075 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [14:57:30] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2071 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [14:57:30] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2064 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [14:57:31] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2076 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [14:57:31] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2072 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [14:57:32] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2085 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [14:57:32] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2068 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [14:57:33] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2101 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [14:57:33] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2066 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [14:57:34] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2069 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [14:57:34] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2107 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [14:57:35] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2110 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [14:57:35] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2113 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [14:57:36] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2109 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [14:57:36] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2115 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [14:57:37] PROBLEM - PyBal backends health check on lvs2014 is CRITICAL: PYBAL CRITICAL - CRITICAL - search-psi-https_9643: Servers cirrussearch2091.codfw.wmnet, cirrussearch2082.codfw.wmnet, cirrussearch2113.codfw.wmnet, cirrussearch2079.codfw.wmnet, cirrussearch2094.codfw.wmnet, cirrussearch2089.codfw.wmnet, cirrussearch2068.codfw.wmnet, cirrussearch2066.codfw.wmnet, cirrussearch2078.codfw.wmnet, cirrussearch2101.codfw.wmnet, cirrussearch2075.codf [14:57:37] cirrussearch2064.codfw.wmnet, cirrussearch2108.codfw.wmnet, cirrussearch2110.codfw.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [14:57:38] PROBLEM - PyBal backends health check on lvs2013 is CRITICAL: PYBAL CRITICAL - CRITICAL - search-psi-https_9643: Servers cirrussearch2062.codfw.wmnet, cirrussearch2071.codfw.wmnet, cirrussearch2072.codfw.wmnet, cirrussearch2068.codfw.wmnet, cirrussearch2079.codfw.wmnet, cirrussearch2085.codfw.wmnet, cirrussearch2069.codfw.wmnet, cirrussearch2082.codfw.wmnet, cirrussearch2080.codfw.wmnet, cirrussearch2066.codfw.wmnet, cirrussearch2078.codf [14:57:38] cirrussearch2075.codfw.wmnet, cirrussearch2064.codfw.wmnet, cirrussearch2094.codfw.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [14:57:42] PROBLEM - ElasticSearch health check for shards on 9643 on search.svc.codfw.wmnet is CRITICAL: CRITICAL - elasticsearch https://search.svc.codfw.wmnet:9643/_cluster/health error while fetching: HTTPSConnectionPool(host=search.svc.codfw.wmnet, port=9643): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [14:58:51] !log dzahn@cumin2002 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment [14:59:03] (03CR) 10Blake: [C:03+2] httpbb: Add a --request-header argument. [software/httpbb] - 10https://gerrit.wikimedia.org/r/1318682 (https://phabricator.wikimedia.org/T428972) (owner: 10Blake) [14:59:28] !log dzahn@cumin2002 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment [14:59:51] !log dzahn@cumin2002 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment [14:59:59] ^^ sorry for the spam, we have depooled the DC but I guess the pybal alerts are active? [15:00:05] jelto, arnoldokoth, mutante, and arnaudb: SRE Collaboration Services office hours (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260804T1500). Please do the needful. [15:00:18] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch2095 is OK: OK - elasticsearch status production-search-psi-codfw: cluster_name: production-search-psi-codfw, status: yellow, timed_out: False, number_of_nodes: 28, number_of_data_nodes: 28, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1739, active_shards: 5029, relocating_shards: 0, initializing_shards: 3, unassigned_shards: 172, del [15:00:18] ssigned_shards: 172, number_of_pending_tasks: 2, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 103, active_shards_percent_as_number: 96.63720215219062 https://wikitech.wikimedia.org/wiki/Search%23Administration [15:00:18] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch2096 is OK: OK - elasticsearch status production-search-psi-codfw: cluster_name: production-search-psi-codfw, status: yellow, timed_out: False, number_of_nodes: 28, number_of_data_nodes: 28, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1739, active_shards: 5029, relocating_shards: 0, initializing_shards: 3, unassigned_shards: 172, del [15:00:18] ssigned_shards: 172, number_of_pending_tasks: 2, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 117, active_shards_percent_as_number: 96.63720215219062 https://wikitech.wikimedia.org/wiki/Search%23Administration [15:00:18] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch2089 is OK: OK - elasticsearch status production-search-psi-codfw: cluster_name: production-search-psi-codfw, status: yellow, timed_out: False, number_of_nodes: 28, number_of_data_nodes: 28, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1739, active_shards: 5029, relocating_shards: 0, initializing_shards: 3, unassigned_shards: 172, del [15:00:18] ssigned_shards: 172, number_of_pending_tasks: 2, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 119, active_shards_percent_as_number: 96.63720215219062 https://wikitech.wikimedia.org/wiki/Search%23Administration [15:00:18] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch2091 is OK: OK - elasticsearch status production-search-psi-codfw: cluster_name: production-search-psi-codfw, status: yellow, timed_out: False, number_of_nodes: 28, number_of_data_nodes: 28, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1739, active_shards: 5029, relocating_shards: 0, initializing_shards: 5, unassigned_shards: 170, del [15:01:41] !log brennen@deploy1003 Started deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for T433981 [15:01:45] T433981: Deploy Phab/Phorge 2026-08-04 - https://phabricator.wikimedia.org/T433981 [15:02:02] (03Merged) 10jenkins-bot: httpbb: Add a --request-header argument. [software/httpbb] - 10https://gerrit.wikimedia.org/r/1318682 (https://phabricator.wikimedia.org/T428972) (owner: 10Blake) [15:02:33] !log brennen@deploy1003 Finished deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for T433981 (duration: 00m 51s) [15:04:26] !log aaron@deploy1003 helmfile [codfw] START helmfile.d/services/rest-gateway: apply [15:05:22] !log aaron@deploy1003 helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply [15:05:38] !log brennen@deploy1003 Started deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for T433981 [15:06:21] !log brennen@deploy1003 Finished deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for T433981 (duration: 00m 43s) [15:11:31] FIRING: ProbeDown: Service gerrit2003:443 has failed probes (http_gerrit_tls_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#gerrit2003:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [15:11:59] (03CR) 10BBlack: [C:03+1] "(fix voting on the wrong row)" [puppet] - 10https://gerrit.wikimedia.org/r/1260730 (https://phabricator.wikimedia.org/T355446) (owner: 10Slyngshede) [15:13:26] PROBLEM - OpenSearch health check for shards on 9200 on cirrussearch2095 is CRITICAL: CRITICAL - elasticsearch http://localhost:9200/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9200): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:13:26] PROBLEM - OpenSearch health check for shards on 9200 on cirrussearch2090 is CRITICAL: CRITICAL - elasticsearch http://localhost:9200/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9200): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:13:26] PROBLEM - OpenSearch health check for shards on 9200 on cirrussearch2097 is CRITICAL: CRITICAL - elasticsearch http://localhost:9200/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9200): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:13:26] PROBLEM - OpenSearch health check for shards on 9200 on cirrussearch2096 is CRITICAL: CRITICAL - elasticsearch http://localhost:9200/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9200): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:13:26] PROBLEM - OpenSearch health check for shards on 9200 on cirrussearch2089 is CRITICAL: CRITICAL - elasticsearch http://localhost:9200/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9200): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:13:27] PROBLEM - OpenSearch health check for shards on 9200 on cirrussearch2083 is CRITICAL: CRITICAL - elasticsearch http://localhost:9200/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9200): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:13:27] PROBLEM - OpenSearch health check for shards on 9200 on cirrussearch2064 is CRITICAL: CRITICAL - elasticsearch http://localhost:9200/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9200): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:14:05] ^^ expected, sorry for the spam [15:14:18] RECOVERY - OpenSearch health check for shards on 9200 on cirrussearch2096 is OK: OK - elasticsearch status production-search-codfw: cluster_name: production-search-codfw, status: yellow, timed_out: False, number_of_nodes: 55, number_of_data_nodes: 55, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1263, active_shards: 3733, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 28, delayed_unas [15:14:18] hards: 28, number_of_pending_tasks: 1, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 99.25551714969423 https://wikitech.wikimedia.org/wiki/Search%23Administration [15:14:18] RECOVERY - OpenSearch health check for shards on 9200 on cirrussearch2097 is OK: OK - elasticsearch status production-search-codfw: cluster_name: production-search-codfw, status: yellow, timed_out: False, number_of_nodes: 55, number_of_data_nodes: 55, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1263, active_shards: 3733, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 28, delayed_unas [15:14:18] hards: 28, number_of_pending_tasks: 1, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 99.25551714969423 https://wikitech.wikimedia.org/wiki/Search%23Administration [15:14:18] RECOVERY - OpenSearch health check for shards on 9200 on cirrussearch2089 is OK: OK - elasticsearch status production-search-codfw: cluster_name: production-search-codfw, status: yellow, timed_out: False, number_of_nodes: 55, number_of_data_nodes: 55, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1263, active_shards: 3733, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 28, delayed_unas [15:14:58] (03PS34) 10Federico Ceratto: mysql: update replication source [cookbooks] - 10https://gerrit.wikimedia.org/r/1238368 (https://phabricator.wikimedia.org/T373436) [15:15:07] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [15:15:16] !log fceratto@cumin1003 END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) [15:15:38] (03CR) 10AOkoth: [C:03+1] gerrit: concatenation typo on dns wipe-cache [cookbooks] - 10https://gerrit.wikimedia.org/r/1320921 (owner: 10Arnaudb) [15:16:12] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [15:16:18] !log jayme@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1154.eqiad.wmnet with OS trixie [15:16:21] !log fceratto@cumin1003 END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) [15:16:31] RESOLVED: ProbeDown: Service gerrit2003:443 has failed probes (http_gerrit_tls_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#gerrit2003:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [15:16:58] PROBLEM - Postgres Replication Lag on puppetdb2003 is CRITICAL: POSTGRES_HOT_STANDBY_DELAY CRITICAL: DB puppetdb (host:localhost) 40112128 and 1 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [15:17:02] PROBLEM - Check unit status of push_cross_cluster_settings_9200 on cirrussearch2111 is CRITICAL: CRITICAL: Status of the systemd unit push_cross_cluster_settings_9200 https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [15:17:35] 06SRE, 10SRE-Access-Requests, 06Release-Engineering-Team: Requesting access to Jenkins server for PWangai-WMF - https://phabricator.wikimedia.org/T433563#12183884 (10pwangai) @Dzahn This will not involve sudo, I will be running a few scripts under my username to summarize the log data we need. [15:17:58] RECOVERY - Postgres Replication Lag on puppetdb2003 is OK: POSTGRES_HOT_STANDBY_DELAY OK: DB puppetdb (host:localhost) 2680656 and 0 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [15:17:58] (03CR) 10MVernon: [C:03+2] swift: add two new codfw backends [puppet] - 10https://gerrit.wikimedia.org/r/1320950 (https://phabricator.wikimedia.org/T424892) (owner: 10MVernon) [15:18:30] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2110 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:18:30] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2109 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:18:30] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2113 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:18:30] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2107 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:18:32] PROBLEM - PyBal backends health check on lvs2014 is CRITICAL: PYBAL CRITICAL - CRITICAL - search-psi-https_9643: Servers cirrussearch2071.codfw.wmnet, cirrussearch2094.codfw.wmnet, cirrussearch2113.codfw.wmnet, cirrussearch2079.codfw.wmnet, cirrussearch2085.codfw.wmnet, cirrussearch2069.codfw.wmnet, cirrussearch2089.codfw.wmnet, cirrussearch2082.codfw.wmnet, cirrussearch2075.codfw.wmnet, cirrussearch2076.codfw.wmnet, cirrussearch2101.codf [15:18:32] cirrussearch2066.codfw.wmnet, cirrussearch2064.codfw.wmnet, cirrussearch2108.codfw.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [15:18:44] PROBLEM - ElasticSearch health check for shards on 9643 on search.svc.codfw.wmnet is CRITICAL: CRITICAL - elasticsearch https://search.svc.codfw.wmnet:9643/_cluster/health error while fetching: HTTPSConnectionPool(host=search.svc.codfw.wmnet, port=9643): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:18:45] (03CR) 10MVernon: [C:03+2] swift: restore ms-be10[69-71] to rings, drain last 3 old-style nodes [puppet] - 10https://gerrit.wikimedia.org/r/1320955 (https://phabricator.wikimedia.org/T429630) (owner: 10MVernon) [15:19:26] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2096 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:19:26] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2095 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:19:26] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2089 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:19:26] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2091 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:19:26] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2068 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:19:27] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2085 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:19:27] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2064 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:19:28] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2078 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:19:28] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2083 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:19:29] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2069 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:19:29] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2080 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:19:30] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2108 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:19:30] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2076 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:19:31] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2101 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:19:31] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2079 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:19:32] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2103 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:19:32] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2071 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:19:33] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2075 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:19:33] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2082 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:19:34] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2062 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:19:34] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2094 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:19:35] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2072 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:19:35] PROBLEM - OpenSearch health check for shards on 9600 on cirrussearch2066 is CRITICAL: CRITICAL - elasticsearch http://localhost:9600/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9600): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:19:36] PROBLEM - PyBal backends health check on lvs2013 is CRITICAL: PYBAL CRITICAL - CRITICAL - search-psi-https_9643: Servers cirrussearch2071.codfw.wmnet, cirrussearch2062.codfw.wmnet, cirrussearch2082.codfw.wmnet, cirrussearch2079.codfw.wmnet, cirrussearch2085.codfw.wmnet, cirrussearch2069.codfw.wmnet, cirrussearch2068.codfw.wmnet, cirrussearch2080.codfw.wmnet, cirrussearch2072.codfw.wmnet, cirrussearch2066.codfw.wmnet, cirrussearch2096.codf [15:19:36] cirrussearch2064.codfw.wmnet, cirrussearch2094.codfw.wmnet, cirrussearch2078.codfw.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [15:19:38] (03PS35) 10Federico Ceratto: mysql: update replication source [cookbooks] - 10https://gerrit.wikimedia.org/r/1238368 (https://phabricator.wikimedia.org/T373436) [15:19:42] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [15:19:46] (03CR) 10JHathaway: memcached: reload_certs on certificate renewal (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1320940 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [15:19:51] !log fceratto@cumin1003 END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) [15:21:08] (03CR) 10Cathal Mooney: [C:03+2] eqsin: bump sr-linux verion to v26 after upgrade [homer/public] - 10https://gerrit.wikimedia.org/r/1320891 (https://phabricator.wikimedia.org/T418439) (owner: 10Cathal Mooney) [15:22:56] (03CR) 10CI reject: [V:04-1] mysql: update replication source [cookbooks] - 10https://gerrit.wikimedia.org/r/1238368 (https://phabricator.wikimedia.org/T373436) (owner: 10Federico Ceratto) [15:23:04] (03Merged) 10jenkins-bot: eqsin: bump sr-linux verion to v26 after upgrade [homer/public] - 10https://gerrit.wikimedia.org/r/1320891 (https://phabricator.wikimedia.org/T418439) (owner: 10Cathal Mooney) [15:23:56] ^^ again, sorry for the spam. CODFW is depooled so no user impact [15:24:27] (03CR) 10JHathaway: memcached: implement refresh_certs (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1320809 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [15:26:45] (03CR) 10AOkoth: [C:03+2] site: apply migration role to phab1005 [puppet] - 10https://gerrit.wikimedia.org/r/1320217 (https://phabricator.wikimedia.org/T377889) (owner: 10AOkoth) [15:26:46] 10ops-codfw, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 10netops: codfw: pod AB switches upgrade (2026) - https://phabricator.wikimedia.org/T426197#12183919 (10cmooney) [15:27:02] RECOVERY - Check unit status of push_cross_cluster_settings_9200 on cirrussearch2111 is OK: OK: Status of the systemd unit push_cross_cluster_settings_9200 https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [15:27:49] RESOLVED: HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs-next - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [15:29:17] jayme@cumin1003 renumber-node (PID 467549) is awaiting input [15:29:40] !log bking@cumin2003 END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - T324335 [15:29:43] T324335: Index the OpenSearch server application logs in the central observability cluster, using ECS format - https://phabricator.wikimedia.org/T324335 [15:32:02] PROBLEM - Check unit status of push_cross_cluster_settings_9600 on cirrussearch2108 is CRITICAL: CRITICAL: Status of the systemd unit push_cross_cluster_settings_9600 https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [15:32:37] 06SRE, 10LDAP-Access-Requests: Grant Access to ciadmin for Vaughn Walters - https://phabricator.wikimedia.org/T433615#12183950 (10Aklapper) Tyler and Dan are both out this week; CC'ing @Lferreira for potential approval not to have this stuck for too long ideally [15:32:42] RECOVERY - ElasticSearch health check for shards on 9643 on search.svc.codfw.wmnet is OK: OK - elasticsearch status production-search-psi-codfw: cluster_name: production-search-psi-codfw, status: green, timed_out: False, number_of_nodes: 28, number_of_data_nodes: 28, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1739, active_shards: 5204, relocating_shards: 0, initializing_shards: 0, unassigned_shards: [15:32:42] ed_unassigned_shards: 0, number_of_pending_tasks: 6733, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 979868, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [15:33:24] !log aaron@deploy1003 helmfile [eqiad] START helmfile.d/services/rest-gateway: apply [15:33:46] !log aaron@deploy1003 helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply [15:35:48] PROBLEM - ElasticSearch health check for shards on 9643 on search.svc.codfw.wmnet is CRITICAL: CRITICAL - elasticsearch https://search.svc.codfw.wmnet:9643/_cluster/health error while fetching: HTTPSConnectionPool(host=search.svc.codfw.wmnet, port=9643): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [15:38:20] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch2096 is OK: OK - elasticsearch status production-search-psi-codfw: cluster_name: production-search-psi-codfw, status: yellow, timed_out: False, number_of_nodes: 27, number_of_data_nodes: 27, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1739, active_shards: 4651, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 553, del [15:38:20] ssigned_shards: 186, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 89.37355880092237 https://wikitech.wikimedia.org/wiki/Search%23Administration [15:38:20] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch2095 is OK: OK - elasticsearch status production-search-psi-codfw: cluster_name: production-search-psi-codfw, status: yellow, timed_out: False, number_of_nodes: 27, number_of_data_nodes: 27, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1739, active_shards: 4651, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 553, del [15:38:20] ssigned_shards: 186, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 89.37355880092237 https://wikitech.wikimedia.org/wiki/Search%23Administration [15:38:20] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch2089 is OK: OK - elasticsearch status production-search-psi-codfw: cluster_name: production-search-psi-codfw, status: yellow, timed_out: False, number_of_nodes: 27, number_of_data_nodes: 27, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1739, active_shards: 4651, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 553, del [15:39:00] !log bking@cumin2003 START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - T324335 [15:39:04] T324335: Index the OpenSearch server application logs in the central observability cluster, using ECS format - https://phabricator.wikimedia.org/T324335 [15:40:20] jayme@cumin1003 renumber-node (PID 467549) is awaiting input [15:40:50] RECOVERY - PyBal backends health check on lvs2013 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [15:41:38] (03CR) 10BCornwall: [C:03+1] Geo-map: August update for Meta [dns] - 10https://gerrit.wikimedia.org/r/1320821 (owner: 10Slyngshede) [15:43:17] (03Abandoned) 10Dzahn: gerrit: add replica and spare servers to ssh host key aliases [puppet] - 10https://gerrit.wikimedia.org/r/1314095 (owner: 10Dzahn) [15:44:48] !log add php8.5 packages to component/php85 - T432983 [15:44:52] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [15:44:53] T432983: [PHP 8.5] Create Debian packages for PHP 8.5 - https://phabricator.wikimedia.org/T432983 [15:49:06] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [15:49:17] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [15:50:34] 10ops-drmrs, 06DC-Ops, 06Traffic: hw troubleshooting: Memory failure for cp6008.drmrs.wmnet - https://phabricator.wikimedia.org/T431651#12184021 (10RobH) Updates: Valerie shipped (3) replacement dimms, (1) for immediate use and (2) spares. https://www.dhl.com/us-en/home/tracking.html?tracking-id=15%204938%... [15:50:41] 10ops-drmrs, 06DC-Ops, 06Traffic: hw troubleshooting: Memory failure for cp6008.drmrs.wmnet - https://phabricator.wikimedia.org/T431651#12184022 (10RobH) a:03RobH [15:50:50] RECOVERY - PyBal backends health check on lvs2014 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [15:51:18] (03PS1) 10Dzahn: gerrit: drop RSA SSH host key [puppet] - 10https://gerrit.wikimedia.org/r/1320990 (https://phabricator.wikimedia.org/T240266) [15:52:02] RECOVERY - Check unit status of push_cross_cluster_settings_9600 on cirrussearch2108 is OK: OK: Status of the systemd unit push_cross_cluster_settings_9600 https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [15:52:35] 10ops-eqiad, 06SRE, 06DC-Ops: Unresponsive management for an-worker1147.mgmt:22 - https://phabricator.wikimedia.org/T433348#12184057 (10VRiley-WMF) @bking Hey, I wanted to reach out about this server. Is there a good time to reboot this? We're having some issues with the iDRAC and I'd like to try to power dr... [15:54:39] (03CR) 10Dzahn: [C:03+2] "this is not deleted from the private repo so far - so it can be added back by simply reverting this one" [puppet] - 10https://gerrit.wikimedia.org/r/1320990 (https://phabricator.wikimedia.org/T240266) (owner: 10Dzahn) [15:54:56] (03PS2) 10Dzahn: gerrit: drop RSA SSH host key [puppet] - 10https://gerrit.wikimedia.org/r/1320990 (https://phabricator.wikimedia.org/T240266) [15:55:33] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [15:55:35] (03PS6) 10AOkoth: phabricator: add multi-replica support [puppet] - 10https://gerrit.wikimedia.org/r/1306283 (https://phabricator.wikimedia.org/T377889) [15:55:42] !log fceratto@cumin1003 END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1) [15:55:51] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [15:56:02] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [15:56:54] PROBLEM - Check unit status of push_cross_cluster_settings_9600 on cirrussearch2076 is CRITICAL: CRITICAL: Status of the systemd unit push_cross_cluster_settings_9600 https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [15:57:12] 10ops-eqiad, 06DC-Ops, 06Data-Platform-SRE (2026-07-03 - 2026-07-31): Unresponsive management for an-worker1147.mgmt:22 - https://phabricator.wikimedia.org/T433348#12184083 (10bking) [15:57:43] 10ops-codfw, 06DC-Ops: Power Supply - PS Redundancy - issue on wikikube-worker2187:9290 - https://phabricator.wikimedia.org/T433994 (10phaultfinder) 03NEW [15:57:47] (03CR) 10Dzahn: [C:03+2] gerrit: drop RSA SSH host key [puppet] - 10https://gerrit.wikimedia.org/r/1320990 (https://phabricator.wikimedia.org/T240266) (owner: 10Dzahn) [15:58:40] (03PS7) 10AOkoth: phabricator: add multi-replica support [puppet] - 10https://gerrit.wikimedia.org/r/1306283 (https://phabricator.wikimedia.org/T377889) [15:59:06] !log robh@cumin1003 START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] [15:59:09] !log robh@cumin1003 END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] [15:59:41] !log robh@cumin1003 START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] [15:59:45] !log robh@cumin1003 END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet'] [16:00:05] jhathaway and rzl: Your horoscope predicts another Puppet request window deploy. May Zuul be (nice) with you. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260804T1600). [16:00:05] No Gerrit patches in the queue for this window AFAICS. [16:00:14] 10ops-eqiad, 06DC-Ops, 06Data-Platform-SRE (2026-07-03 - 2026-07-31): Unresponsive management for an-worker1147.mgmt:22 - https://phabricator.wikimedia.org/T433348#12184101 (10bking) @VRiley-WMF I don't interact with the `an-worker` servers much, so [[ https://wikimedia.slack.com/archives/C055QGPTC69/p178585... [16:00:18] RECOVERY - OpenSearch health check for shards on 9600 on cirrussearch2076 is OK: OK - elasticsearch status production-search-psi-codfw: cluster_name: production-search-psi-codfw, status: yellow, timed_out: False, number_of_nodes: 28, number_of_data_nodes: 28, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1739, active_shards: 4651, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 553, del [16:00:18] ssigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 89.37355880092237 https://wikitech.wikimedia.org/wiki/Search%23Administration [16:01:49] (03CR) 10BCornwall: [C:03+2] varnish: Split out media functions from upload [puppet] - 10https://gerrit.wikimedia.org/r/1320759 (https://phabricator.wikimedia.org/T427465) (owner: 10Slyngshede) [16:02:44] (03CR) 10Jgiannelos: [C:03+2] mobileapps: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320429 (owner: 10PipelineBot) [16:03:03] (03Restored) 10Jgiannelos: mobileapps: Pin staging to latest master [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320201 (owner: 10Jgiannelos) [16:03:24] (03PS2) 10Jgiannelos: Revert "mobileapps: Pin staging to latest master" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320201 [16:04:41] (03CR) 10Jgiannelos: [C:03+2] Revert "mobileapps: Pin staging to latest master" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320201 (owner: 10Jgiannelos) [16:04:58] !log gerrit2002/gerrit1003/gerrit2003 - rm /etc/gerrit/ssh_host_rsa_key [16:05:00] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [16:05:12] !log aokoth@cumin1003 START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet [16:05:21] (03Merged) 10jenkins-bot: mobileapps: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320429 (owner: 10PipelineBot) [16:06:54] RECOVERY - Check unit status of push_cross_cluster_settings_9600 on cirrussearch2076 is OK: OK: Status of the systemd unit push_cross_cluster_settings_9600 https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [16:06:55] (03Merged) 10jenkins-bot: Revert "mobileapps: Pin staging to latest master" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320201 (owner: 10Jgiannelos) [16:07:21] !log jgiannelos@deploy1003 helmfile [staging] START helmfile.d/services/mobileapps: apply [16:07:28] !log aokoth@cumin1003 END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet [16:07:30] !log jgiannelos@deploy1003 helmfile [staging] DONE helmfile.d/services/mobileapps: apply [16:07:33] !log jgiannelos@deploy1003 helmfile [staging] START helmfile.d/services/mobileapps: apply [16:07:37] !log jgiannelos@deploy1003 helmfile [staging] DONE helmfile.d/services/mobileapps: apply [16:07:38] jouncebot: nowandnext [16:07:39] For the next 0 hour(s) and 52 minute(s): Puppet request window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260804T1600) [16:07:39] In 0 hour(s) and 52 minute(s): MediaWiki infrastructure (UTC late) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260804T1700) [16:07:50] !log jgiannelos@deploy1003 helmfile [staging] START helmfile.d/services/mobileapps: apply [16:07:54] (03PS1) 10Ladsgroup: thumbor: Use ImageMagick's -background=none flag for Alpha WebP images [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320998 (https://phabricator.wikimedia.org/T283646) [16:07:56] !log jgiannelos@deploy1003 helmfile [staging] DONE helmfile.d/services/mobileapps: apply [16:08:01] mutante: nothing in the puppet window, all yours [16:08:06] i see the puppet window is empty. [16:08:10] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:08:17] rzl: thanks!ok, sneaking in gerrit service restart [16:08:17] !log jgiannelos@deploy1003 helmfile [staging] START helmfile.d/services/mobileapps: apply [16:08:21] !log jgiannelos@deploy1003 helmfile [staging] DONE helmfile.d/services/mobileapps: apply [16:08:39] !log jgiannelos@deploy1003 helmfile [eqiad] START helmfile.d/services/mobileapps: apply [16:09:22] !log jgiannelos@deploy1003 helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply [16:09:26] !log jgiannelos@deploy1003 helmfile [codfw] START helmfile.d/services/mobileapps: apply [16:09:54] (03PS1) 10Scott French: php8.3: Rebuild to pick up new PHP packages (8.3.33) [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1320997 [16:10:09] !log jgiannelos@deploy1003 helmfile [codfw] DONE helmfile.d/services/mobileapps: apply [16:10:40] (03CR) 10BCornwall: [C:03+1] php8.3: Rebuild to pick up new PHP packages (8.3.33) [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1320997 (owner: 10Scott French) [16:10:45] !log jgiannelos@deploy1003 helmfile [eqiad] START helmfile.d/services/mobileapps: apply [16:10:56] well, or not, since Arnaudb is debugging something elese [16:11:28] !log jgiannelos@deploy1003 helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply [16:11:44] (03PS6) 10Effie Mouzeli: api-gateway:: remove API Gateway specific ratelimiting logic #3 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319563 (https://phabricator.wikimedia.org/T428625) (owner: 10Aaron Schulz) [16:12:22] (03CR) 10Kamila Součková: [C:03+1] php8.3: Rebuild to pick up new PHP packages (8.3.33) [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1320997 (owner: 10Scott French) [16:12:59] (03CR) 10AOkoth: "https://puppet-compiler.wmflabs.org/output/1306283/9112/" [puppet] - 10https://gerrit.wikimedia.org/r/1306283 (https://phabricator.wikimedia.org/T377889) (owner: 10AOkoth) [16:13:57] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:14:19] mutante: I might sneak in a changeprop update in that case, if that won't conflict [16:16:33] (03PS8) 10Effie Mouzeli: memcached: implement refresh_certs [puppet] - 10https://gerrit.wikimedia.org/r/1320809 (https://phabricator.wikimedia.org/T353511) [16:16:59] (03CR) 10Effie Mouzeli: memcached: implement refresh_certs (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1320809 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [16:17:05] rzl: go for it [16:17:14] right now I am waiting for a bit [16:17:19] (03PS9) 10Effie Mouzeli: memcached: implement refresh_certs [puppet] - 10https://gerrit.wikimedia.org/r/1320809 (https://phabricator.wikimedia.org/T353511) [16:17:31] cool thanks, coordinating [16:17:41] (03PS10) 10Effie Mouzeli: memcached: implement refresh_certs [puppet] - 10https://gerrit.wikimedia.org/r/1320809 (https://phabricator.wikimedia.org/T353511) [16:17:41] (if you get there before me, go ahead) [16:17:42] !log reprepro include php8.3_8.3.33-1+wmf12u1 into component/php83 for bookworm-wikimedia [16:17:44] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [16:17:50] PROBLEM - PyBal backends health check on lvs2014 is CRITICAL: PYBAL CRITICAL - CRITICAL - search-omega-https_9443: Servers cirrussearch2088.codfw.wmnet, cirrussearch2111.codfw.wmnet, cirrussearch2100.codfw.wmnet, cirrussearch2073.codfw.wmnet, cirrussearch2097.codfw.wmnet, cirrussearch2092.codfw.wmnet, cirrussearch2063.codfw.wmnet, cirrussearch2098.codfw.wmnet, cirrussearch2112.codfw.wmnet, cirrussearch2104.codfw.wmnet, cirrussearch2099.co [16:17:50] t, cirrussearch2102.codfw.wmnet, cirrussearch2093.codfw.wmnet, cirrussearch2087.codfw.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [16:17:50] PROBLEM - PyBal backends health check on lvs2013 is CRITICAL: PYBAL CRITICAL - CRITICAL - search-omega-https_9443: Servers cirrussearch2088.codfw.wmnet, cirrussearch2074.codfw.wmnet, cirrussearch2061.codfw.wmnet, cirrussearch2100.codfw.wmnet, cirrussearch2073.codfw.wmnet, cirrussearch2092.codfw.wmnet, cirrussearch2070.codfw.wmnet, cirrussearch2098.codfw.wmnet, cirrussearch2112.codfw.wmnet, cirrussearch2104.codfw.wmnet, cirrussearch2077.co [16:17:50] t, cirrussearch2111.codfw.wmnet, cirrussearch2093.codfw.wmnet, cirrussearch2087.codfw.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [16:17:55] !log reprepro include php8.3_8.3.33-1+wmf11u1 into component/php83 for bullseye-wikimedia [16:17:58] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [16:18:02] PROBLEM - ElasticSearch health check for shards on 9443 on search.svc.codfw.wmnet is CRITICAL: CRITICAL - elasticsearch https://search.svc.codfw.wmnet:9443/_cluster/health error while fetching: HTTPSConnectionPool(host=search.svc.codfw.wmnet, port=9443): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:18:09] (03CR) 10Effie Mouzeli: memcached: implement refresh_certs (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1320809 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [16:18:26] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2105 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:18:26] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2090 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:18:26] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2106 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:18:26] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2104 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:18:26] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2097 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:18:27] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2102 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:18:27] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2099 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:18:28] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2100 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:18:28] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2070 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:18:28] (03CR) 10Clément Goubert: [C:03+1] "Prod no-op, ship it." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319563 (https://phabricator.wikimedia.org/T428625) (owner: 10Aaron Schulz) [16:18:29] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2098 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:18:29] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2061 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:18:30] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2081 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:18:30] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2063 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:18:31] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2088 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:18:31] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2073 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:18:32] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2093 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:18:32] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2092 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:18:33] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2077 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:18:33] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2084 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:18:34] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2067 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:18:34] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2074 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:18:35] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2065 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:18:35] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2087 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:18:36] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2114 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:18:36] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2112 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:18:37] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2111 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [16:19:21] (03PS1) 10Cathal Mooney: Codfw: enable OSPF on new Arelion transport sub-ints [homer/public] - 10https://gerrit.wikimedia.org/r/1321000 (https://phabricator.wikimedia.org/T424839) [16:20:32] (03CR) 10RLazarus: [C:03+2] Revert^2 "Add CacheAbstractContentFragment to high_traffic_jobs with concurrency=1" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320238 (https://phabricator.wikimedia.org/T430898) (owner: 10Jforrester) [16:20:36] (03CR) 10Clément Goubert: [C:03+2] deployment_server: Install make for test suite [puppet] - 10https://gerrit.wikimedia.org/r/1320119 (owner: 10Clément Goubert) [16:20:56] mutante: okay, going ahead actually :) but no conflict with a gerrit restart, as long as that submit goes through first [16:21:07] rzl: go !:) [16:21:34] ayy what's up with cirrussearch? [16:21:50] !log ladsgroup@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance [16:21:55] inflatador: ^ I believe it's depooled in codfw but not downtimed [16:22:03] (03PS15) 10BCornwall: varnish: Configure text for thumbnails [puppet] - 10https://gerrit.wikimedia.org/r/1306896 (https://phabricator.wikimedia.org/T427465) (owner: 10Slyngshede) [16:22:37] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Depooling db1179 (T433990)', diff saved to https://phabricator.wikimedia.org/P95865 and previous config saved to /var/cache/conftool/dbconfig/20260804-162236-ladsgroup.json [16:22:41] T433990: Optimize echo tables in x1 - https://phabricator.wikimedia.org/T433990 [16:22:46] (03Merged) 10jenkins-bot: Revert^2 "Add CacheAbstractContentFragment to high_traffic_jobs with concurrency=1" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320238 (https://phabricator.wikimedia.org/T430898) (owner: 10Jforrester) [16:23:18] RECOVERY - OpenSearch health check for shards on 9400 on cirrussearch2097 is OK: OK - elasticsearch status production-search-omega-codfw: cluster_name: production-search-omega-codfw, status: yellow, timed_out: False, number_of_nodes: 25, number_of_data_nodes: 25, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 4810, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 385, [16:23:18] _unassigned_shards: 385, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 92.58902791145333 https://wikitech.wikimedia.org/wiki/Search%23Administration [16:23:18] RECOVERY - OpenSearch health check for shards on 9400 on cirrussearch2104 is OK: OK - elasticsearch status production-search-omega-codfw: cluster_name: production-search-omega-codfw, status: yellow, timed_out: False, number_of_nodes: 25, number_of_data_nodes: 25, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 4810, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 385, [16:23:18] _unassigned_shards: 385, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 92.58902791145333 https://wikitech.wikimedia.org/wiki/Search%23Administration [16:23:18] RECOVERY - OpenSearch health check for shards on 9400 on cirrussearch2093 is OK: OK - elasticsearch status production-search-omega-codfw: cluster_name: production-search-omega-codfw, status: yellow, timed_out: False, number_of_nodes: 25, number_of_data_nodes: 25, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 4810, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 385, [16:23:27] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-private-users for aramirezwmf - https://phabricator.wikimedia.org/T433999 (10ARamirez_WMF) 03NEW [16:23:33] (03PS2) 10Cathal Mooney: Codfw: enable OSPF on new Arelion transport sub-ints [homer/public] - 10https://gerrit.wikimedia.org/r/1321000 (https://phabricator.wikimedia.org/T424839) [16:23:49] !log rzl@deploy1003 helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply [16:23:50] RECOVERY - PyBal backends health check on lvs2014 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [16:23:50] RECOVERY - PyBal backends health check on lvs2013 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [16:23:56] RECOVERY - ElasticSearch health check for shards on 9443 on search.svc.codfw.wmnet is OK: OK - elasticsearch status production-search-omega-codfw: cluster_name: production-search-omega-codfw, status: yellow, timed_out: False, number_of_nodes: 25, number_of_data_nodes: 25, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 4810, relocating_shards: 0, initializing_shards: 0, unassigned_sha [16:23:56] , delayed_unassigned_shards: 385, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 92.58902791145333 https://wikitech.wikimedia.org/wiki/Search%23Administration [16:23:57] rzl: i got the ok too to go ahead [16:24:00] !log rzl@deploy1003 helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply [16:24:12] mutante: go for it, I'm helmfile applying but through with gerrit unless I need to roll back [16:24:25] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Repooling after maintenance db1179 (T433990)', diff saved to https://phabricator.wikimedia.org/P95866 and previous config saved to /var/cache/conftool/dbconfig/20260804-162424-ladsgroup.json [16:24:26] (03CR) 10Scott French: [V:03+2] "Built locally:" [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1320997 (owner: 10Scott French) [16:24:28] ack, thanks, using cookbook for it [16:25:23] !log dzahn@cumin1003 START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2003 [16:25:39] (03CR) 10Cathal Mooney: [C:03+2] Codfw: enable OSPF on new Arelion transport sub-ints [homer/public] - 10https://gerrit.wikimedia.org/r/1321000 (https://phabricator.wikimedia.org/T424839) (owner: 10Cathal Mooney) [16:25:54] !log restarting gerrit - dropped outdated RSA host key [16:25:56] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [16:26:36] staging looks fine (as last time), moving on to codfw [16:26:38] !log bking@cumin2003 END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - T324335 [16:26:42] T324335: Index the OpenSearch server application logs in the central observability cluster, using ECS format - https://phabricator.wikimedia.org/T324335 [16:26:51] !log ladsgroup@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance [16:26:54] PROBLEM - Check unit status of push_cross_cluster_settings_9400 on cirrussearch2073 is CRITICAL: CRITICAL: Status of the systemd unit push_cross_cluster_settings_9400 https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [16:27:19] !log rzl@deploy1003 helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply [16:27:38] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Depooling db1179 (T433990)', diff saved to https://phabricator.wikimedia.org/P95867 and previous config saved to /var/cache/conftool/dbconfig/20260804-162736-ladsgroup.json [16:27:41] T433990: Optimize echo tables in x1 - https://phabricator.wikimedia.org/T433990 [16:27:57] !log rzl@deploy1003 helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply [16:27:58] !log dzahn@cumin1003 END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2003 [16:28:40] !log bking@cumin2003 conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad [16:28:58] !log bking@cumin2003 conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw [16:28:58] !log bking@cumin2003 conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw [16:28:59] !log bking@cumin2003 conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw [16:29:14] (03CR) 10Jforrester: "CI equivalent: https://gerrit.wikimedia.org/r/c/integration/config/+/1321008" [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1320997 (owner: 10Scott French) [16:29:28] James_F: LGTY so far? [16:29:35] rzl: Yes, all quiet. [16:30:28] the gerrit-ssh RSA host key is now gone. hopefully we have no ancient clients older than like 12 years or so. if anyone complains about gerrit-ssh, lmk. [16:30:36] closing out ticket from around 2019 [16:30:38] <3 [16:32:24] deploying in eqiad [16:32:28] !log rzl@deploy1003 helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply [16:33:10] can't wait to be reminded about some gerrit-pushing bot i run but have completely forgotten about that's about to break [16:33:16] !log rzl@deploy1003 helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply [16:33:42] (i did update the keys on libup and checked that the cloud/instance-puppet pusher script doesn't need to have it manually updated) [16:35:06] we did serve both old and new key in parallel for a bit - as it is called industry standard. in hindsight though.. does it really change anything? old clients would negotiate down and just notice it a week or 2 later. [16:35:59] James_F: looks like that Just Worked this time, any concerns? [16:35:59] mutante: It lets you see if the new key worked or just flaked in new clients, though. [16:36:12] rzl: Nope, all good from my end. [16:36:19] sweet [16:36:57] taavi: btw, I did see your comment about old apache logs on gerrit in old location. very valid. will clean that up and just move it to some "old" dir first. [16:37:17] awesome, thanks [16:37:21] James_F: true, fair enough [16:39:19] 06SRE, 06collaboration-services, 10Gerrit, 06Release-Engineering-Team (Seen), 07Security: gerrit.wikimedia.org:29418 uses 1024-bit RSA key only - https://phabricator.wikimedia.org/T240266#12184308 (10Dzahn) >>! In T240266#12152251, @Bawolff wrote: > The announcement email kind of directs users here.... [16:41:12] 06SRE, 10LDAP-Access-Requests: Grant Access to ciadmin for Vaughn Walters - https://phabricator.wikimedia.org/T433615#12184313 (10Dzahn) a:03Lferreira Levi, if you don't mind.. and we can just close this out. [16:43:31] 06SRE, 10SRE-Access-Requests, 06Release-Engineering-Team: Requesting access to Jenkins server for PWangai-WMF - https://phabricator.wikimedia.org/T433563#12184321 (10Dzahn) @pwangai Cool! in that case I think the simplest path forward is we just add people to "contint-users" which gives unprivileged shell ac... [16:44:35] (03CR) 10JHathaway: [C:03+1] memcached: implement refresh_certs [puppet] - 10https://gerrit.wikimedia.org/r/1320809 (https://phabricator.wikimedia.org/T353511) (owner: 10Effie Mouzeli) [16:45:40] 10ops-codfw, 06SRE, 06DC-Ops: Power Supply - PS Redundancy - issue on wikikube-worker2187:9290 - https://phabricator.wikimedia.org/T433994#12184331 (10Jhancock.wm) @Clement_Goubert I got the power issue resolved on this one but the server is about to have a hard time. This error popped up after getting the... [16:45:42] (03PS4) 10Effie Mouzeli: memcached: reload_certs on certificate renewal [puppet] - 10https://gerrit.wikimedia.org/r/1320940 (https://phabricator.wikimedia.org/T353511) [16:46:00] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 6/8 UP : OSPFv3: 6/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [16:48:30] !log gerrit2003 - moving old apache logfiles older than 60 days from /var/log/apache2 to /srv/gerrit/site_path/review_site/logs/old/ [16:48:32] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [16:48:49] FIRING: HelmReleaseBadStatus: Helm release wdqs/scholarly-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [16:49:02] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [16:50:18] RECOVERY - OpenSearch health check for shards on 9400 on cirrussearch2073 is OK: OK - elasticsearch status production-search-omega-codfw: cluster_name: production-search-omega-codfw, status: green, timed_out: False, number_of_nodes: 27, number_of_data_nodes: 27, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 5195, relocating_shards: 2, initializing_shards: 0, unassigned_shards: 0, de [16:50:18] assigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [16:50:41] (03PS16) 10BCornwall: varnish: Configure text for thumbnails [puppet] - 10https://gerrit.wikimedia.org/r/1306896 (https://phabricator.wikimedia.org/T427465) (owner: 10Slyngshede) [16:52:19] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2187.codfw.wmnet [16:52:51] !log gerrit2003:/var/log/apache2# ln -s /srv/gerrit/site_path/review_site/logs/ gerrit [16:52:52] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2187.codfw.wmnet [16:52:53] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [16:52:58] 10ops-codfw, 06SRE, 06DC-Ops: Power Supply - PS Redundancy - issue on wikikube-worker2187:9290 - https://phabricator.wikimedia.org/T433994#12184407 (10ops-monitoring-bot) Cookbook cookbooks.sre.k8s.pool-depool-node started by cgoubert@cumin2003 depool for host wikikube-worker2187.codfw.wmnet completed: - wik... [16:54:42] !log cgoubert@cumin2003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on wikikube-worker2187.codfw.wmnet with reason: Hardware issue [16:54:53] 10ops-codfw, 06DC-Ops, 06ServiceOps: Power Supply - PS Redundancy - issue on wikikube-worker2187:9290 - https://phabricator.wikimedia.org/T433994#12184410 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=67936647-b793-4d60-b367-b9ecddbf5fe7) set by cgoubert@cumin2003 for 1 day, 0:00:00 on... [16:54:54] 10ops-codfw, 06DC-Ops, 06ServiceOps: Power Supply - PS Redundancy - issue on wikikube-worker2187:9290 - https://phabricator.wikimedia.org/T433994#12184421 (10Clement_Goubert) [16:55:16] 10ops-codfw, 06DC-Ops, 06ServiceOps, 10ServiceOps-Upgrades-Hardware: Power Supply - PS Redundancy - issue on wikikube-worker2187:9290 - https://phabricator.wikimedia.org/T433994#12184426 (10Clement_Goubert) [16:55:28] 10ops-codfw, 06DC-Ops, 06ServiceOps, 10ServiceOps-Upgrades-Hardware: Power Supply - PS Redundancy - issue on wikikube-worker2187:9290 - https://phabricator.wikimedia.org/T433994#12184432 (10Clement_Goubert) Depooled, downtimed, go ahead. [16:57:02] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-private-users for aramirezwmf - https://phabricator.wikimedia.org/T433999#12184439 (10MaCollins-WMF) Hi - confirming we need Aida to have access to the FY26-27 Metrics Superset dashboard. [16:58:49] RESOLVED: HelmReleaseBadStatus: Helm release wdqs/scholarly-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [17:00:04] swfrench-wmf: Your horoscope predicts another MediaWiki infrastructure (UTC late) deploy. May Zuul be (nice) with you. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260804T1700). [17:00:12] o [17:00:14] o/ [17:00:24] I'll be getting started on the infra window shortly [17:00:34] !log cwilliams@cumin1003 START - Cookbook sre.mysql.update-replication [17:00:46] !log cwilliams@cumin1003 END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97) [17:00:53] rzl: did you have anything else you needed to get out related to the changeprop changes? [17:01:45] (03CR) 10Scott French: [V:03+2] "Great, thank you!" [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1320997 (owner: 10Scott French) [17:01:46] (03CR) 10Scott French: [V:03+2 C:03+2] php8.3: Rebuild to pick up new PHP packages (8.3.33) [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1320997 (owner: 10Scott French) [17:03:24] swfrench-wmf: nope thanks, I'm hands off [17:03:35] awesome, thanks [17:05:35] !log swfrench@deploy1003 Started scap sync-world: Pick up new PHP production image [17:06:54] RECOVERY - Check unit status of push_cross_cluster_settings_9400 on cirrussearch2073 is OK: OK: Status of the systemd unit push_cross_cluster_settings_9400 https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [17:07:29] this deployment should take ~ 35m or so, after which I'll have a second helmfile-only change to deploy [17:10:48] (03PS3) 10Mahveotm: decorators: Add UTC timestamp to retry warnings [software/pywmflib] - 10https://gerrit.wikimedia.org/r/1320151 (https://phabricator.wikimedia.org/T433698) [17:16:29] (03CR) 10CWilliams: mysql: update replication source (031 comment) [cookbooks] - 10https://gerrit.wikimedia.org/r/1238368 (https://phabricator.wikimedia.org/T373436) (owner: 10Federico Ceratto) [17:22:38] FIRING: [2x] CertAlmostExpired: gNMI TLS certificate for lsw1-d6-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [17:26:18] 06SRE, 06collaboration-services, 10Gerrit, 06Release-Engineering-Team (Seen), 07Security: gerrit.wikimedia.org:29418 uses 1024-bit RSA key only - https://phabricator.wikimedia.org/T240266#12184535 (10dancy) @dzahn Congratulations! [17:27:23] FIRING: [2x] CertAlmostExpired: gNMI TLS certificate for lsw1-d6-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [17:28:06] !log aokoth@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on phab1005.eqiad.wmnet with reason: Puppet Failure [17:31:36] (03CR) 10Dzahn: [C:03+1] phabricator: add multi-replica support [puppet] - 10https://gerrit.wikimedia.org/r/1306283 (https://phabricator.wikimedia.org/T377889) (owner: 10AOkoth) [17:32:12] (03PS17) 10BCornwall: varnish: Configure text for thumbnails [puppet] - 10https://gerrit.wikimedia.org/r/1306896 (https://phabricator.wikimedia.org/T427465) (owner: 10Slyngshede) [17:32:23] RESOLVED: [2x] CertAlmostExpired: gNMI TLS certificate for lsw1-d6-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [17:33:46] !log swfrench@deploy1003 Finished scap sync-world: Pick up new PHP production image (duration: 28m 32s) [17:34:00] wow, fast! [17:34:00] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Repooling after maintenance db1179 (T433990)', diff saved to https://phabricator.wikimedia.org/P95868 and previous config saved to /var/cache/conftool/dbconfig/20260804-173359-ladsgroup.json [17:34:05] T433990: Optimize echo tables in x1 - https://phabricator.wikimedia.org/T433990 [17:34:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.22% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [17:35:20] ^ "not unexpected" in the sense that mw-web appears to be intermittently running warm over the last 12h [17:35:29] I'll let that settle for a few minutes, and then queue up my next change [17:38:16] (03CR) 10Scott French: [C:03+2] mw-*: Revert temporary mw.mail_timeout increase [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319157 (https://phabricator.wikimedia.org/T383047) (owner: 10Scott French) [17:39:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 25% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [17:42:39] (03Merged) 10jenkins-bot: mw-*: Revert temporary mw.mail_timeout increase [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319157 (https://phabricator.wikimedia.org/T383047) (owner: 10Scott French) [17:44:46] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95869 and previous config saved to /var/cache/conftool/dbconfig/20260804-174445-ladsgroup.json [17:45:02] alright, next deployment starting shortly [17:45:55] !log swfrench@deploy1003 Started scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - T383047 [17:45:59] T383047: Could not send confirmation email: Unknown error in PHP's mail() function. - https://phabricator.wikimedia.org/T383047 [17:46:56] !log swfrench@deploy1003 swfrench: Deploy helmfile-only msmtp timeout override cleanup - T383047 synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [17:47:13] (03CR) 10BCornwall: [V:04-1] "Done" [puppet] - 10https://gerrit.wikimedia.org/r/1306896 (https://phabricator.wikimedia.org/T427465) (owner: 10Slyngshede) [17:47:13] (03PS8) 10AOkoth: phabricator: add multi-replica support [puppet] - 10https://gerrit.wikimedia.org/r/1306283 (https://phabricator.wikimedia.org/T377889) [17:48:13] !log swfrench@deploy1003 swfrench: Continuing with deployment [17:50:00] !log swfrench@deploy1003 Finished scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - T383047 (duration: 04m 05s) [17:52:44] (03CR) 10AOkoth: [C:03+2] phabricator: add multi-replica support [puppet] - 10https://gerrit.wikimedia.org/r/1306283 (https://phabricator.wikimedia.org/T377889) (owner: 10AOkoth) [17:52:47] that should be everything I've got planned for this window. I'll keep an eye on things for a bit, but don't plan to touch anything further. [17:55:21] !log jayme@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1154.eqiad.wmnet [17:55:22] !log jayme@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1154.eqiad.wmnet [17:55:25] !log jayme@cumin1003 END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1154.eqiad.wmnet [17:55:33] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95870 and previous config saved to /var/cache/conftool/dbconfig/20260804-175531-ladsgroup.json [18:00:05] jnuche and jeena: May I have your attention please! MediaWiki train - Utc-0+Utc-7 Version (secondary timeslot). (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260804T1800) [18:02:30] (03CR) 10Ladsgroup: [C:03+2] thumbor: Use ImageMagick's -background=none flag for Alpha WebP images [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320998 (https://phabricator.wikimedia.org/T283646) (owner: 10Ladsgroup) [18:05:09] (03Merged) 10jenkins-bot: thumbor: Use ImageMagick's -background=none flag for Alpha WebP images [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320998 (https://phabricator.wikimedia.org/T283646) (owner: 10Ladsgroup) [18:06:19] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Repooling after maintenance db1179 (T433990)', diff saved to https://phabricator.wikimedia.org/P95871 and previous config saved to /var/cache/conftool/dbconfig/20260804-180618-ladsgroup.json [18:06:23] T433990: Optimize echo tables in x1 - https://phabricator.wikimedia.org/T433990 [18:06:35] !log ladsgroup@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1203.eqiad.wmnet with reason: Maintenance [18:06:59] (03Abandoned) 10Andrew Bogott: cloudnfs: replace wikiqlever with project uuid [puppet] - 10https://gerrit.wikimedia.org/r/1320181 (owner: 10Andrew Bogott) [18:07:02] !log ladsgroup@deploy1003 helmfile [staging] START helmfile.d/services/thumbor: apply [18:07:22] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Depooling db1203 (T433990)', diff saved to https://phabricator.wikimedia.org/P95872 and previous config saved to /var/cache/conftool/dbconfig/20260804-180721-ladsgroup.json [18:08:12] !log ladsgroup@deploy1003 helmfile [staging] DONE helmfile.d/services/thumbor: apply [18:10:42] !log ladsgroup@deploy1003 helmfile [codfw] START helmfile.d/services/thumbor: apply [18:13:40] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. [18:14:25] !log ladsgroup@deploy1003 helmfile [codfw] DONE helmfile.d/services/thumbor: apply [18:14:25] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. [18:17:50] (03CR) 10Aaron Schulz: [C:03+1] rest: Add test server option to REST Sandbox for Wikipedia projects [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1313827 (https://phabricator.wikimedia.org/T408816) (owner: 10KineticPelagic) [18:18:26] !log ladsgroup@deploy1003 helmfile [eqiad] START helmfile.d/services/thumbor: apply [18:20:45] !log ladsgroup@deploy1003 helmfile [eqiad] DONE helmfile.d/services/thumbor: apply [18:22:06] (03PS3) 10Tsevener: Add new Apple app site association file for Test Wiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1315128 (https://phabricator.wikimedia.org/T432412) [18:28:39] (03CR) 10CDobbins: [C:03+1] varnish: Add SPDX license header to VTC files [puppet] - 10https://gerrit.wikimedia.org/r/1313259 (owner: 10BCornwall) [18:29:05] (03CR) 10BCornwall: [V:03+2 C:03+2] varnish: Add SPDX license header to VTC files [puppet] - 10https://gerrit.wikimedia.org/r/1313259 (owner: 10BCornwall) [18:34:29] (03PS1) 10AOkoth: hieradata: reenable scap bootstrapping on phab1005 [puppet] - 10https://gerrit.wikimedia.org/r/1321048 (https://phabricator.wikimedia.org/T434003) [18:37:23] (03CR) 10Tsevener: "@krinkle@fastmail.com" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1315128 (https://phabricator.wikimedia.org/T432412) (owner: 10Tsevener) [18:40:43] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, August 04 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-it" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1313827 (https://phabricator.wikimedia.org/T408816) (owner: 10KineticPelagic) [18:41:39] (03CR) 10Bking: [C:03+1] "Although PCC shows a diff, I confirmed that the default setting of false will not materially change the `cephosd` cluster values via `root" [puppet] - 10https://gerrit.wikimedia.org/r/1318163 (https://phabricator.wikimedia.org/T429387) (owner: 10Andrew Bogott) [18:43:53] (03CR) 10AOkoth: [C:03+2] hieradata: reenable scap bootstrapping on phab1005 [puppet] - 10https://gerrit.wikimedia.org/r/1321048 (https://phabricator.wikimedia.org/T434003) (owner: 10AOkoth) [18:43:57] (03CR) 10Krinkle: [C:03+1] Add configurable RestTermsOfServiceUrl [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320163 (https://phabricator.wikimedia.org/T428147) (owner: 10Milazg) [18:46:04] (03CR) 10Bking: [C:03+1] "As I did with the previous patch in the chain, I've confirmed that the defaults set by this change are no different than the defaults curr" [puppet] - 10https://gerrit.wikimedia.org/r/1313966 (https://phabricator.wikimedia.org/T429387) (owner: 10Andrew Bogott) [18:57:47] (03PS1) 10Andrew Bogott: wmcs-backup: don't error out on a few failed snapshot removals [puppet] - 10https://gerrit.wikimedia.org/r/1321058 [19:02:12] !log gerrit ssh -p 29418 gerrit.wikimedia.org gerrit index changes 1320979 [19:02:15] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [19:02:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 17.77% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [19:03:37] (03CR) 10Andrew Bogott: [C:03+2] P:openstack: designate: Remove absented mcrouter resources [puppet] - 10https://gerrit.wikimedia.org/r/1307344 (https://phabricator.wikimedia.org/T427189) (owner: 10Majavah) [19:06:46] (03CR) 10Andrew Bogott: [C:03+2] Ceph: support setting bdev_enable_discard [puppet] - 10https://gerrit.wikimedia.org/r/1318163 (https://phabricator.wikimedia.org/T429387) (owner: 10Andrew Bogott) [19:07:10] (03CR) 10Andrew Bogott: [C:03+2] ceph.conf: support adjusting slow ops health messages [puppet] - 10https://gerrit.wikimedia.org/r/1313966 (https://phabricator.wikimedia.org/T429387) (owner: 10Andrew Bogott) [19:08:16] (03PS1) 10Ahmon Dancy: Bump buildkit image to wmf-v0.32.2 [puppet] - 10https://gerrit.wikimedia.org/r/1321081 (https://phabricator.wikimedia.org/T433985) [19:16:44] PROBLEM - Ensure acme-chief-backend is running only in the active node on acmechief2002 is CRITICAL: PROCS CRITICAL: 2 processes with args acme-chief-backend https://wikitech.wikimedia.org/wiki/Acme-chief [19:17:44] RECOVERY - Ensure acme-chief-backend is running only in the active node on acmechief2002 is OK: PROCS OK: 1 process with args acme-chief-backend https://wikitech.wikimedia.org/wiki/Acme-chief [19:23:48] 06SRE, 06ServiceOps, 10VisualEditor, 10VisualEditor Suggestion Mode, and 3 others: Deploy Headless VE in k8s for technical pilot - https://phabricator.wikimedia.org/T431497#12185046 (10Ottomata) →14Duplicate dup:03T432715 [19:27:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.24% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [19:27:39] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Repooling after maintenance db1203 (T433990)', diff saved to https://phabricator.wikimedia.org/P95874 and previous config saved to /var/cache/conftool/dbconfig/20260804-192738-ladsgroup.json [19:27:43] T433990: Optimize echo tables in x1 - https://phabricator.wikimedia.org/T433990 [19:32:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.41% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [19:35:50] (03CR) 10Aghirelli: [C:03+1] Add configurable RestTermsOfServiceUrl [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1320163 (https://phabricator.wikimedia.org/T428147) (owner: 10Milazg) [19:37:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.41% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [19:38:26] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95875 and previous config saved to /var/cache/conftool/dbconfig/20260804-193825-ladsgroup.json [19:45:06] PROBLEM - Ensure traffic_manager is running for instance backend on cp2058 is CRITICAL: PROCS CRITICAL: 3 processes with args /usr/bin/traffic_manager --nosyslog https://wikitech.wikimedia.org/wiki/Apache_Traffic_Server [19:45:18] !log jhancock@cumin2002 START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS bookworm [19:45:48] !log jhancock@cumin2002 START - Cookbook sre.hosts.move-vlan for host mc2046 [19:46:06] RECOVERY - Ensure traffic_manager is running for instance backend on cp2058 is OK: PROCS OK: 1 process with args /usr/bin/traffic_manager --nosyslog https://wikitech.wikimedia.org/wiki/Apache_Traffic_Server [19:46:09] !log jhancock@cumin2002 START - Cookbook sre.dns.netbox [19:49:12] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95876 and previous config saved to /var/cache/conftool/dbconfig/20260804-194911-ladsgroup.json [19:50:28] !log jhancock@cumin2002 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" [19:50:33] !log jhancock@cumin2002 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002" [19:50:33] !log jhancock@cumin2002 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [19:50:33] !log jhancock@cumin2002 START - Cookbook sre.dns.wipe-cache mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [19:50:37] !log jhancock@cumin2002 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [19:50:38] !log jhancock@cumin2002 START - Cookbook sre.network.configure-switch-interfaces for host mc2046 [19:50:58] !log jhancock@cumin2002 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046 [19:50:59] !log jhancock@cumin2002 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046 [19:57:10] (03PS1) 10Ahmon Dancy: Add /w/deployment-info.php entrypoint [mediawiki-config] (train-dev) - 10https://gerrit.wikimedia.org/r/1321096 [19:57:53] (03Abandoned) 10Ahmon Dancy: wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE [mediawiki-config] (train-dev) - 10https://gerrit.wikimedia.org/r/1316090 (https://phabricator.wikimedia.org/T433104) (owner: 10Ahmon Dancy) [19:59:58] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Repooling after maintenance db1203 (T433990)', diff saved to https://phabricator.wikimedia.org/P95877 and previous config saved to /var/cache/conftool/dbconfig/20260804-195957-ladsgroup.json [20:00:03] T433990: Optimize echo tables in x1 - https://phabricator.wikimedia.org/T433990 [20:00:05] RoanKattouw, urbanecm, TheresNoTime, kindrobot, and cjming: Time to snap out of that daydream and deploy UTC late backport window. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260804T2000). [20:00:05] hyang: A patch you scheduled for UTC late backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [20:00:11] aloha, i am here [20:00:15] !log ladsgroup@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1216.eqiad.wmnet with reason: Maintenance [20:06:03] would anyone be able to help me deploy my patch? [20:07:25] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [20:09:06] !log jhancock@cumin2002 START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage [20:09:27] TheresNoTime if you are available, perhaps? [20:12:52] hyang: hi, yes one moment :) [20:13:06] !log jhancock@cumin2002 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage [20:13:26] whoot [20:13:54] (03CR) 10TrainBranchBot: [C:03+2] "Approved by samtar@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1313827 (https://phabricator.wikimedia.org/T408816) (owner: 10KineticPelagic) [20:14:41] 06SRE, 06DC-Ops, 10decommission-hardware, 06Data-Platform-SRE (2026-07-03 - 2026-07-31): decommission an-test-master100[1-2] - https://phabricator.wikimedia.org/T433495#12185132 (10VRiley-WMF) [20:15:08] (03Merged) 10jenkins-bot: rest: Add test server option to REST Sandbox for Wikipedia projects [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1313827 (https://phabricator.wikimedia.org/T408816) (owner: 10KineticPelagic) [20:15:16] 06SRE, 06DC-Ops, 10decommission-hardware, 06Data-Platform-SRE (2026-07-03 - 2026-07-31): decommission an-test-coord1001.eqiad.wmnet - https://phabricator.wikimedia.org/T433494#12185134 (10VRiley-WMF) [20:15:32] !log samtar@deploy1003 Started scap sync-world: Backport for [[gerrit:1313827|rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] [20:15:36] T408816: Add 'test' as a server option for Wikipedia projects within REST Sandbox - https://phabricator.wikimedia.org/T408816 [20:16:28] checking [20:18:01] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [20:21:21] !log samtar@deploy1003 samtar, kineticpelagic: Backport for [[gerrit:1313827|rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [20:21:26] T408816: Add 'test' as a server option for Wikipedia projects within REST Sandbox - https://phabricator.wikimedia.org/T408816 [20:21:42] hyang: should be okay to test now [20:21:49] checking! [20:24:04] (03PS1) 10Bking: cirrussearch: Enable performance governor on all master-eligible hosts [puppet] - 10https://gerrit.wikimedia.org/r/1321102 (https://phabricator.wikimedia.org/T324335) [20:24:35] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1321102 (https://phabricator.wikimedia.org/T324335) (owner: 10Bking) [20:26:17] (03CR) 10CI reject: [V:04-1] cirrussearch: Enable performance governor on all master-eligible hosts [puppet] - 10https://gerrit.wikimedia.org/r/1321102 (https://phabricator.wikimedia.org/T324335) (owner: 10Bking) [20:27:20] hyang: hows it looking? [20:28:18] looks fine so far. confirming with my teammate AaronSchulz who is helping me out [20:28:39] !log jhancock@cumin2002 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS bookworm [20:28:48] (ack) [20:29:08] (03PS2) 10Bking: cirrussearch: Enable performance governor on all master-eligible hosts [puppet] - 10https://gerrit.wikimedia.org/r/1321102 (https://phabricator.wikimedia.org/T324335) [20:30:01] (03PS3) 10Bking: cirrussearch: Enable performance governor on all master-eligible hosts [puppet] - 10https://gerrit.wikimedia.org/r/1321102 (https://phabricator.wikimedia.org/T324335) [20:30:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-int releases routed via main at eqiad: 21.44% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-int&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [20:30:57] FIRING: ProbeDown: Service chart-renderer:30443 has failed probes (http_chart-renderer_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#chart-renderer:30443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [20:33:22] verified [20:33:35] TheresNoTime: thank you for your patience [20:33:41] !log samtar@deploy1003 samtar, kineticpelagic: Continuing with deployment [20:33:55] no worries, continuing :) [20:34:05] PROBLEM - MariaDB Replica Lag: s8 on db1285 is CRITICAL: CRITICAL slave_sql_lag Replication lag: 2035.14 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [20:35:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-int releases routed via main at eqiad: 21.44% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-int&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [20:35:41] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1321102 (https://phabricator.wikimedia.org/T324335) (owner: 10Bking) [20:35:57] RESOLVED: ProbeDown: Service chart-renderer:30443 has failed probes (http_chart-renderer_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#chart-renderer:30443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [20:40:13] !log samtar@deploy1003 Finished scap sync-world: Backport for [[gerrit:1313827|rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] (duration: 24m 40s) [20:40:17] T408816: Add 'test' as a server option for Wikipedia projects within REST Sandbox - https://phabricator.wikimedia.org/T408816 [20:40:24] hyang: live on prod now [20:40:48] 10ops-codfw, 06SRE, 06DC-Ops, 06ServiceOps: Rerack mc-gp2006 and mc2046 - https://phabricator.wikimedia.org/T432716#12185216 (10Jhancock.wm) i got it done. ty for given me the chance to learn a new thing! [20:42:01] huzzah, thanks sammy! [20:42:20] (03PS4) 10Bking: cirrussearch: Enable performance governor on all master-eligible hosts [puppet] - 10https://gerrit.wikimedia.org/r/1321102 (https://phabricator.wikimedia.org/T324335) [20:43:11] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1321102 (https://phabricator.wikimedia.org/T324335) (owner: 10Bking) [20:49:51] np! :) [21:00:05] Deploy window Readers deployment window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260804T2100) [21:02:01] 10ops-eqiad, 06SRE, 10Data-Persistence-Backup, 06DC-Ops: db1245 crashed - https://phabricator.wikimedia.org/T431115#12185255 (10VRiley-WMF) I was given an action plan and will be following through with these steps. Reseat all fans, reseat the heatsink and cpu. update bios firmware. Firmware is now out of d... [21:03:28] (03PS5) 10Bking: cirrussearch: Enable performance governor on master-eligible hosts [puppet] - 10https://gerrit.wikimedia.org/r/1321102 (https://phabricator.wikimedia.org/T324335) [21:04:16] (03CR) 10CI reject: [V:04-1] cirrussearch: Enable performance governor on master-eligible hosts [puppet] - 10https://gerrit.wikimedia.org/r/1321102 (https://phabricator.wikimedia.org/T324335) (owner: 10Bking) [21:05:42] (03PS6) 10Bking: cirrussearch: Enable performance governor on master-eligible hosts [puppet] - 10https://gerrit.wikimedia.org/r/1321102 (https://phabricator.wikimedia.org/T324335) [21:06:49] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1321102 (https://phabricator.wikimedia.org/T324335) (owner: 10Bking) [21:07:54] (03CR) 10Krinkle: [C:03+1] "LGTM. Should be safe to rollout ahead of the apache change anytime." [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1315128 (https://phabricator.wikimedia.org/T432412) (owner: 10Tsevener) [21:12:05] RECOVERY - MariaDB Replica Lag: s8 on db1285 is OK: OK slave_sql_lag Replication lag: 0.03 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [21:13:28] !log ladsgroup@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1225.eqiad.wmnet with reason: Maintenance [21:14:19] 06SRE, 10Maps, 06Traffic: Possibility to allow Wikimedia Maps usage on all Wikibase Cloud instances - https://phabricator.wikimedia.org/T429191#12185282 (10MSantos) @Anton.Kokh, @ssingh or @Lydia_Pintscher do we have a sense of how much traffic that is going to bring to the maps servers? Usually I approve ad... [21:19:02] (03PS2) 10Majavah: Undeploy WP25EasterEggs (II) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1318203 (https://phabricator.wikimedia.org/T418134) [21:19:05] (03CR) 10Jforrester: [C:03+1] Undeploy WP25EasterEggs (II) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1318203 (https://phabricator.wikimedia.org/T418134) (owner: 10Majavah) [21:29:25] (03CR) 10Dzahn: [C:03+2] Bump buildkit image to wmf-v0.32.2 [puppet] - 10https://gerrit.wikimedia.org/r/1321081 (https://phabricator.wikimedia.org/T433985) (owner: 10Ahmon Dancy) [21:38:53] (03PS2) 10Jcrespo: mariadb: Remove all references on puppet to db1150 & db1171 [puppet] - 10https://gerrit.wikimedia.org/r/1320923 (https://phabricator.wikimedia.org/T433825) [21:38:53] (03PS3) 10Jcrespo: bacula: Remove last references to backup1003 & backup2003 [puppet] - 10https://gerrit.wikimedia.org/r/1320928 (https://phabricator.wikimedia.org/T420506) [21:38:53] (03PS1) 10Jcrespo: mediabackups: Update versitygw systemd unit file to SIGHUP on reload [puppet] - 10https://gerrit.wikimedia.org/r/1321123 [21:39:38] 10ops-eqiad, 06SRE, 10Data-Persistence-Backup, 06DC-Ops: db1245 crashed - https://phabricator.wikimedia.org/T431115#12185362 (10jcrespo) Thank you for the update. [22:10:28] 06SRE, 10SRE-Access-Requests: Grant access to analytics-privatedata-users for EAlbizzati-WMF - https://phabricator.wikimedia.org/T434034 (10EAlbizzati-WMF) 03NEW [22:11:31] 06SRE, 10SRE-Access-Requests: Grant access to analytics-privatedata-users for EAlbizzati-WMF - https://phabricator.wikimedia.org/T434034#12185447 (10EAlbizzati-WMF) [22:21:53] 06SRE, 06ServiceOps, 10ServiceOps-Upgrades-Hardware, 07Kubernetes, 13Patch-For-Review: wikikube-ctrl2006 implementation tracking - https://phabricator.wikimedia.org/T406596#12185454 (10RobH) a:05jasmine_→03Jhancock.wm @Jhancock.wm, This host was racked by you, but is currently reporting issues with... [22:23:00] !log ladsgroup@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1237.eqiad.wmnet with reason: Maintenance [22:23:46] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Depooling db1237 (T433990)', diff saved to https://phabricator.wikimedia.org/P95878 and previous config saved to /var/cache/conftool/dbconfig/20260804-222345-ladsgroup.json [22:23:50] T433990: Optimize echo tables in x1 - https://phabricator.wikimedia.org/T433990 [22:26:36] (03PS1) 10Scott French: Temporarily point all etcd client SRV records to codfw [dns] - 10https://gerrit.wikimedia.org/r/1321090 (https://phabricator.wikimedia.org/T428495) [22:26:39] (03PS2) 10Scott French: hieradata: Temporarily point eqiad PyBals at codfw etcd [puppet] - 10https://gerrit.wikimedia.org/r/1321089 (https://phabricator.wikimedia.org/T428495) [22:35:36] (03PS1) 10Cwhite: profile::opensearch: add scap to roles [puppet] - 10https://gerrit.wikimedia.org/r/1321141 (https://phabricator.wikimedia.org/T350516) [22:39:56] (03PS1) 10Scott French: hieradata: Use component/zookeeper34 on all configcluster hosts [puppet] - 10https://gerrit.wikimedia.org/r/1321106 (https://phabricator.wikimedia.org/T428495) [22:39:57] (03PS1) 10Scott French: hieradata: Prepare conf1007 for reimage to bookworm [puppet] - 10https://gerrit.wikimedia.org/r/1321107 (https://phabricator.wikimedia.org/T428495) [22:39:59] (03PS1) 10Scott French: hieradata: Temporarily move etcd replication to conf1007 [puppet] - 10https://gerrit.wikimedia.org/r/1321108 (https://phabricator.wikimedia.org/T428495) [22:40:01] (03PS2) 10Scott French: hieradata: Enable the v2 API and set cluster_bootstrap false in eqiad [puppet] - 10https://gerrit.wikimedia.org/r/1321109 (https://phabricator.wikimedia.org/T428495) [22:53:57] FIRING: CoreRouterInterfaceDown: Core router interface down - cr2-eqiad:xe-3/2/1 (Transport: cr1-esams:xe-0/0/7 (Colt, 445419311 80ms 10Gbps wave)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [22:58:01] FIRING: [2x] CoreRouterInterfaceDown: Core router interface down - cr1-esams:xe-0/0/7 (Transport: cr2-eqiad:xe-3/2/1 (Colt, 445419311 80ms 10Gbps wave)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [23:03:01] RESOLVED: [2x] CoreRouterInterfaceDown: Core router interface down - cr1-esams:xe-0/0/7 (Transport: cr2-eqiad:xe-3/2/1 (Colt, 445419311 80ms 10Gbps wave)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [23:07:16] 06SRE, 06App Experience, 06Content-Platform-Team, 06ServiceOps, and 2 others: Code 414 error when selecting zh-min-nan (nan) language from article list - https://phabricator.wikimedia.org/T434037 (10cooltey) 03NEW [23:08:01] FIRING: [2x] CoreRouterInterfaceDown: Core router interface down - cr1-esams:xe-0/0/7 (Transport: cr2-eqiad:xe-3/2/1 (Colt, 445419311 80ms 10Gbps wave)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [23:11:45] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Repooling after maintenance db1237 (T433990)', diff saved to https://phabricator.wikimedia.org/P95879 and previous config saved to /var/cache/conftool/dbconfig/20260804-231144-ladsgroup.json [23:11:50] T433990: Optimize echo tables in x1 - https://phabricator.wikimedia.org/T433990 [23:14:11] 06SRE, 06App Experience, 06Content-Platform-Team, 06ServiceOps, and 2 others: Code 414 error when selecting zh-min-nan (nan) language from article list - https://phabricator.wikimedia.org/T434037#12185568 (10cooltey) [23:16:31] 06SRE, 06App Experience, 06Content-Platform-Team, 06ServiceOps, and 2 others: Code 414 error when selecting zh-min-nan (nan) language from article list - https://phabricator.wikimedia.org/T434037#12185577 (10GitHubPRBot) cooltey opened https://github.com/wikimedia/apps-android-wikipedia/pull/6770 [23:18:01] RESOLVED: [2x] CoreRouterInterfaceDown: Core router interface down - cr1-esams:xe-0/0/7 (Transport: cr2-eqiad:xe-3/2/1 (Colt, 445419311 80ms 10Gbps wave)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [23:22:32] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95880 and previous config saved to /var/cache/conftool/dbconfig/20260804-232230-ladsgroup.json [23:31:23] (03CR) 10Ladsgroup: [C:03+1] "I merge it tomorrow morning unless people object" [puppet] - 10https://gerrit.wikimedia.org/r/1290104 (https://phabricator.wikimedia.org/T107188) (owner: 10Krinkle) [23:33:18] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95881 and previous config saved to /var/cache/conftool/dbconfig/20260804-233317-ladsgroup.json [23:34:21] (03CR) 10Samwilson: [C:03+1] "Looks good. I guess this is conflicting with the addition of `WebP_Flags` in I6651c9d040fd316e690cf9f1756c414cff45e30e." [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1270150 (https://phabricator.wikimedia.org/T290345) (owner: 10TheDJ) [23:35:16] (03CR) 10Samwilson: [C:03+1] "Oh oops, just noticed that this is from ages ago! But yeah, looks good. Do you want me to rebase it?" [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1270150 (https://phabricator.wikimedia.org/T290345) (owner: 10TheDJ) [23:35:34] (03PS4) 10Krinkle: mediawiki: Disable legacy `short_urls` on vhosts where it does not work [puppet] - 10https://gerrit.wikimedia.org/r/1290104 (https://phabricator.wikimedia.org/T107188) [23:42:21] (03PS1) 10TrainBranchBot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1321156 [23:42:21] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1321156 (owner: 10TrainBranchBot) [23:44:06] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Repooling after maintenance db1237 (T433990)', diff saved to https://phabricator.wikimedia.org/P95882 and previous config saved to /var/cache/conftool/dbconfig/20260804-234405-ladsgroup.json [23:44:10] T433990: Optimize echo tables in x1 - https://phabricator.wikimedia.org/T433990 [23:44:22] !log ladsgroup@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1264.eqiad.wmnet with reason: Maintenance [23:45:09] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Depooling db1264 (T433990)', diff saved to https://phabricator.wikimedia.org/P95883 and previous config saved to /var/cache/conftool/dbconfig/20260804-234508-ladsgroup.json [23:53:57] (03CR) 10CI reject: [V:04-1] Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1321156 (owner: 10TrainBranchBot)