[00:47:26] 10Continuous-Integration-Infrastructure, 06Reader Experience Team, 07ci-test-error (WMF-deployed Build Failure), 07Essential-Work, 10MediaWiki-Core-Skin-Architecture (Menus 2.0): CI tests fail with TypeError: array_map(): Argument #2 ($array) must be of... - https://phabricator.wikimedia.org/T422907#12126657 [07:25:33] (03PS1) 10Kevin Bazira: inference-services: Add CI pipeline jobs for tts-section-generator server [integration/config] - 10https://gerrit.wikimedia.org/r/1311317 (https://phabricator.wikimedia.org/T432308) [08:13:10] (03CR) 10Hashar: [C:03+2] "INFO:jenkins_jobs.builder:Number of jobs generated: 4" [integration/config] - 10https://gerrit.wikimedia.org/r/1311317 (https://phabricator.wikimedia.org/T432308) (owner: 10Kevin Bazira) [08:15:07] (03Merged) 10jenkins-bot: inference-services: Add CI pipeline jobs for tts-section-generator server [integration/config] - 10https://gerrit.wikimedia.org/r/1311317 (https://phabricator.wikimedia.org/T432308) (owner: 10Kevin Bazira) [08:17:00] (03CR) 10Hashar: [C:03+2] Fix sonar authentication [integration/gearman-java] - 10https://gerrit.wikimedia.org/r/1310210 (https://phabricator.wikimedia.org/T429547) (owner: 10Pwangai) [08:18:46] (03CR) 10Hashar: [C:03+2] Zuul: [BlueSpiceDistributionConnector] Add BlueSpiceSmartList [integration/config] - 10https://gerrit.wikimedia.org/r/1310056 (owner: 10Hslater) [08:20:15] (03CR) 10CI reject: [V:04-1] Fix sonar authentication [integration/gearman-java] - 10https://gerrit.wikimedia.org/r/1310210 (https://phabricator.wikimedia.org/T429547) (owner: 10Pwangai) [08:20:34] (03Merged) 10jenkins-bot: Zuul: [BlueSpiceDistributionConnector] Add BlueSpiceSmartList [integration/config] - 10https://gerrit.wikimedia.org/r/1310056 (owner: 10Hslater) [08:22:26] (03CR) 10Hashar: [C:03+2] "deployed" [integration/config] - 10https://gerrit.wikimedia.org/r/1310056 (owner: 10Hslater) [09:50:44] Project mediawiki-core-doxygen build #21518: 04FAILURE in 32 min: https://integration.wikimedia.org/ci/job/mediawiki-core-doxygen/21518/ [09:50:49] Project mwcore-phpunit-coverage-master build #5338: 04FAILURE in 50 min: https://integration.wikimedia.org/ci/job/mwcore-phpunit-coverage-master/5338/ [10:39:16] FIRING: DiskSpace: Disk space contint1003:9100:/srv 0% free [10:39:16] There is no more Disk space on contint1003 in /srv left. It looks like /srv/docker is using most of the space. [10:44:00] is there some non-destructive way to clean up old files? also it looks like vfs is used on contint which is wasting a lot of space [10:50:31] Project mediawiki-core-doxygen build #21519: 04STILL FAILING in 32 min: https://integration.wikimedia.org/ci/job/mediawiki-core-doxygen/21519/ [11:17:16] Zuul gate appears stuck. 3 jobs there all three idle for 45min with no execution having begun [11:18:15] I guess because of DiskSpace [11:18:31] * Krinkle cancels deployment and goes for lunch [11:18:49] https://integration.wikimedia.org/zuul/ [12:01:59] I freed up disk space but I think the jobs have to be canceled and re-trigged [12:02:25] jelto: ok, what did you remove :D [12:02:59] I pruned the build cache [12:03:52] https://grafana.wikimedia.org/d/000000377/host-overview?orgId=1&from=now-3h&to=now&timezone=utc&var-server=contint1003 [12:04:41] 500GB -> 1.4 TB -> 280GB [12:05:55] or for now-6h, 280GB -> 1.4 TB -> 280GB [12:06:29] quite sudden growth [12:06:34] seems to be a recent phenomenon [12:06:44] yeah something started to fill up the disk at 08:44 UTC [12:06:48] none as bad as this one, but its' been happening a few times [12:07:25] 12 days: https://grafana.wikimedia.org/d/000000377/host-overview?orgId=1&from=now-12d&to=now&timezone=utc&var-server=contint1003&var-datasource=000000026&var-cluster=ci&refresh=5m&viewPanel=panel-28 [12:08:17] * Krinkle tries again [12:10:22] still clogged [12:10:49] !log krinkle@contint1002: zuul restart [12:10:50] Logged the message at https://wikitech.wikimedia.org/wiki/Release_Engineering/SAL [12:13:17] Error: NOT_REGISTERED [12:13:33] https://phabricator.wikimedia.org/T75658#770074 [12:13:41] looks like that is because Jenkins is unreachable [12:15:09] Jul 16 12:12:12 contint1002 zuul-server[1913017]: 2026-07-16 12:12:12,572 ERROR zuul.Gearman: Job [12:16:46] dancy: hashar: any ideas :) [12:17:55] the jenkins webserver is up, but contint1002 has no jenkins process running [12:18:59] contint 5M IN CNAME contint1002.wikimedia.org; thats's where zuul commands are sent when using fab, which I assume is up to date [12:19:58] jenkins 300 IN CNAME contint1003.wikimedia.org. [12:20:11] I dont' know if jenkins.eqiad.wmnet is something we use, but if yes, that means it's not on the same host [12:20:20] https://wikitech.wikimedia.org/wiki/Jenkins says codfw is primary [12:21:10] T432326 [12:21:11] T432326: DiskSpace (contint1003) - https://phabricator.wikimedia.org/T432326 [12:21:41] 06Release-Engineering-Team (Radar), 06collaboration-services: DiskSpace (contint1003) - https://phabricator.wikimedia.org/T432326#12128054 (10Krinkle) [12:21:51] 10Continuous-Integration-Infrastructure, 06Release-Engineering-Team (Radar), 06collaboration-services: DiskSpace (contint1003) - https://phabricator.wikimedia.org/T432326#12128056 (10Krinkle) [12:22:42] T418521 confirms zuul is on the old one, jenkins on the new one [12:22:43] T418521: setup 2 contint machines for jenkins - https://phabricator.wikimedia.org/T418521 [12:23:55] Jul 16 12:13:53 contint1003 jenkins[1119]: WARNING: [hudson.plugins.sshslaves.SSHLauncher launch] SSH Launch of contint2003 on contint2003.wikimedia.org failed in 60,064 ms [12:24:38] probably not important since that's not primary? [12:30:57] ok, jenkins restat seems to have fixed it [12:31:31] !log krinkle@contint1003:~$ sudo /usr/sbin/service jenkins restart [12:31:32] Logged the message at https://wikitech.wikimedia.org/wiki/Release_Engineering/SAL [12:31:52] https://www.mediawiki.org/w/index.php?title=Continuous_integration%2FJenkins&diff=8508987&oldid=8462095 [12:54:18] (03PS1) 10Gerrit Patch Uploader: Zuul: Add User:1F616EMO to CI allowlist Change-Id: I105eb704da4eb96fe4037da1111c861391c18907 [integration/config] - 10https://gerrit.wikimedia.org/r/1311455 [12:54:18] (03CR) 10Gerrit Patch Uploader: "This commit was uploaded using the Gerrit Patch Uploader [1]." [integration/config] - 10https://gerrit.wikimedia.org/r/1311455 (owner: 10Gerrit Patch Uploader) [12:54:45] (03PS2) 10Neriah: Zuul: Add User:1F616EMO to CI allowlist [integration/config] - 10https://gerrit.wikimedia.org/r/1311455 (owner: 10Gerrit Patch Uploader) [12:59:12] Project mediawiki-core-doxygen build #21520: 04STILL FAILING in 32 min: https://integration.wikimedia.org/ci/job/mediawiki-core-doxygen/21520/ [13:01:49] Did https://gerrit.wikimedia.org/r/c/mediawiki/core/+/1281992 get hit by the diskspace issues earlier? Seems to suggest it went to start jobs but no jobs were triggered? [13:12:35] Yeah looks like it, I've triggered +2 again and it seems to have started the jobs again [13:16:07] (03PS1) 10Majavah: Review access change [jquery-client] (refs/meta/config) - 10https://gerrit.wikimedia.org/r/1311458 [13:32:13] Project mediawiki-core-doxygen build #21521: 04STILL FAILING in 32 min: https://integration.wikimedia.org/ci/job/mediawiki-core-doxygen/21521/ [13:39:28] (03CR) 10Jforrester: "I vaguely recall us saying we were going to drop this library, but until we do I suppose this makes sense. Timo, do you agree?" [jquery-client] (refs/meta/config) - 10https://gerrit.wikimedia.org/r/1311458 (owner: 10Majavah) [13:47:39] 10Continuous-Integration-Infrastructure, 06Release-Engineering-Team (Radar), 06collaboration-services: DiskSpace (contint1003) - https://phabricator.wikimedia.org/T432326#12128371 (10Jelto) The disk usage is going up again https://grafana.wikimedia.org/d/000000377/host-overview?orgId=1&from=now-6h&to=now&ti... [13:58:22] (03CR) 10Krinkle: "LGTM. How is this configured on other repos? I don't see a libup grant on WrappedString, mediawiki/libs/less.php, etc. I don't see it in t" [jquery-client] (refs/meta/config) - 10https://gerrit.wikimedia.org/r/1311458 (owner: 10Majavah) [13:59:53] (03CR) 10Majavah: "LibUp is in the `mediawiki` group so that's where the access comes from in repos inheriting `mediawiki` or some of its subdirectories. +1 " [jquery-client] (refs/meta/config) - 10https://gerrit.wikimedia.org/r/1311458 (owner: 10Majavah) [14:05:01] Project mediawiki-core-doxygen build #21522: 04STILL FAILING in 32 min: https://integration.wikimedia.org/ci/job/mediawiki-core-doxygen/21522/ [14:20:49] Yippee, build fixed! [14:20:49] Project mediawiki-core-doxygen build #21523: 09FIXED in 15 min: https://integration.wikimedia.org/ci/job/mediawiki-core-doxygen/21523/ [14:22:13] 10Continuous-Integration-Config, 13Patch-For-Review, 06Quality-and-Test-Engineering-Team (SonarCloud Admin), 06Test Platform (Aktau 28): Migrate sonar analysis from deprecated sonar.login to sonar.token - https://phabricator.wikimedia.org/T429547#12128510 (10pwangai) [14:26:58] (03CR) 10Krinkle: "Ha, that works. I didn't know we added bots to that group (given Jenkins, i18n-bot etc separate as rules under the mediawiki repo parent, " [jquery-client] (refs/meta/config) - 10https://gerrit.wikimedia.org/r/1311458 (owner: 10Majavah) [14:29:06] (03PS1) 10Krinkle: Review access change [jquery-client] (refs/meta/config) - 10https://gerrit.wikimedia.org/r/1311473 [14:30:01] (03CR) 10Krinkle: "I can't save this myself (access denied) but here you go: https://gerrit.wikimedia.org/r/c/jquery-client/+/1311473" [jquery-client] (refs/meta/config) - 10https://gerrit.wikimedia.org/r/1311458 (owner: 10Majavah) [14:30:15] (03CR) 10Majavah: [C:03+2] Review access change [jquery-client] (refs/meta/config) - 10https://gerrit.wikimedia.org/r/1311473 (owner: 10Krinkle) [14:30:18] (03CR) 10Majavah: [V:03+2 C:03+2] Review access change [jquery-client] (refs/meta/config) - 10https://gerrit.wikimedia.org/r/1311473 (owner: 10Krinkle) [14:30:28] (03Abandoned) 10Majavah: Review access change [jquery-client] (refs/meta/config) - 10https://gerrit.wikimedia.org/r/1311458 (owner: 10Majavah) [14:35:28] Krinkle: I'm at my desk now if there are still problems. [14:35:55] dancy: nothing active but it may be worth looking at what happened [14:36:43] I suppose we're reasonably confident zuul/jenkins froze up due to lack of space, so it'd be more about what happened a few days ago that started this suden growth in disk space in short bursts throughout recent days, each time worse than the previous one [14:36:54] 1TB in an hour or so. [14:37:08] oh wow. OK I'll look around. [14:38:24] from backscroll: https://grafana.wikimedia.org/d/000000377/host-overview?orgId=1&from=now-12d&to=now&timezone=utc&var-server=contint1003&var-datasource=000000026&var-cluster=ci&refresh=5m&viewPanel=panel-28 [14:38:54] Krinkle: According to https://phabricator.wikimedia.org/T432326 it's because Docker on that system isn't using the right storage driver. [14:39:34] dancy: hm.. maybe it's correlated with the debian upgrade? [14:40:30] ref T418521 [14:40:31] T418521: setup 2 contint machines for jenkins - https://phabricator.wikimedia.org/T418521 [14:40:45] dancy: Krinkle: most certainly [14:40:57] that happened a while ago due to overlayfs kernel module not being available https://gerrit.wikimedia.org/r/c/operations/puppet/+/402797 [14:41:00] * Krinkle steps back into the bushes. Have fun :) [14:41:02] https://gerrit.wikimedia.org/r/c/operations/puppet/+/402797 [14:41:28] hieradata/role/common/ci.yaml:profile::base::overlayfs: true [14:41:36] that is what we had on the old hosts (contint1002 / contint2002) [14:41:40] Ah, missing for contint1003 [14:41:42] ? [14:41:46] the new hosts (contint1003 / contint2003) have the role `jenkins` [14:42:46] so maybe the kernel module is missing [14:43:08] (I am braindumping that after you mentioned T432326 "Docker does not have the right storage driver" [14:43:09] T432326: DiskSpace (contint1003) - https://phabricator.wikimedia.org/T432326 [14:43:49] `/etc/modprobe.d/blacklist-wmf_overlay.conf` does exist on contint1003. [14:45:28] 10Continuous-Integration-Infrastructure, 06Release-Engineering-Team (Radar), 06collaboration-services: DiskSpace (contint1003) - https://phabricator.wikimedia.org/T432326#12128622 (10dancy) >>! In T432326#12127808, @Jelto wrote: > I'll leave this task open to either replace `vfs` with a proper storage driver... [14:45:45] and when changing the storage driver, I think the directory hsa to be nuked when Docker is stopped. On the next stop it will catch up and recreate the layer/fs/storage system with whatever it found [14:45:53] Nod [14:45:55] OR maybe the storage system has to be explicitly added to the daemon conf [14:45:58] I don't remember [14:46:37] whoever has mentioned `Storage Driver: vfs`Β  most certainly have found the root cause [14:47:01] Nod. The giant /srv/docker/vfs directory was a smoking gun [14:47:02] and my guess that is Docker ultimate fallback when it can't use any other driver (eg overlayfs because of lack of a kernel module) [14:47:12] Agreed [14:47:58] it might be necessary to nuke the image and volumes sub dir [14:48:14] or maybe easier: stop docker, nuke everything below /srv/docker, start docker and it should recreate everything it needs [14:48:29] the only loss are the pipelinelib caching layers which is imho not a big deal (they are caches) [14:49:04] [I am just brain dumping everything I can remember] [14:51:02] I am off again, thank you dancy ! [14:51:12] Thanks hashar! I'll push it along [14:51:28] πŸ‘ πŸŽ–οΈ [14:55:06] 10Continuous-Integration-Infrastructure, 06Release-Engineering-Team (Radar), 06collaboration-services: DiskSpace (contint1003) - https://phabricator.wikimedia.org/T432326#12128660 (10dancy) >>! In T432326#12128622, @dancy wrote: >>>! In T432326#12127808, @Jelto wrote: >> I'll leave this task open to either r... [14:58:17] "Network is unreachable" from scap backport [14:58:25] mergeMessageFileList.php generated PHP notices/warnings: [14:58:25] Warning: socket_sendto(): Unable to write to socket [101]: Network is unreachable in /srv/mediawiki-staging/php-1.47.0-wmf.10/includes/Logger/Monolog/LegacyHandler.php on line 215 [15:00:01] I guess it's not meant to be today. [15:01:07] ref T410738 [15:01:07] T410738: pretrain failing when calling mergeMessageFileList.php - https://phabricator.wikimedia.org/T410738 [15:01:12] I'll self-revert for now as I need to do something else now [15:03:17] Krinkle: Did you get a stack trace? [15:04:50] Krinkle: oh, nevermind. I see that the same thing is happening to wmf-beta-update-all , so somebody merged something broken recently. [15:04:57] about 2 hours ago. [15:05:25] or, beta logstash is having its own problem. [15:14:58] Ah, you mean there's a genuine error but it's not making it to stderr because it tries to send to Logstash first which we block in this container [15:15:05] Right [15:15:19] The most recent wmf-beta-update-all run succeeded (the one with your revert) [15:16:35] is train log triage happening today? or perhaps after a little while, once things here are resolved? [15:27:37] Yippee, build fixed! [15:27:38] Project mwcore-phpunit-coverage-master build #5339: 09FIXED in 27 min: https://integration.wikimedia.org/ci/job/mwcore-phpunit-coverage-master/5339/ [15:35:44] 10GitLab (CI & Job Runners), 06Release-Engineering-Team, 07Essential-Work: Buildkit v0.31.2 released - https://phabricator.wikimedia.org/T432360 (10dancy) 03NEW [15:40:09] The beta mediawiki update task seems so much flakier now than when it was run via Jenkins, but I cannot explain why that would be. The things that are failing are commands from inside `scap` that are no different at all from the prior triggering mechanism. Same commands, same user, same runtime host, I think even same external target hosts mostly. Β―\_(ツ)_/Β― [15:40:48] hard to tell when it's no longer alerting here [15:42:23] I was not smart enough to understand your advice taavi. The alerts do go to the releng@lists.wikimedia.org mailing list right now. [15:45:19] which part of that? or the whole process? [15:45:51] bd808: I think it just feels that way because we get an alert for every failure, not just alerts for the transitions between ok and not-ok. [15:46:24] And the alerts show up in our inbox rather than as IRC messages flowing by [15:46:44] taavi: I will write up a new task and then maybe can ask you more informed questions. :) [15:47:19] dancy: that actually is a really good observation. [15:51:19] 10Beta-Cluster-Infrastructure: Restore IRC alerting for failed MediaWiki updates in Beta Cluster - https://phabricator.wikimedia.org/T432364 (10bd808) 03NEW [15:53:22] Speaking of flaky, I'm considering restarting Jenkins. There are a bunch of puppet changes with queued operations-puppet-tests-bullseye jobs that are never processing. [15:53:47] Tons of mediawiki/extensions/CampaignEvents patches are lurking [15:54:21] I'll give it an hour to see how things change. [15:54:47] did someone push a massive chain again? The https://gerrit.wikimedia.org/r/c/operations/deployment-charts/+/1311413 merge in ops/puppet was mentioned as having problems in slack earlier today. [15:55:10] there is an absolutely massive chain in CampaignEvents [15:55:26] 26 stacked patches. ouch [15:55:41] seems like Reedy already asked them to not do that but it was ignored [15:55:47] Nod. [15:56:02] Maybe too nicely? "Can we please be careful..." [15:57:22] 21 patchsets on the patch at the bottom of the stack. We need the new zuul so this doesn't keep being the cause of self denial of service attacks. [16:15:28] in the meanwhile we may need to be a bit more direct: knock that shit off. [16:16:53] last time it did start to clear on its own right when we were on the point of either a restart or trying this: https://www.mediawiki.org/wiki/Continuous_integration/Zuul#Very_high_queue_of_merger:merge_functions [16:18:19] I added a bit more direct comment [16:20:53] thx taavi [16:22:28] (03CR) 10Jforrester: "WFM, thanks!" [jquery-client] (refs/meta/config) - 10https://gerrit.wikimedia.org/r/1311473 (owner: 10Krinkle) [16:24:11] 10Continuous-Integration-Infrastructure, 06Release-Engineering-Team (Radar), 06collaboration-services, 13Patch-For-Review: DiskSpace (contint1003) - https://phabricator.wikimedia.org/T432326#12129176 (10Dzahn) [16:26:30] 10Continuous-Integration-Infrastructure, 06Release-Engineering-Team (Radar), 06collaboration-services, 13Patch-For-Review: DiskSpace (contint1003) - https://phabricator.wikimedia.org/T432326#12129186 (10Dzahn) thanks all! patch has been deployed ` [contint1003:~] $ lsmod | grep overlay overlay... [16:26:55] queue is at 436, down from ~480 5 minutes ago or so. [16:27:34] i guess my inclination is to let this ride but fast-fail the CampaignEvents ones if it starts spiking up again. [16:27:47] 10Beta-Cluster-Infrastructure: Restore IRC alerting for failed MediaWiki updates in Beta Cluster - https://phabricator.wikimedia.org/T432364#12129190 (10bd808) [16:27:48] 10Beta-Cluster-Infrastructure, 10observability, 07Tracking-Neverending: Setup monitoring for Beta Cluster (tracking) - https://phabricator.wikimedia.org/T53497#12129189 (10bd808) [16:27:56] (above cc: dancy) [16:28:19] taavi: You were extremely polite. No suggestion of remove C+2 rights that I'd expect from h.ashar. ;-) [16:30:08] brennen: Confirmed. Thanks [16:30:10] ohhi! [16:30:27] James_F: sadly you can push massive sets of patches without +2 and I fear that me showing up and threatening developer account blocks on the first interaction would be seen as "unconstructive" [16:30:50] queue now at 400 [16:30:53] taavi: ISTR the limit is 5 not 20 unless you're in ldap/mw. [16:31:12] is it? [16:31:14] maybe the limit should just be 5? [16:31:29] ^ [16:31:30] having some sort of a limit would be certainly a smart thing to do [16:31:44] I feel like I'm regularly pushing things that are in the range of over 5 but under 10 [16:31:46] James_F: what if I stop, remove the +2, and try again in a bit? https://gerrit.wikimedia.org/r/c/operations/mediawiki-config/+/1311466 [16:31:58] effie: It'll go later, not sooner. [16:32:03] I know there is a limit. I used to get told by git-review that I made too big of a mess to upload. [16:32:20] effie: See the long threads of commits on https://integration.wikimedia.org/zuul/ [16:32:45] !log Restarted docker on contint1003 [16:32:46] effie: if that's an urgent change, i can try and clear out the queue with a hack that fails the ones clogging things up. [16:32:46] Logged the message at https://wikitech.wikimedia.org/wiki/Release_Engineering/SAL [16:32:57] I don't see a limit configured at least on https://gerrit.wikimedia.org/r/admin/repos/All-Projects,access or https://gerrit.wikimedia.org/r/admin/repos/mediawiki,access and I don't know where else it would be on [16:33:03] !log Restarted docker on contint1003 (T432326) [16:33:05] Logged the message at https://wikitech.wikimedia.org/wiki/Release_Engineering/SAL [16:33:05] T432326: DiskSpace (contint1003) - https://phabricator.wikimedia.org/T432326 [16:33:36] taavi: It's in zuul-cloner itself, I think? Or something like that. [16:33:40] Anyway, meetings. :-( [16:34:17] !log Restarted docker on contint2003 (T432326) [16:34:20] Logged the message at https://wikitech.wikimedia.org/wiki/Release_Engineering/SAL [16:34:53] brennen: it is not urgent per se, I am rebooting some servers and I need to deploy wmf-config each time [16:35:02] 10Continuous-Integration-Infrastructure, 06Release-Engineering-Team (Radar), 06collaboration-services: DiskSpace (contint1003) - https://phabricator.wikimedia.org/T432326#12129256 (10dancy) 05Openβ†’03Resolved p:05Triageβ†’03High a:03dancy [16:35:19] brennen: but given I am mid-scapping, may be we should? [16:36:59] yeah, i don't really like having scap jammed up and this could realistically take another 30-60 minutes. [16:37:11] let me try this thing and see what happens. [16:37:46] effie: In this case I'd just force-land the config patch. [16:37:46] 10Continuous-Integration-Infrastructure, 06Release-Engineering-Team (Radar), 06collaboration-services: DiskSpace (contint1003) - https://phabricator.wikimedia.org/T432326#12129263 (10Dzahn) {F94033905} [16:37:54] thank you, it is 19:37 here and even if I want to leave things as they are, I will need to push yet another patch [16:39:20] James_F: I would like to be polite to computers, hopefully when the rise of the machines happens, they will remember [16:40:34] !log contint1002: chown and chmod on /srv/zuul/git/mediawiki/extensions/CampaignEvents/.git per https://www.mediawiki.org/wiki/Continuous_integration/Zuul#Very_high_queue_of_merger:merge_functions [16:40:40] Logged the message at https://wikitech.wikimedia.org/wiki/Release_Engineering/SAL [16:41:10] that knocked the queue down to 165 and dropping relatively quickly [16:41:30] fingers crossed [16:42:48] 107 now [16:45:09] 06Release-Engineering-Team (Priority Backlog πŸ“₯), 07Epic, 06Quality-and-Test-Engineering-Team (Test Infrastructure): Pretrain (nΓ©e Group -1) QTE validation environment - https://phabricator.wikimedia.org/T369112#12129315 (10bd808) [16:46:36] 06Release-Engineering-Team (Priority Backlog πŸ“₯): Collect data on merge to deploy cycle times for the rolling `wmf/next` branch so we can report on Pretrain impacts - https://phabricator.wikimedia.org/T420790#12129322 (10bd808) [16:46:38] 06Release-Engineering-Team (Priority Backlog πŸ“₯), 07Epic, 06Quality-and-Test-Engineering-Team (Test Infrastructure): Pretrain (nΓ©e Group -1) QTE validation environment - https://phabricator.wikimedia.org/T369112#12129323 (10bd808) [16:47:34] and I've just seen the title T424736 so immediately my thought is that this is a plot to get merges to happen slower so that the time between merge and deployment gets shorter [16:47:34] T424736: ST5.1 - Time from code merge to production testing availability - https://phabricator.wikimedia.org/T424736 [16:49:39] there are indeed often many ways to influence a single point metric [16:51:39] brennen: I have a feeling, my patch is not in the queue to be submitted [16:52:17] hrm, yeah [16:52:18] oh no wait [16:52:21] maybe? [16:52:39] yea it is not [16:52:46] sigh [16:53:19] I do not see 1311466 in the queue [16:54:58] effie: I've re-C+2'ed it. [16:55:11] I'm writing a quick patch (ha!) for integration/zuul. [16:56:35] despite our brave +2 efforts, jenkins will not badge [16:57:10] can't we just, merge it ? [16:57:20] (thank you folks so much!) [16:59:10] I think https://gerrit-review.googlesource.com/Documentation/config-gerrit.html#change.maxSubmittableAtOnce is maybe the setting to control the stacked change depth? I can't find it set anywhere yet (codesearch, github search, git grep in ops/puppet), but I'd swear that I have had to adjust patch chains before when git-review would refuse to send a stack up to gerrit. [17:00:09] effie: it's finally in the queue. things may have just been kind of jammed up on CampaignEvents [17:00:31] should be pretty quick at this point. [17:00:38] 10Beta-Cluster-Infrastructure: Restore IRC alerting for failed MediaWiki updates in Beta Cluster - https://phabricator.wikimedia.org/T432364#12129375 (10bd808) >>! In T256168#11960239, @taavi wrote: >>>! In T256168#11960221, @bd808 wrote: >> On my side we are still missing any notification if the sync jobs break... [17:00:50] thank you for coming to my rescue folks [17:01:04] bd808: That's on the gerrit side, we want to do this on the zuul side I think. [17:01:28] ... was not enjoying my time stuck in the Jenkins tower [17:01:46] effie: apologies for the amount of legacy zuul jank currently leaking into the deployment user experience. [17:01:53] James_F: the zuul side "fix" is the newer zuul. h.ashar was explaining it in a meeting recently [17:02:15] work is underway to upgrade all this but it's... a nontrivial lift for all the reasons it hasn't happened yet. [17:02:21] cheers brennen! [17:02:32] FIRING: InstanceDown: Project deployment-prep instance deployment-cache-text08 is down - https://prometheus-alerts.wmcloud.org/?q=alertname%3DInstanceDown [17:02:34] bd808: Yes, I am monkey-patching an equivalent for our current Zuul v2 world. [17:02:39] 10Beta-Cluster-Infrastructure: Project deployment-prep instance deployment-cache-text08 is down - https://phabricator.wikimedia.org/T432374 (10wmcs-alerts) 03NEW [17:03:33] James_F: yeah, that seems reasonable in the likely event that it will save a bunch of hours putting out fires that could go to work on the newer zuul. [17:04:06] brennen: Exactly. Let's not leave the rake in the grass whilst we're (you're) meant to be busy on the… shiny new field? Metaphor fun. [17:04:33] i got up quite early this morning to mow and nearly had a literal rake-in-grass experience. [17:04:59] feels emblematic of the moment. [17:05:46] * bd808 has a sudden memory of a lawn mower flinging a dog bone through a basement window [17:06:24] i can hear that sound [17:07:04] !log Hard reboot of deployment-cache-text08.deployment-prep.eqiad1.wikimedia.cloud via Horizon; console shows OOM (T432374) [17:07:06] Logged the message at https://wikitech.wikimedia.org/wiki/Release_Engineering/SAL [17:07:07] T432374: Project deployment-prep instance deployment-cache-text08 is down - https://phabricator.wikimedia.org/T432374 [17:07:32] RESOLVED: InstanceDown: Project deployment-prep instance deployment-cache-text08 is down - https://prometheus-alerts.wmcloud.org/?q=alertname%3DInstanceDown [17:08:52] taavi: https://phabricator.wikimedia.org/T432364#12129373 has my uninformed questions for subject matter experts such as yourself on prometheus/alertmanager/grafana things in WMCS. [17:09:29] πŸŽ‰ [17:17:37] 10Beta-Cluster-Infrastructure: Restore IRC alerting for failed MediaWiki updates in Beta Cluster - https://phabricator.wikimedia.org/T432364#12129461 (10taavi) >>! In T432364#12129373, @bd808 wrote: > * Is there already a prometheus instance collecting the `node_systemd_unit_state` metric for the `wmf-beta-updat... [17:23:30] 10Beta-Cluster-Infrastructure, 06Data-Platform-SRE (2026-07-03 - 2026-07-31): No Puppet resources found on instance deployment-cirrussearch13 on project deployment-prep - https://phabricator.wikimedia.org/T428822#12129478 (10bking) 05In progressβ†’03Resolved Instance `deployment-cirrussearch13` no longer... [17:29:49] 10Beta-Cluster-Infrastructure: Project deployment-prep instance deployment-cache-text08 is down - https://phabricator.wikimedia.org/T432374#12129505 (10bd808) 05Openβ†’03Resolved a:03bd808 [17:36:08] James_F: I had been wondering about spending a few hours hacking on the existing zuul to avoid this situation. I'm glad you decided to pick it up. [17:36:39] Let me know if I can assist. [17:40:30] There is a stuck job at the top of the operations/deployment-charts gate-and-submit queue. Zuul is reporting "https://integration.wikimedia.org/ci/job/helm-lint/None/console" as the queued job. This is borked, but I can't remember how to make zuul give up on that one change so it can move on with the rest of the queue. [17:40:37] https://gerrit.wikimedia.org/r/c/operations/deployment-charts/+/1311418 is the gerrit change. [17:43:26] scrounging around for any docs on this [17:44:06] https://www.mediawiki.org/wiki/Continuous_integration/Zuul#Replay_events maybe? [17:44:39] maybe? [17:45:01] (03PS1) 10Jforrester: WMF: Cap dependency-stack depth to throttle oversized patch series [integration/zuul] (patch-queue/debian/jessie-wikimedia) - 10https://gerrit.wikimedia.org/r/1311501 [17:45:26] a naive dance of +2,-2,+2 did nothing visible [17:45:40] dancy: ^ That might fix it. [17:45:49] * dancy takes a look [17:45:59] clicked "rebuild" on job 34038 [17:46:13] which is helm-lint and linked from that change [17:46:35] Thanks mutante! [17:46:45] https://gerrit.wikimedia.org/r/c/operations/deployment-charts/+/1311418 [17:46:48] ready to submit ^ [17:48:10] it has been in ready to submit for 3 hours. The bug is that the gate-and-submit job has a broken jenkins job spec that is never going to run. [17:48:54] bd808: It's a not-based-off-master issue again, I think. [17:49:09] it's starting gate-and-submit now [17:49:10] Now I rebased 1311418 it's running CI. [17:49:18] oh! It moved to the bottom of the queue. You may have done the needful mutante [17:49:33] maybe the "rebuild" followed by "recheck" [17:49:51] or James_F or whoever. anyway thanks universe [17:50:26] rebuild,recheck,rebase - rejoice [17:50:56] poke stick in zuul's cage, rattle [17:50:58] recheck shouldn't do anything for the gate-and-submit jobs [17:51:12] wait for growling [17:51:36] (except if you already have a +2 vote on the patch, at which point it sees it as a comment with +2 which may trigger gate-and-submit again :D how clear and simple) [17:51:53] getting up to speed, is there a specific incantation of these words that's sure to get your gate-and-submit going? we're in kind of a fiddly situation with poolcounter at the moment, and waiting on the backport of https://gerrit.wikimedia.org/r/c/operations/mediawiki-config/+/1311468 [17:52:09] * bd808 has the "everything moved on the page" eye twitch with a https://integration.wikimedia.org/zuul/ refresh [17:54:19] swfrench-wmf: I don't see that one in the queue yet, but why... [17:54:48] puzzling [17:55:03] try rebasing? [17:55:17] yeah, lemme cycle the +2 with a rebase in between [17:58:32] "Waiting for jobs" - I like the sound of that [17:58:40] thanks for the tip :) [17:59:05] I see the jobs running there now too swfrench-wmf [17:59:18] and off we go ... thanks, folks [18:27:12] (03CR) 10Ahmon Dancy: "I love this idea. The implementation looks straightforward. I just have some comments about text." [integration/zuul] (patch-queue/debian/jessie-wikimedia) - 10https://gerrit.wikimedia.org/r/1311501 (owner: 10Jforrester) [18:27:49] I think we have a stuck operations-puppet-tests-bullseye job holding things up too. Looking into that. [18:28:47] thanks! yeah, it looks like the last build was at 14:41 UTC [18:28:59] ... which was a while ago :) [18:33:38] !log Jenkins placed in shutdown mode in preparation for a restart [18:33:39] Logged the message at https://wikitech.wikimedia.org/wiki/Release_Engineering/SAL [18:37:04] !log Restarted Jenkins to unstick jobs. [18:37:05] Logged the message at https://wikitech.wikimedia.org/wiki/Release_Engineering/SAL [18:37:40] btw mutante: contint.wikimedia.org still resolves to contint1002 [18:38:07] `operations-puppet-tests-bullseye` jobs are running again. [18:38:27] I see progress! :) [18:38:29] thank you [18:40:42] 06Release-Engineering-Team, 10Catalyst (Luka Ijo Pimeja Jan), 07Essential-Work, 13Patch-For-Review: Add RSS extension to PatchDemo - https://phabricator.wikimedia.org/T426130#12129785 (10jeena) >>! In T426130#12126523, @Danielyepezgarces wrote: > I don't have a specific one in mind right now, but it would... [18:51:29] !log Updated buildkit to v0.31.2 gitlab-cloud-runners (staging and production) (T432360) [18:51:31] Logged the message at https://wikitech.wikimedia.org/wiki/Release_Engineering/SAL [18:51:31] T432360: Buildkit v0.31.2 released - https://phabricator.wikimedia.org/T432360 [18:53:44] (03open) 10dancy: staging and prod: Use buildkitd wmf-v0.31.2 [repos/releng/gitlab-cloud-runner] - 10https://gitlab.wikimedia.org/repos/releng/gitlab-cloud-runner/-/merge_requests/622 (https://phabricator.wikimedia.org/T432360) [18:54:18] (03update) 10dancy: staging and prod: Use buildkitd wmf-v0.31.2 [auto-approve] [repos/releng/gitlab-cloud-runner] - 10https://gitlab.wikimedia.org/repos/releng/gitlab-cloud-runner/-/merge_requests/622 (https://phabricator.wikimedia.org/T432360) [18:54:20] (03update) 10dancy: staging and prod: Use buildkitd wmf-v0.31.2 [auto-approve] [repos/releng/gitlab-cloud-runner] - 10https://gitlab.wikimedia.org/repos/releng/gitlab-cloud-runner/-/merge_requests/622 (https://phabricator.wikimedia.org/T432360) [18:55:24] (03merge) 10dancy: staging and prod: Use buildkitd wmf-v0.31.2 [auto-approve] [repos/releng/gitlab-cloud-runner] - 10https://gitlab.wikimedia.org/repos/releng/gitlab-cloud-runner/-/merge_requests/622 (https://phabricator.wikimedia.org/T432360) [19:05:01] 06Release-Engineering-Team (Doing 😎), 13Patch-For-Review: Allow configuration of canary and production checks based on deployment target - https://phabricator.wikimedia.org/T428971#12129857 (10dancy) 05In progressβ†’03Resolved [19:05:16] (03PS1) 10Fomafix: Zuul: [mediawiki/extensions/GlobalContributions] Add CentralAuth Phan dependency [integration/config] - 10https://gerrit.wikimedia.org/r/1311525 [19:11:41] (03PS2) 10Jforrester: WMF: Cap dependency-stack depth to throttle oversized patch series [integration/zuul] (patch-queue/debian/jessie-wikimedia) - 10https://gerrit.wikimedia.org/r/1311501 [19:11:41] (03CR) 10Jforrester: WMF: Cap dependency-stack depth to throttle oversized patch series (034 comments) [integration/zuul] (patch-queue/debian/jessie-wikimedia) - 10https://gerrit.wikimedia.org/r/1311501 (owner: 10Jforrester) [19:14:54] (03CR) 10Ahmon Dancy: [C:03+1] WMF: Cap dependency-stack depth to throttle oversized patch series [integration/zuul] (patch-queue/debian/jessie-wikimedia) - 10https://gerrit.wikimedia.org/r/1311501 (owner: 10Jforrester) [19:38:55] brennen: Just checking to see if you released that Gerrit repo from Zuul jail. [19:42:34] 10Gerrit, 06collaboration-services: Rename of a repository is not replicated - https://phabricator.wikimedia.org/T398401#12129968 (10Dzahn) I got the rename replication working!:) with the patch(es) above. But it's in a weird state. As in _it actually does the rename on all 3 servers_ but ALSO still logs the... [20:17:01] dancy: yeah, i did once the queue dropped to 0. [20:17:02] (03open) 10dancy: DNM: Testing [repos/releng/scap] - 10https://gitlab.wikimedia.org/repos/releng/scap/-/merge_requests/1225 [20:17:11] (apologies for not logging.) [20:17:24] thx [20:30:20] (03update) 10dancy: DNM: Testing [repos/releng/scap] - 10https://gitlab.wikimedia.org/repos/releng/scap/-/merge_requests/1225 [20:33:47] (03update) 10dancy: DNM: Testing [repos/releng/scap] - 10https://gitlab.wikimedia.org/repos/releng/scap/-/merge_requests/1225 [20:46:32] 10GitLab (CI & Job Runners), 06Release-Engineering-Team, 07Essential-Work: Buildkit v0.31.2 released - https://phabricator.wikimedia.org/T432360#12130307 (10dancy) 05Openβ†’03Resolved p:05Triageβ†’03High a:03dancy [22:49:42] 10Gerrit, 06Release-Engineering-Team, 10Catalyst (PatchDemo): Creating a Patchdemo wiki for VisualEditor fails on missing ref from gerrit-replica - https://phabricator.wikimedia.org/T432413 (10brennen) 03NEW [22:51:40] ^ kind of wondering if gerrit replication has stopped or is just lagging. [22:53:53] :(( I have been trying to _fix_ replication today. [22:54:14] effectively fixed ssh between servers that did _not_ work before [22:54:22] seeing this now is frustrating [22:56:04] ah, yeah, sorry, just seeing there's been replication-related work happening [22:56:15] also because there is nothing about it that would obviously break it [22:56:32] status before: one server could not ssh to the other. status after: they now can [22:57:36] one part of it was dropping a .ssh config file into the gerrit home. I will revert that one. [22:59:08] hold on.. looking at replication log and so on [23:00:00] there is regular replication and then "replication of renamed projects" in this whole story [23:00:02] no rush from this end. i see a bunch of stuff with what looks like retry counts and "pending" statuses from `ssh -p 29418 gerrit.wikimedia.org replication list --detail` but i don't think i've ever looked at that before so i'm not sure if it's abnormal. [23:00:28] well, let me test what happens if I revert one specific thing [23:00:39] I see the failures too now [23:04:05] brennen: yea. I think it's already replicating again now [23:04:11] as soon as I removed the config file [23:05:13] thanks mutante [23:06:27] this is going on my rapidly expanding list of "stuff tyler would probably have known what was up with immediately if he weren't on sabbatical" [23:06:56] !log gerrit1003/2002/2003: rm /srv/gerrit/.ssh/config and revert gerrit:1311577 which fixed T398401 but caused T432413 - replication is working again [23:07:01] Logged the message at https://wikitech.wikimedia.org/wiki/Release_Engineering/SAL [23:07:01] T398401: Rename of a repository is not replicated - https://phabricator.wikimedia.org/T398401 [23:07:02] T432413: Creating a Patchdemo wiki for VisualEditor fails on missing ref from gerrit-replica - https://phabricator.wikimedia.org/T432413 [23:07:25] fixed one thing - broke another [23:07:49] back to it later then :) [23:07:59] appreciate the quick fix [23:08:00] this was about the rename-plugin [23:08:18] a friend of mine recently was talking about transmissions on old hondas [23:08:23] [2026-07-16 23:08:09,697] Replication to gerrit@gerrit2002.wikimedia.org:/srv/gerrit/git/operations/puppet.git completed in 118065ms, 141080ms delay, 0 retries [CONTEXT pushOneId="89b5f8e6" ] [23:08:51] he was like "at a certain point, it starts to kind of live on its own problems and you don't really want to mess with it" [23:09:23] repos/VisualEditor$ git fetch origin refs/changes/70/1311570/1 master [23:09:31] * branch refs/changes/70/1311570/1 -> FETCH_HEAD [23:09:37] that'll do it. [23:09:55] yea, I freshly cloned it from the replica. confirmed [23:10:23] haha, yea, it is like the Honda [23:11:59] 10Gerrit, 06Release-Engineering-Team, 10Catalyst (PatchDemo), 13Patch-For-Review: Creating a Patchdemo wiki for VisualEditor fails on missing ref from gerrit-replica - https://phabricator.wikimedia.org/T432413#12130856 (10brennen) 05Openβ†’03Resolved a:03Dzahn Looks like this should be fixed now.... [23:12:31] 10Gerrit, 06Release-Engineering-Team, 10Catalyst (PatchDemo), 13Patch-For-Review: Creating a Patchdemo wiki for VisualEditor fails on missing ref from gerrit-replica - https://phabricator.wikimedia.org/T432413#12130860 (10Dzahn) ` 23:08 < mutante> [2026-07-16 23:08:09,697] Replication to gerrit@gerrit2... [23:13:56] brennen: I used to dislike that we are using the replica when we had only 2 servers. I did not like that it effectively made both machines production. but I have stopped argueing like that since we have 3 servers and there is also gerrit-spare.wikimedia.org now [23:15:00] yeah, i guess that makes it less tangled between "this is a fallback machine" and "we can send traffic here to reduce load on the main one" [23:15:36] yea. except in this case.. both would have been the same [23:16:05] different failure modes different story [23:19:29] cu tomorrow. at least we avoided others wondering about all this when Europe wakes up [23:25:46] yeah, for sure [23:25:54] * brennen back to mowing the yard