[07:30:41] greetings [08:02:28] morning! [08:08:47] morning, I'm back! [08:10:59] wb dhinus [08:13:30] welcome back! \o/ [09:10:13] re: T411248 I'm ready to flip the switch on toolsbeta, i.e. dumps_use_nfs_lb: true in hiera, unless there are objections ? [09:10:14] T411248: Plan to make clouddumps more resilient and easier to operate - https://phabricator.wikimedia.org/T411248 [09:10:43] I'll turn it on, run puppet and then do a dumps-nfs failover [09:14:19] dhinus: we also hit paws nfs 100% usage during the weekend, is there any magic wand you know of to free some space? (the cron that populates the files-to-delete fails running, so that's empty mostly) [09:14:40] (I'm du'ing stuff, but takes a lot of time) [09:15:27] dcaro: yes saw that :( no magic wand that I know of, I know a.ndrew was working on that script, let's ask him later today [09:15:40] 👍 [09:16:11] a.ndrew is OOO and back on the 15th btw [09:16:17] or the 16th actually [09:20:05] ack, we might have to check that sooner then. I'm currently catching up with the backlog, I can have a look later today [09:53:58] dhinus: got the du command output back, I can try removing a few freeroot directories to free some space now, not sure how much will last though [10:09:12] also nudging the set of tofu patches starting at https://gitlab.wikimedia.org/repos/cloud/toolforge/tofu-provisioning/-/merge_requests/107 [10:13:35] I'm going to lunch shortly, will take a look afterwards [10:13:44] * dcaro lunch too [10:13:49] cya in a bit [11:00:08] You should run a new pipeline, because the target branch has changed for this merge request. [11:00:42] ^ the most annoying gitlab warning ever. it really can detect that a pipeline could be run but can't automatically trigger it? [11:01:10] or just notice that the new target branch points to the exact same ref as the old one so a new build isn't really needed :P [11:15:53] hmm, my toolsbeta-redis instances failed to schedule due to "Anti-affinity instance group policy was violated" [11:16:04] godog: we should have more than 3 hypervisors where VMs can be scheduled, right? [11:25:05] hrm, manually scheduling a single VM in that group succeeds [11:39:20] yeah, doing the same thing via tofu but one at a time succeeds [11:45:57] agree (re: gitlab) [12:26:53] thilp: did you get around deploying your first patch? if not, I'm free until our daily if you want to pair on it [12:27:41] thanks! I’m seeing Raymond.Ndibe before the daily for that [12:28:13] is it blocking anything? [12:30:13] no no, just wondering as the other day we were not able to get to it [13:25:21] taavi: interesting, which server group is that trying to use? I see toolsbeta-redis server group empty [13:30:19] godog: empty? I see all 3 new redis VMs in there, although the Horizon page took a bit to load [13:30:24] (that's indeed the group) [13:31:06] taavi: ah yes my bad, I didn't realize I had to wait [13:31:31] to answer your question yes there should definitely be enough cloudvirts [13:34:33] yeah, indeed, I noticed [13:38:28] I'll check the logs rq [13:39:17] I quickly didn't find anything more than the error I pasted and "too many retries, giving up" or something similar [13:40:09] so I suspect it's some sort of a race condition where scheduling multiple VMs that have an anti-affinity rule with each other will end up in a loop of both attempting to schedule to the same hypervisor which fails due to the other one trying the same [13:43:36] that seems likely yeah [13:50:48] speaking of logs, I just opened T432012 [13:50:49] T432012: Get rid of neutronclient deprecation warnings - https://phabricator.wikimedia.org/T432012 [14:52:44] thilp: if/when you start working on the logs-api stuff, let me know if you want a quick intro [14:53:17] dhinus: if/when you start working on PAWS, I can give you 2cents also of what I've been doing so far [14:53:27] dcaro: if you’re available now, I’d be happy to? Otherwise let me know and I’ll find some time in your calendar [14:53:45] thilp: +1 works for me [14:54:56] dcaro: thanks, I will probably ping you about PAWS tomorrow... but I see we're at 98% disk right now, do you think we need to do something today? [14:56:10] we might be ok, I'll check in a bit [14:56:12] if you have a quick fix to avoid reaching 100% today, I'll let you do it and we can discuss tomorrow about the next steps [15:15:02] thilp: the admin docs show a nice diagram for logs-api https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Logs_Service [15:33:19] * thilp appreciates it! [15:53:58] godog: if you're still around, tools-redis-10 fails to schedule with the same error even as a standalone instance and after a couple of retries [17:17:52] * dcaro off [17:17:55] cya tomorrow!