[06:59:52] greetings [07:03:11] taavi: hah, I checked and tools-redis server group uses anti-affinity whereas e.g. tools-k8s-worker uses soft-anti-affinity, my understanding is that anti-affinity (not soft) is working as intended, the scheduler picks an host then nova-compute on the host checks server groups and refuses the allocation, rinse and repeat three times which is the maximum scheduler attempts [07:04:05] whereas soft-anti-affinity weights the cloudvirts and is more lenient [10:08:19] I feel like I'm re-discovering this fact: we have a 'tools' project id and name :( https://phabricator.wikimedia.org/P94814 [10:11:10] and why there's "toolsbeta" inside the "tools" one? :D [10:11:51] excellent point! I didn't notice [10:12:04] uhhhhhhh [10:12:23] mmhh os-project-name does a prefix match? wmcs-openstack server group list --os-project-name toolsbe [10:12:28] returns same results [10:12:56] ok nevermind maybe more a confusing behavior rather than my initial hypothesis [10:18:38] odd, sorry for the potential heart attacks [10:19:10] taavi: what do you think re: my soft-anti-affinity theory ? given that we have a mixture of policies I suspect we've been here before [10:25:05] godog: mhm, maybe, although it feels a bit silly to not take that constraint into account before picking a candidate hypervisor :/ [10:29:45] taavi: I agree, I wasn't aware of this mechanism (i.e. server groups are checked after scheduling) and found out from reading https://docs.openstack.org/nova/2026.1/admin/troubleshooting/affinity-policy-violated.html [10:39:10] I tried to deploy PAWS and Ansible failed with "Repository already have a repository named jupyterhub" [10:40:32] trying to workaround it with "helm repo remove" :P [10:40:58] that worked! [10:41:37] folks sorry on the late notice on this one, I'd missed it on the list [10:41:47] off and on again ftw dhinus [10:41:58] we are going to be upgrading the switch in codfw rack B5 later today: https://phabricator.wikimedia.org/T430918 [10:42:04] cloudweb2002-dev is in this rack [10:42:23] do we need to take any specific action here? or can the host survive an interruption? [10:42:34] topranks: ok no problem, yes host can survive no problem [10:42:41] topranks: check the depool info of that host again :P [10:43:16] taavi: ah apologies, many thanks for doing that!! [11:13:49] I see cloudbackup2003 is in rack B7 we have on the list for Thursday, should I ask data persistence about that one? [11:15:09] topranks: that's also ours and the current theory is that it needs no action as well [11:15:29] taavi: ok thanks, I'll make a note on the task [11:29:53] for whatever reason T350687 left the old bullseye instance in toolsbeta (and toolsbeta) only in place.. is there anything more to upgrading to trixie (which is running in toolforge already) than spinning up a new instance with the roles and flipping the web proxy over? [11:29:54] T350687: [harbor] Move harbor data to object storage service - https://phabricator.wikimedia.org/T350687 [11:32:55] afternoon! [11:33:43] Raymond_Ndibe: ^ [11:39:47] taavi: fyi. in case you missed it, there's failed probes for haproxy Service tools-k8s-haproxy-8:443 has failed probes (http_this_tool_does_not_exist_toolforge_org_ip6) [11:49:51] dcaro: yeah, the redis cluster was a bit unhappy during the switchover, now fixed [11:50:31] 👍 [12:31:01] trying to move toolsbeta harbor to a new machine [12:36:35] taavi: Error: unexpected status from HEAD request to https://toolsbeta-harbor.wmcloud.org/v2/toolforge/jobs-api/manifests/0.0.541-dev-mr-351: 502 Bad Gateway [12:36:53] failed a few times in a row, not sure if something is goingon [12:36:53] see my previous message [12:36:59] aahh, ok ohh [12:38:21] dcaro: try now? [12:38:49] 👍 [13:19:20] dhinus: I'm around if you want to check the paws stuff [13:30:48] dcaro: meeting, I'll ping you later [13:32:10] ack, np [13:37:18] taavi: I'm sorry I forget, but thanks for the tagging on the tasks xd [13:38:31] dcaro: re: T432115 JFYI earlier this morning I deployed dumps-nfs lb to toolsbeta and did a worker roll-reboot + a bunch of reboots for standalone hosts, not sure tbh how/if that would affect the mailer [13:38:32] T432115: [jobs-emailer] No emails sent in the last hour - https://phabricator.wikimedia.org/T432115 [13:39:19] godog: ack, I'll add it to the task, but at a first glance does not seem related (unless restarting pods got the emailer stuck somehow, that should not have) [13:39:36] *nod* makes sense [13:40:08] did you do the same on tools? [13:40:16] (it also got stopped there it seems) [13:40:38] no I didn't deploy to tools [13:40:43] ok so definitely not related [13:41:52] okok, yep [13:42:08] wait no, it's only toolsbeta [13:42:19] I though the double alert was one for tools and one for toolsbeta [13:42:31] lol ok downgranding from 'definitely' to 'maybe' [13:45:06] oh, fourohfour has a big spike in 500 errors [13:45:41] ah no, that was before, it went out, my grafana did not refresh [13:45:54] * dcaro needs some coffee [13:49:57] codfw has issues it seems? anyone doing onything there? [13:50:22] T430918 maybe [13:50:23] T430918: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918 [13:52:21] not sure, that matches timing wise but shouldn't have affected our stuff [13:52:37] seems like ~most recovered already [13:52:41] the alert went away yep [13:52:55] agree it seems like a side effect maybe? [14:17:16] if there are no objections I'll roll-reboot cloudwebs in eqiad [14:22:36] guys one thing, prometheus2005 was in that rack. in theory that's fine, the other promehtues nodes scrape endpoints still, but there is a discussion in #mediawiki_security about whether it triggered a load of alerts as soon as it was able to talk to the network again [14:22:41] could be a red heering but fyi [14:22:46] *herring [14:23:04] thanks for the update! [14:50:25] dcaro: about paws, it looks like your cleanup was very effective! can you please write an update in T429578? [14:50:26] T429578: PAWS disk space is running out - https://phabricator.wikimedia.org/T429578 [14:50:48] dhinus: ack [14:50:59] godog: it might be the nfs stuff, it's failing to mount: https://phabricator.wikimedia.org/T432115#12119941 [14:51:41] *sigh* [14:53:10] I forget, is it weld in charge of setting up those mounts in the manifest ? [14:53:32] volume-admission [14:53:39] thank you [14:53:41] (a k8s admission controller) [14:54:47] ok I wonder if "just" creating the directories is enough for now [14:54:53] or rather, telling puppet not to remove them [14:55:44] ok won't be enough probably, as in the new directory also needs to be mounted for the symlink to work (thinking out loud) [15:01:02] dcaro: I'll undo my nfs change now for toolforge worker-nfs and fix puppet + volume-admission properly tomorrow [15:01:23] ack, let me know if you want help or reviews and such [15:02:16] somehow I'd expected that "your tool is failing to start entirely" would be something that the emailer would pick on :D [15:03:56] lol [15:04:10] dcaro: thank you, rollback is only unsetting a hiera variable and running puppet [15:04:20] should be complete shortly [15:04:49] every time I stub my toe on the fact that cumin 'hostname glob' doesn't work as expected on cloudcumin :( [15:04:49] taavi: yep, it kinda hangs in "waiting for pod to be created" sort of state [15:05:09] says 'running for 5h' [15:06:33] dcaro: any better now ? [15:19:32] godog: oh yes, everything started running [15:21:56] dcaro: nice, I see one cronjob pod in Error? with 'kubectl get all' that is [15:22:00] pod/cronjob-29734041-scj77 0/1 Error 0 27s [15:23:15] looking, it comes and goes [15:26:40] I think the issue now is on the job side [15:26:41] `tee: /data/project/sample-complex-app/www/static/cronjob_status.json: No such file or directory ` [15:26:44] wait [15:26:47] that's the home dir [15:26:52] that's not mounted either? [15:27:19] on which host ? [15:28:10] but yeah non-dumps mounts should not be affected at all [15:28:34] toolsbeta-test-k8s-worker-nfs-10 I think [15:29:33] hmm, the dir does not exist `/data/project/sample-complex-app/www/static` but the homedir is mounted [15:29:49] was this working before? [15:29:59] (it should be sending emails anyhow id) [15:30:01] *xd [15:30:16] yep, the alert is gone :facepalm: [15:30:19] heh no idea if it was working before tbh, but yes homes are mounted indeed [15:30:28] ok I'm logging off, see you tomorrow [15:31:06] ack, thanks! [17:47:22] * dcaro off [17:47:25] cya tomorrow!