[08:57:00] o/ dcausse did you already reach out to ML (SREs) regarding access to ML studio machines? I reached out to Sucheta and Guillaume. Sucheta supports the request and asks if there’s anything else they could assist us with. Do you already know the stack required for the highlighter inference service? If so, we could ask ML if they could set it up for us. [08:58:53] pfischer: not yet but based Erik testing it's likely to be a hugging face model [08:59:23] there's a meeting with ML later this afternoon perhaps that'd be a good place to ask [09:00:17] No worries. Yes, good idea, I’ll bring it up in the ML/Research/Search exchange. [09:00:25] thanks! [10:20:19] lunch [12:46:45] o/ [12:54:24] \o [13:06:23] o/ [13:40:14] pfischer: once you have a moment could you approve T432347 & T432347? [13:40:15] T432347: Requesting access to ml-lab-users for dcausse - https://phabricator.wikimedia.org/T432347 [13:40:30] and T432345 I mean [13:40:31] T432345: Requesting access to ml-lab-users for ebernardson - https://phabricator.wikimedia.org/T432345 [13:40:49] sigh made a typo in Erik's name :/ [14:04:27] hmm, two out of three reindex processes failed with "cirrus_reindexer.k8s.KubectlException: release not found in list of pods" [14:04:51] there are a lot of letters, can't blame you :P [14:05:32] do we have to expect that our mwscripts can just disappear and fail without cleanup? [14:11:18] weird... the resource should still be there... [14:11:23] or it was not even started? [14:13:38] not entirely sure, this is from validate_pod_state_consistency when it runs MWScript.pod_name [14:15:06] looks like commonswiki failed in eqiad :S 4 separate times. [14:15:13] wondering if "list" can be lagging a bit and perhaps a actual get might be more accurate [14:15:16] sigh :( [14:15:22] cloudelastic claims to still be waiting for the reindex, but the release is not found in pod list either [14:15:41] it's like (unverified) something cleaered out the active pods [14:16:07] there should be logs hopefully somewhere to confirm it actually started? [14:16:21] it certainly started, it was running for ~6 days [14:16:38] oh ok [14:16:58] the report with --verbose reads logs and reports on state, it was running for some time but now when i try and get a report it doesn't find a pod to read logs for [14:17:18] * ebernhardson sighs [14:21:01] * ebernhardson should also change these logs to print the cluster...kinda annoying having three tmux panes and not sure which is which cluster [14:21:33] Hello ! Just as a reminder, we're having a post mortem on the OpenSearch upgrade, 40' from now. You should have received an invite. [14:21:43] gehel: can't make it, have another mtng [14:22:01] Trey314159: I did not invite you (I expect that you've been less involved in the project), but feel free to join anyway! [14:22:26] damn! I hope that dcausse is joining, otherwise, it might make more sense to delay the post mortem... [14:23:43] gehel: yes I'll be around [14:24:03] gehel: yeah i just did the analysis analysis as usual and made a few analyzer tweaks. Happy to show up if you need me, but I don't think I have much insight. [14:24:40] dcausse: Good! I hope you'll relay whatever Erik has in mind. [14:25:07] Trey314159: I suspect you have better things to do! But we should have a coffee one of these days, we haven't talked for a looong time! [14:30:45] I added a thanos-swift storage panel to the Cirrus Streaming updater dashboard: https://grafana.wikimedia.org/goto/bfsahfc9amb5se?orgId=default [14:31:02] gehel: it has been a while! [14:40:06] Trey314159: that reminds me! While looking into opensearch 3 i found lucene made changes to nori, they removed some POS tags and replaced them with fine-grained tags [14:40:24] i don't entirely even know what those words mean :P But for my test i had to adjust the nori_posfilter bits [14:42:05] ebernhardson: POS is "part of speech"; they had some aggressive filters by default, so I customized the list. Finer-grained tags sound good, but it may take some fiddling to figure out exactly what to filter on. [14:49:58] we have a few alerts firing for SUP in cloudelastic, are they expected? [14:51:14] not expected [14:52:46] something odds indeed [14:54:30] it's recovering, the producer sent very few updates, seems like it buffered many events, weird... [16:00:44] hmm, i wrote a thing that does per-item retries in SUP and let fable do a max-effort code review...it found 15 findings (and the code review skill limits it to finding 15 things :P) [16:01:26] * ebernhardson wonders how many are actual problems...tbd [16:06:39] :) [16:11:20] dinner [16:52:52] * ebernhardson wants some sort of magical cross-host bash history [17:22:30] isn't that fun...trying to get logs and understand why the previous reindexes failed...but kubectl says: kubectl logs job/mw-script.eqiad.sbw1r92b: error: timed out waiting for the condition [17:22:44] * ebernhardson may have the wrong command, my history disappeared :P [19:45:10] sigh...it seems like reindexing. I just started the reindexing bits back up and cirrus alterted on memory pressure in codfw [19:46:26] That's not good [19:49:16] except...morelike traffic is also way up today [19:49:41] although it's down now from earlier peaks [19:50:14] Out of curiosity, what does `mediawiki_CirrusSearch_backend_failures_total{type="memory_issue"}` represent? [19:50:49] inflatador: usually the parent circuit breaker: https://docs.opensearch.org/latest/install-and-configure/configuring-opensearch/circuit-breaker/ [19:51:09] but it could be any of those [19:52:10] ebernhardson ah OK, that makes sense [19:53:53] maybe we consider allocating more memory to the instances again...not entirely sure. Since the proximal cause seems to be "you upgraded the opensearch/lucene version" i'm not sure what an alternate fix would be [19:55:38] we still hold the main cluster at 30G it looks like, so we would have to also pay the no-more compressed pointers tax [19:56:16] boo [19:57:23] lol at claude: I'm noticing some of the web sources are unreliable—they contain AI-generated content with inaccurate details. I need to verify the actual status of JEP 534 [20:01:18] altohuhg, hmm maybe reindexing is not related? Looking at it codfw is almost done and mw@codfw->dnsdisc is the one that alerted. But codfw is almost done and was just finishing reindexing a few tiny indices [20:01:34] maybe it would be better to guess it's related to the related-articles cache apparently getting busted today and having to refill [20:03:13] (as opposed to eqiad and cloudelastic...which are both basically starting over at commonswiki_file) [20:06:34] "they contain AI-generated content with inaccurate details. " - I'm sure you had nothing to do with that, Claude ;P