[08:04:20] good morning! [08:41:19] hi! [10:07:39] lunch [13:19:14] \o [13:19:20] o/ [13:22:25] .o/ [13:41:52] dcausse I won't be able to make pairing today but LMK if I can help w/anything [13:51:35] inflatador: ping [13:51:47] i shut down cluster manager for omega, cirrussearch2092 , and omega is not responding now. need help [13:52:01] i brought the node back without upgrade [13:53:44] :S [13:55:01] ok so looks like 9443 (omega) isn't responding to /_cat/health, so indeed it's not sure about a quorum [13:55:10] no user traffic is here (afaik) so not big rush [13:57:53] it was last master in the cluster, all other were upgraded already [13:58:22] last master at all, or last master on the old version? There should have still been multiple masters [13:58:44] health returns this after waiting `{"error":{"root_cause":[{"type":"cluster_manager_not_discovered_exception","reason":null}],"type":"cluster_manager_not_discovered_exception","reason":null},"status":503}` [13:59:17] this is one of those things where if we had logs flowing into logstash it would help :P [13:59:33] (not directed at you, just general because logs came up recently) [13:59:40] this is a omega cluster status before shutting down the node https://www.irccloud.com/pastebin/HcjdHIQZ/ [14:00:07] hmm, ok so it thinks it had 4 master capable iiuc [14:03:49] back [14:03:53] inflatador: np! [14:05:06] not seeing great answers for omega right now...nodes are not behaving :S [14:06:12] part of it is that opensearch logs are terrible to read...every log line is like 300+ chars [14:10:11] reminds of the master election slow down we had when there was too many banned nodes [14:13:52] from the logs, it's almost like it's electing too-fast [14:16:05] might need to stop a couple of the master capables so there are only 3...maybe [14:16:26] yes [14:17:11] i'm going to stop the ones with the most recent startup times, correct? [14:20:23] https://www.irccloud.com/pastebin/xLwa7OV2/ [14:20:31] atsukoito: yea should be reasonable [14:21:43] one of those perhaps? https://gerrit.wikimedia.org/g/operations/puppet/+/6638d02719961a8207430a501e667099c959a015/hieradata/role/codfw/cirrus/opensearch.yaml#11 [14:23:14] from `tail -n 1000 production-search-omega-codfw.log | grep -oE 'failed to join \{cirrussearch[0-9]+' | sort | uniq -c` on the master capable list the nodes trying to join are all master nodes, it's not even clear who they would be joining :S [14:23:28] hotthreads is not particularly useful mainly shuffling cluster states and index metadata [14:23:33] atsukoito just got back from dropping off kids, looking [14:24:17] atsukoito I'm up in https://meet.google.com/aod-fbxz-joy if you want to join [14:24:50] maybe 2073 and 2086 complaining the most [14:25:05] stopped 2092 and 2106 as those were most recently started [14:25:09] also, the list is https://gerrit.wikimedia.org/g/operations/puppet/+/6638d02719961a8207430a501e667099c959a015/hieradata/role/codfw/cirrus/opensearch.yaml#104 [14:25:38] atsukoito: oops my bad, thanks [14:26:27] i suppose my guess would be to stop 2 nodes, it that doesn't help stop all 5 masters and bring them up one at a time? [14:27:07] i suppose the logs on 2106 are a little weird [14:27:31] on other nodes the log regex `failed to join \{cirrussearch[0-9]+'` matches mostly other nodes, but on 2106 it 99% matched itself [14:27:56] might stop 2106 first? [14:28:18] oh i see you did [14:32:19] seems like it worked, it's back [14:32:49] indeed, it's back at yellow [14:33:08] so far...i have to say we've had more operational problems with opensearch and cluster stability in 1 year than in 10 of elastic... [14:33:48] routing.allocation.exclude is empty so not related to dynamic replica settings... [14:56:57] dcausse: thanks for checking! i was afraid we'll need to pool back codfw if eqiad won't be working [14:57:45] inflatador: could we try removing the voting exclusion to see? [14:58:39] also wondering how to see current exclusions, _cluster/settings has nothing and https://docs.opensearch.org/latest/api-reference/cluster-api/cluster-voting-configuration-exclusions/ only has POST requests... [14:59:01] according to "OpenSearch expert" agent it can be found in cluster state, GET /_cluster/state/metadata [14:59:42] dcausse: 'https://search.svc.codfw.wmnet:9443/_cluster/state?filter_path=metadata.cluster_coordination.voting_config_exclusions' [14:59:51] at least, that looks plausible to me [14:59:51] dcausse do you want me to remove the exclusion now? I was gonna wait until after the reimage. If it will give us some insight though I'm fine with it [15:00:16] thanks! [15:00:36] meh okta wants me to log in again...logged in like 2 days ago [15:01:10] inflatador: well.. just curious to see if the problem is still there or if it was just something related to transitioning the last opensearch 1x masters? [15:03:55] dcausse cool, I can wipe them out now and see what happens then ;) [15:04:10] ^^ atsukoito just so you know, I will be watching closely [15:11:36] atsukoito just ran `curl -XDELETE http://0:9400/_cluster/voting_config_exclusions?wait_for_removal=false` against omega CODFW [15:12:39] inflatador: thanks [15:21:59] NP, looks like it didn't cause any issues [15:38:00] https://phabricator.wikimedia.org/P94669 [15:57:34] There are some search-update pipeline SLO alerts at the moment. Are you aware of them? https://alerts.wikimedia.org/?q=%40state%3Dactive&q=team%3Ddata-platform&q=alertname%3DSLOBudgetBurn [16:33:54] btullis: thanks for bringing it up! with `eqiad`, it is linked to the network outage. with `codfw`, we had quorum issue on `omega` after last cluster manager was switched off for re-imaging. we are back now (copying from slack for visibility) [16:36:28] inflatador: ebernhardson: we lost bunch of machines in codfw that was already provisioned to trixie right about now [16:37:12] atsukoito ACK, I bet that is related to a switch maintenance DC Ops just pinged me about [16:37:44] ref T429861 [16:37:44] T429861: codfw: rack B2 maintenance 2026-07-01 11:00 am CT - https://phabricator.wikimedia.org/T429861 [16:40:10] yup, those 7 went off https://www.irccloud.com/pastebin/d1hT7Dwq/ [16:41:22] For context, we typically allow DC Ops/NetOps to take out an entire row of CirrusSearch without telling us, although that might not be the best idea when we're actively doing our own maintenance ;) [16:43:58] it is alright in this case, we only lost 2 masters in main search (port 9200), which should be okay, clusters are still yellow [16:58:26] i wish openseach could return non-active (historic) nodes in the output.. i don't want to go ask external system to fill out the gaps while the nodes are reimaging [17:08:25] atsukoito that reminds me, I was working on a dashboard at https://grafana.wikimedia.org/d/e7219710-20bb-4007-8af8-c9d6dd7e0b35/cirrussearch-opensearch-ops-dashboard-wip?orgId=1&from=now-7d&to=now&timezone=utc&var-cluster=elasticsearch&var-site=eqiad&var-cluster_shortname=chi&var-exported_cluster=production-search&refresh=30s . I believe there is a way to get at least some of that info from prometheus [17:16:07] Puppet is a good secondary source of truth if you just need to know which nodes are in which cluster, which are masters, etc [17:16:53] * inflatador will update the docs to make this more clear [17:19:45] dinner [17:20:29] inflatador: added you to CC for an active SLOBurn alert T430827 while it is active [17:20:29] T430827: SLOBudgetBurn - https://phabricator.wikimedia.org/T430827 [17:23:01] ebernhardson: actually, we have a few graphs that aren't converging https://w.wiki/Rwed [17:24:42] atsukoito ACK, I am happy to take this one over [17:24:57] thanks! [17:26:24] atsukoito: hmm, looking [17:27:26] and the goal should be to hold at 1.0 on this graph? [17:31:51] ebernhardson I'm not sure myself. ryankemper are you around? We might need some help with ^^ [17:36:14] the grafana dashboard for codfw SUP looks reasonable to me, [17:41:26] unrelated, but cirrussearch1086 got stuck again. I restarted opensearch, but if it happens again I'm gonna ban/unban in the hope that it will lose some of its more popular shards [17:41:34] +1 [17:53:45] inflatador: catching up here [17:53:57] re the slo, it looks ok. the 2h metric has gone back to 0. I don't understand what the flat values are in the 3d graph [17:55:18] it feels like it's perhaps not aggregated across the right columns? Like it shouldn't be possible for the 3-day average to go to 80% in a moments notice. Or i'm misunderstanding what this represents [18:00:16] if task_attempt_num is part of the aggregation (looks like it is) it probably shouldnt be [18:32:16] k8s port-forwarded to flink and hit it with curl, and I see that current codfw consumer-search job_id is `c8255cffdb9e0864a3ab9623b0fdc324`, vertex/task_id for the ES sink chain is `85bc842284628bb3dd502e791460192a`, and the live subtasks are attempt 17. The high 6h/3d SLO series are old attempts `8..16`, so the flat long-window graph must be stale attempt-label series rather than active burn [18:32:34] I think the SLO query should aggregate away (`task_attempt_id`, `task_attempt_num`, `kubernetes_pod_name`, `tm_id`, `host`, `instance`, and probably `job_id`) before the burn-rate evaluation, keeping (`task_id`/`operator_id`/`operator_name`) [18:33:14] tl;dr: im with erik on this one, there's no active issue with the pipeline, but we're doing the aggregation somewhat wrong