[11:58:15] nemo-yiannis: so sorry I broke rt-testing <3 [11:58:28] no worries [14:08:20] Hello. I wonder if anyone would please be able to have a look at this new entry to the service catalog please, if you have time: https://gerrit.wikimedia.org/r/c/operations/puppet/+/1311036 [14:16:08] btullis: I will take a look though I will cc traffic as they are the owners there [14:23:49] effie: Many thanks indeed. [18:45:50] hey folks. was going through some unrelated alerts and saw this [18:45:52] > k8s-ingress-staging_30443: Servers kubestage2003.codfw.wmnet, kubestage2004.codfw.wmnet are marked down but pooled [18:45:55] expected? [18:46:02] this is pybal complaining of course [19:02:04] sukhe: can you see how far back that's been firing? if it's a week or so, I wonder if it might be related to the work in https://phabricator.wikimedia.org/T427401 [19:02:16] swfrench-wmf: yeah, 5 days or so [19:02:30] https://alerts.wikimedia.org/?q=alertname%3DPyBal%20backends%20health%20check&q=team%3Dsre&q=%40receiver%3Dirc-spam [19:02:45] "irc-spam" wonder what that means but yeah [19:05:49] cool, thanks for flagging! yeah, it seems that k8s ingress in codfw staging (typically where the "more disruptive" forms of prep work happens ahead of wider upgrades) is completely down [19:06:14] I'll follow up on the task tracking the istio upgrade [19:06:23] swfrench-wmf: thanks! [19:07:01] no problem! FYI, it *might* take until j.ayme is back to get it resolved [19:07:49] no worries, we should just mark the host not pooled I guess. that will prevent pybal from complaining. [19:08:23] I think the problem is that _none_ of the 4 backends work [19:08:28] I see [19:08:47] so unless we make the depool threshold 100%(?) we'll get some form of alert =/ [19:08:57] yeah... if all of them are down :> [19:09:03] ok, let's not worry too much I guess, as long as it's known [19:09:04] thanks!