[00:22:50] After eliminating know-uncuminable hosts, I see 9 total VMs that don't respond to cumin. Now I'm going through to see what's going on with each... [01:29:15] Reading syslogs, the first occurance of 'network is unreachable' happens at different times on different VMs, and doesn't correspond to live migrations. [01:29:32] So something is going on but it's infrequent and apparently not related to migration. [01:30:24] Different kernels though, so not an obvious issue with the OS running on affected VMs [04:15:39] huh...would it make sense to run something daily to identify these and notify us, at least until a pattern is discerned? [04:17:45] yes, although I'm not sure how to automaticlaly distinguish from the outside whether a given host 'fell off the network' or is just broken in some prosaic internal way. So it would likely require by-hand followup to verify which reports are valid. [04:18:02] Current (negative) findings are on https://phabricator.wikimedia.org/T432426 [08:32:19] morning [13:00:55] taavi, my latest tofu challenge: tofu init -upgrade gets me "Could not retrieve the list of available versions for provider hashicorp/openstack" -- this happens on the branch with my changes but doesn't happen on the main branch. Is that a red herring or is it choking because it doesn't know what a openstack_containerinfra_clustertemplate_v1 is? [13:01:43] andrewbogott: your project folder is probably missing the `required_providers {}` block? [13:01:59] * andrewbogott tries [13:02:21] so just symlink providers.tf from the root to the module [13:02:58] yep, that was exactly it. Thank you! [13:53:22] can I have a +1 for T432153 and T431590? [13:53:22] T432153: Quota increase request for project osmit - https://phabricator.wikimedia.org/T432153 [13:53:23] T431590: Quota increase request for project kafka-infrastructure - https://phabricator.wikimedia.org/T431590 [13:55:10] taavi: +1d [13:56:14] ty [13:58:33] I still feel a bit sketched out about the new workflow's "here's a ready command you can paste into a privileged shell to do this", even if it is super convinient [20:06:22] jhathaway, I have a very frivolous request: I'm trying to track VMs that spontaneously fall off the network, and for this reason I'm planning to watch the hosts-down dashboard for the next while. That's easier if there aren't already a bunch of false reports on the dashboard; right now half of the false reports are in the 'puppet-dev' project because puppet is jammed up on the asbestos-* hosts. [20:07:17] Would you be willing to repair or delete those hosts to clean up my dashboard? If that sounds like work then you should say 'no' and I'll just get ack them and remember that I ack'd them :) [20:08:40] andrewbogott: sure I can take care of those! [20:08:46] thank you! [20:10:25] done! [20:14:15] thanks again [22:34:36] a.ndrewbogott: I hope I didn't send you on a snipe hunt there, but thank you for poking to see if there is something weird happening.