[02:57:43] FIRING: KeyholderUnarmed: 1 unarmed Keyholder key(s) on netmon2002:9100 - https://wikitech.wikimedia.org/wiki/Keyholder - TODO - https://alerts.wikimedia.org/?q=alertname%3DKeyholderUnarmed [06:57:43] FIRING: KeyholderUnarmed: 1 unarmed Keyholder key(s) on netmon2002:9100 - https://wikitech.wikimedia.org/wiki/Keyholder - TODO - https://alerts.wikimedia.org/?q=alertname%3DKeyholderUnarmed [08:39:05] Hi folks, is puppet struggling? I'm seeing puppet runs failing with timeouts to puppetserver1002... [08:39:17] e.g 'Could not evaluate: Could not retrieve file metadata for puppet:///modules/puppet/facter.conf: Request to https://puppetserver1002.eqiad.wmnet:8140/puppet/v3/file_metadata/modules/puppet/facter.conf?links=manage&checksum_type=sha256&source_permissions=ignore&environment=production timed out connect operation after 60.062 seconds' [08:42:01] And https://puppetboard.wikimedia.org/nodes?status=failed is quite a long list, many of which seem to have similar issues [09:46:37] 10netops, 06Infrastructure-Foundations, 06SRE: Don't announce OSPF routes in unicast BGP on Nokia SR-Linux - https://phabricator.wikimedia.org/T423430#12108326 (10cmooney) I was able to test this in containerlab, which then gave me the confidence to try in prod and it's simple enough. Resulting policy is si... [10:57:43] FIRING: KeyholderUnarmed: 1 unarmed Keyholder key(s) on netmon2002:9100 - https://wikitech.wikimedia.org/wiki/Keyholder - TODO - https://alerts.wikimedia.org/?q=alertname%3DKeyholderUnarmed [11:14:53] I armed the keyholder for rancid on netmon2002 [11:15:39] Emperor: that is indeed a long list of failures [11:16:36] sorry I only seen your message, the grown ups are off today unfortunately, I'll see if I can spot anything, or maybe Jsse will have some ideas when he gets online [11:17:28] RESOLVED: KeyholderUnarmed: 1 unarmed Keyholder key(s) on netmon2002:9100 - https://wikitech.wikimedia.org/wiki/Keyholder - TODO - https://alerts.wikimedia.org/?q=alertname%3DKeyholderUnarmed [11:18:54] seems to be intermittent failures, so yeah does suggest some resource exhaustion or similar [11:21:45] topranks: I spoke to s.obanski who said he'd ask j.hathaway to have a look once he's online [11:22:23] yeah it doesn't seem bad enough to warrant waking/paging him right now [11:22:58] I'll take a look anyway, it's my team and I'm on call, even if I'm not a puppet expert [11:23:07] doubt it's the network but I can validate that [11:41:42] load-average on puppetserver1002 did go up a little from ~03:25 UTC [11:42:03] something happened at that stage there is a burst of traffic on all of them. but overall the basic health stuff seems ok across the fleet [11:46:49] no signs of packet loss or similar either [12:02:48] FIRING: PuppetFailure: Puppet has failed on ping1004:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [12:07:48] RESOLVED: PuppetFailure: Puppet has failed on ping1004:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [13:16:29] not the first time that puppetserver1002 has gotten some weird tummyache iirc [13:17:55] https://phabricator.wikimedia.org/T373527 [13:18:06] okay that's an older task than i remembered [13:26:49] FIRING: PuppetFailure: Puppet has failed on ldap-maint1001:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [13:51:00] o/, i'll take a look at the puppetserver woes [14:01:49] RESOLVED: PuppetFailure: Puppet has failed on ldap-maint1001:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [14:50:01] 10netops, 10Cloud-VPS, 06Infrastructure-Foundations, 06SRE, 06tools-infrastructure-team: Upgrade cloudsw1-e4-eqiad - https://phabricator.wikimedia.org/T429013#12109624 (10fgiunchedi) >>! In T429013#12109258, @cmooney wrote: > Thanks @fgiunchedi. > > This isn't so urgent we want to cause stress for you g... [14:54:56] 10netops, 06Infrastructure-Foundations, 06SRE: Edge BGP: Change order we apply export policies - https://phabricator.wikimedia.org/T431849 (10cmooney) 03NEW p:05Triage→03Low [15:18:25] FIRING: SystemdUnitFailed: requestctl-credential-refresh.service on puppetserver2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [16:18:25] RESOLVED: SystemdUnitFailed: requestctl-credential-refresh.service on puppetserver2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed