[10:24:01] elukey: o/ puppet is failing with 'parameter 'management_config_data' expects a Sensitive[Hash] value, got Undef' on cloudcumins [10:25:28] taavi: o/ yes yes I know I am working on it, IIUC https://gerrit.wikimedia.org/r/c/operations/puppet/+/1307098/ should fix it (just filed it, the previous attempt was not right ) [10:25:35] sorry for the failures [10:25:48] my knowledge of sensitive is clearly not great [10:26:32] elukey: you need sensitivity training is what you're saying? [10:28:42] claime: exactly yes [10:29:09] claime: I need to send more <3 [10:29:21] :D [10:39:20] there is still a little bug, hope to fix it asap [10:57:24] phabricator master went down [10:57:26] I am on it [10:58:09] alright, thx! [10:58:18] great to see a new promoted master go down [10:58:25] why is that not showing u on -operations ? [10:58:30] p [10:59:15] Manuel is the Lucky Luke of DB monitoring, he detects them faster than his shadow [10:59:26] also it could have waited 4min for the end of our oncall shifts [11:00:20] can someone create a task please? [11:00:52] marostegui: on it [11:01:27] the host is going to be back in a moment, but I will switch it to the previous master [11:01:42] so phab is going to stay RO for a few more minutes [11:01:54] I guess I can't create a task then? [11:02:01] XioNoX: true XD [11:02:03] yeah I was about to say :D [11:02:10] etherpad is up though! [11:02:57] moritzm, Raine I see 3 acked pages in https://portal.victorops.com/ui/wikimedia/incidents that will probably page you again at the same time [11:03:19] I guess they didn't auto-resolve? [11:04:49] no new tasks == no new bugs.. not saying *don't* fix it, but.. /j [11:05:47] quick time to spin up a bugzilla instance >.> [11:06:57] taavi: fixed! Sorry for the trouble [11:07:03] we already have this wonderful system called "mediawiki", just create a wiki page with a section per bug [11:07:04] elukey: thanks! [11:09:38] we should be good now [11:09:48] I will work on puppet now, as I did all the changes live on proxies [11:09:58] can someone test phab for me please? [11:11:02] Yes, I just made an edit successfully. No read-only DB message. [11:11:12] thanks [11:13:51] lol :-D ack, thanks XioNoX :-D [11:23:24] marostegui: getting `MariaDB server is running with the --read-only option` again, if that is unexpected [11:23:29] yeah, puppet [11:23:31] I am on it [11:23:35] :) [11:24:09] TheresNoTime: fixed [11:24:20] working, thankyou! [11:28:29] marostegui: are you done with the Puppet followup step for db1228? I need to reboot two puppet servers and would disable Puppet merges for about 10 minutes [11:28:41] moritzm: go for it! [11:28:51] ok,doing that now [11:42:23] done, puppet merges are re-enabled [12:40:58] I am disabling puppet on all cp hosts to rollout a new webrequest-based tagging for x-provenance [13:10:38] I see that work is currently active on T429699 - I thought I'd let you know that I'm seeing spicerack auth errors from the reimage cookbook and I'm pretty sure I'm copy/pasting the password correctly. [13:10:39] T429699: Add the management password from pwstore to the cumin hosts - https://phabricator.wikimedia.org/T429699 [13:11:42] I'm currently trying to reimage an-test-master100[3-4] to bookworm - please feel free to try these out if it helps (cc: elukey) - They're insetup, so fair game. [13:15:05] btullis: I am in the middle of a rollout, will check in 10/15 mins! [13:15:14] the only thing that I did was adding a new config file [13:15:19] nothing more [13:15:33] but it may be the last change that I merged last week [13:18:14] puppet re-enabled on cp hosts, tested on magru and all went fine. The new changes will be picked up as puppet rolls out them [13:18:19] all right back to reimage [13:21:29] so an-test-master1003 is a supermicro host bought a year ago afaics https://netbox.wikimedia.org/dcim/devices/6332/ [13:21:33] so not something super new [13:22:13] the last change that I introduced is https://gerrit.wikimedia.org/r/c/operations/cookbooks/+/1305602/1/cookbooks/sre/hosts/reimage.py [13:23:32] ok found the issue [13:23:38] ssh ADMIN@an-test-master1003.mgmt.eqiad.wmnet doesn't work [13:23:42] root@ works [13:23:58] so provisioning failed to set up the pwstore's mgmt password for ADMIN [13:24:00] * elukey sigh [13:24:53] the task is https://phabricator.wikimedia.org/T393030, maybe provisioning didn't run correctly for those [13:25:00] I'll re-run it for an-test-master1003 [13:29:57] OK, cool. Sorry to add to your workload. I was going to re-run provisioning and omit `--no-users` but the cookbook complained because it was already active in netbox. So I bailed out. [13:30:34] yes yes please ping me if you see these things so I am aware :D [13:33:47] I assume someone is working on `PyBalBGPUnstable lvs sre (pybal 64600 208.80.154.196 eqiad)`, but I don't know who :D can you please let me know? [13:34:23] Raine: that should be from yesterday. did it page again? [13:34:29] yes, the ack expired [13:34:44] the question is why didn't it resolve though [13:34:49] it should have by now [13:35:09] mhm [13:35:27] but it didn't page now right? I didn't get anything [13:35:31] it did [13:35:36] I acked it too fast maybe? :D [13:35:56] no wait my bad, nevermind :) [13:36:01] well it paged in the app, not in irc, interestingly [13:37:22] did it page in the app 24h after it paged originally? [13:38:07] yes [13:39:01] that's the resolution message getting dropped before VO gets it, then [13:39:15] I believe it also happened in the early EU shift today [13:39:59] with the es1039 alert I think? [13:40:01] https://portal.victorops.com/ui/wikimedia/incidents [13:40:07] it says "ACK" here though interestingly [13:40:13] sorry better link https://portal.victorops.com/ui/wikimedia/incident/8123/details [13:40:14] I can just manually resolve the pybal incident in the splunk web UI ? [13:40:17] yes [13:40:27] sukhe: yeah, that means not-resolved [13:40:45] but I wonder why though? [13:40:59] it happens several times a quarter iirc [13:41:09] ok,I've just resolved "PyBalBGPUnstable lvs sre (pybal 64600 208.80.154.196 eqiad)" in the web UI [13:44:07] https://phabricator.wikimedia.org/search/query/qI353NhVV_G4/#R [14:07:28] btullis: you should be unblocked now! [14:16:44] elukey: thank you [14:54:31] wondering - yesterday when the initial PyBal alerts fired they paged. that was understandable, they fired when CR1 was first rebooted, and this wasn't service affecting as CR2 was handling all traffic [14:54:40] the mistake was I had not downtimed time [14:54:55] I did downtime them at that point though - so when the RECOVERY came in they were downtimed [14:55:07] not sure if that might affect the message propagation to victorops? [14:57:29] it definitely could [15:08:49] <_joe_> and this is why, kids, XA transactions are a pain in the arse. [16:55:51] I got an unexpected diff on a netbox-sync cookbook when reimaging. topranks XioNoX am I OK to say go? [16:56:10] btullis: what's the diff? it's best to check [16:56:49] (I am not Cathal or Arzhel obviously but perhaps it was some Traffic stuff we were doing :) [16:56:50] https://usercontent.irccloud-cdn.com/file/MK7AahnB/image.png [16:57:02] oh yeah, this is topranks [16:57:15] https://gerrit.wikimedia.org/r/plugins/gitiles/operations/dns/+/313be1731970f0a802f4dbca8f9ad38bb54d1206%5E%21/#F0 [16:57:18] definitely @ him :D [16:57:25] btullis: yes good to go, though I think I just ok'd it myself [16:57:29] but yes that's fine [16:57:52] sukhe: oh damn I am caught red handed [16:58:05] Cool, I thought so. I just looked away from a reimage, so it might have been there a few minutes. [16:58:30] yep np thanks :) [16:58:31] topranks: ! [16:59:31] All good. My cookbook failed with this anyway, but the reimage is done. `spicerack.reposync.RepoSyncPushError: Error pushing to origin: bitflags 1032: [rejected] (fetch first)` [16:59:42] I guess that shows that you already pushed it. [17:02:51] btullis: yeah, it's odd at first the cookbook wouldn't run for me, there was a lock, but it then proceeded obviously before yours finished [17:03:05] anyway no harm, I'll give it a manual re-run now just to make sure everything has been pulled through [17:03:47] no changes, all good [19:33:21] arnaudb: I opened https://phabricator.wikimedia.org/T431030 for the 5:30am page. [19:49:21] thanks XioNoX !