[07:45:06] 10netops, 06Traffic, 10Cloud-VPS, 06Data-Platform-SRE, and 4 others: Plan to make clouddumps more resilient and easier to operate - https://phabricator.wikimedia.org/T411248#12118080 (10fgiunchedi) I forgot to attach a task number, however: 07:44 -stashbot:#wikimedia-cloud- Logged the message at https://w... [08:38:35] Good morning! [08:39:06] hello [08:40:15] fabfur: I'm planning to roll out a few PyBal changes, https://gerrit.wikimedia.org/r/1310129, on `high-traffic2`in eqiad and `low-traffic` in eqiad and codfw, starting in about 30 minutes. Is is a good time [08:40:35] yes for me! [08:41:07] just alert oncallers right before [08:41:39] is there a way to figure out who's oncall? [08:42:27] victorops? [08:42:35] checkout the topic on -sre-private [08:42:57] or write directly on -sre-private like "oncallers: beware I'm doing...." [08:43:10] I usually do this way without mentioning explicitly people [08:45:44] thanks! [09:00:37] 10netops, 06Traffic, 10Cloud-Services, 06collaboration-services, and 7 others: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918#12118300 (10CWilliams-WMF) [09:05:17] 10netops, 06Traffic, 10Cloud-VPS, 06Data-Platform-SRE, and 4 others: Plan to make clouddumps more resilient and easier to operate - https://phabricator.wikimedia.org/T411248#12118309 (10fgiunchedi) As of now `dumps_use_nfs_lb: true` is deployed to toolsbeta, unfortunately for k8s workers that means a roll-... [09:05:29] 10netops, 06Traffic, 10Cloud-Services, 06collaboration-services, and 7 others: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918#12118310 (10CWilliams-WMF) [09:50:46] fabfur: starting rollout for https://gerrit.wikimedia.org/r/1310129, https://phabricator.wikimedia.org/T431538 [09:50:55] ack [09:51:25] Rollout plan: lvs1020 -> 1018 -> 1019 -> 2014 -> 2013 https://www.irccloud.com/pastebin/kY77mDOf/ [10:00:32] healthcheck has failed, reverting [10:09:13] fabfur: restarted it back, pools has returned, `ip addr` and `ipvsadm -L -n` has same content [10:09:23] ack [10:09:37] find out why the healthchecks were failing? [10:09:49] yeah, my log is full of openssl errors [10:10:13] `Jul 14 10:00:04 lvs1020 pybal[3315535]: OpenSSL.SSL.Error: [('SSL routines', 'ssl3_get_record', 'wrong version number')]` [10:11:56] is the backend ready to answer to requests with the vip name? [10:12:42] going to check it, it seems like cloudelastic and cirrussearch answers differently [10:16:07] i found the mistake, there was "http://localhost/" check that I replaced with https, fatfingered it [10:19:48] :) [10:24:16] So yeah, only single service was affected, and it was secondary [10:27:43] fabfur: gonna roll out second attempt after lunch, https://gerrit.wikimedia.org/r/1310535 [10:27:49] ack [11:08:21] 10netops, 06Traffic, 10Cloud-Services, 06collaboration-services, and 7 others: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918#12118704 (10cmooney) [11:09:00] 10netops, 06Traffic, 10Cloud-VPS, 06collaboration-services, and 8 others: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918#12118705 (10taavi) [11:42:22] fabfur: rolling out https://gerrit.wikimedia.org/r/1310535, take two, starting from lvs1020 [11:44:12] atsukoito: ok (lunch) [11:44:26] Bon appetite! [11:45:38] 10netops, 06Traffic, 10Cloud-VPS, 06collaboration-services, and 8 others: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918#12118824 (10cmooney) [11:49:01] 10netops, 06Traffic, 10Cloud-VPS, 06collaboration-services, and 8 others: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918#12118846 (10BTullis) I have fully drained dse-k8s-worker2003 and depooled cirrussearch2110. Good to go from our side. [11:50:14] restarted lvs1020, `ip addr` and `ipvsadm -L -n` are same [12:01:26] restarted lvs1018, `ip addr` is the same, `ipvsadm -L -n` peers are the same [12:10:09] restarted lvs1019, `ip addr` is the same, `ipvsadm -L -n` peers are the same [12:16:44] restarted lvs2014, `ip addr` is the same, `ipvsadm -L -n` almost are the same [12:23:42] restarted lvs2013, `ip addr` is the same, `ipvsadm -L -n` almost are the same [12:24:01] fabfur: i'm finished with restarts [12:55:39] 10netops, 06Traffic, 10Cloud-VPS, 06collaboration-services, and 8 others: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918#12119177 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=8b130fc6-8cf0-455b-b0d3-cfde4861384c) set by cmooney@cumin1003 for 2:00:00 on 5 host(... [12:57:33] 10netops, 06Traffic, 10Cloud-VPS, 06collaboration-services, and 8 others: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918#12119202 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=bdeaaae5-a950-451d-a6e0-365a469fd764) set by cmooney@cumin1003 for 2:00:00 on 30 host... [13:11:01] 06Traffic, 10MediaWiki-Core-AuthManager, 10MediaWiki-extensions-OAuth, 06Product Safety and Integrity, and 5 others: Support platform attestation (App Attest / Play Integrity) for requests from the official mobile apps - https://phabricator.wikimedia.org/T431910#12119282 (10ssingh) [13:27:25] FIRING: [3x] SystemdUnitFailed: haproxy.service on cp7001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:28:41] reboot [13:42:43] 06Traffic, 06cloud-services-team, 06collaboration-services, 10Infrastructure Security, and 7 others: Reboot cirrussearch in CODFW - https://phabricator.wikimedia.org/T432119 (10atsuko) 03NEW [13:43:13] 06Traffic, 06cloud-services-team, 06collaboration-services, 10Infrastructure Security, and 7 others: Reboot cirrussearch in CODFW - https://phabricator.wikimedia.org/T432119#12119478 (10atsuko) [13:49:09] FIRING: [3x] HaproxyKafkaExporterDown: HaproxyKafka on cp2049 is down - https://wikitech.wikimedia.org/wiki/HAProxyKafka#HaproxyKafkaExporterDown - https://alerts.wikimedia.org/?q=alertname%3DHaproxyKafkaExporterDown [13:49:18] ^ this is depooled [13:50:43] FIRING: [7x] HaproxyKafkaExporterDown: HaproxyKafka on cp2043 is down - https://wikitech.wikimedia.org/wiki/HAProxyKafka#HaproxyKafkaExporterDown - https://alerts.wikimedia.org/?q=alertname%3DHaproxyKafkaExporterDown [13:51:10] FIRING: [15x] VarnishPrometheusExporterDown: Varnish Exporter on instance cp2043:9331 is unreachable - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/000000304/varnish-dc-stats?viewPanel=17 - https://alerts.wikimedia.org/?q=alertname%3DVarnishPrometheusExporterDown [13:51:51] 10netops, 06Traffic, 10Cloud-VPS, 06collaboration-services, and 8 others: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918#12119568 (10cmooney) Upgrade is now complete, all looks healthy. Thanks everyone for their assistance. [13:53:41] RESOLVED: [15x] VarnishPrometheusExporterDown: Varnish Exporter on instance cp2043:9331 is unreachable - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/000000304/varnish-dc-stats?viewPanel=17 - https://alerts.wikimedia.org/?q=alertname%3DVarnishPrometheusExporterDown [13:54:08] RESOLVED: [8x] HaproxyKafkaExporterDown: HaproxyKafka on cp2043 is down - https://wikitech.wikimedia.org/wiki/HAProxyKafka#HaproxyKafkaExporterDown - https://alerts.wikimedia.org/?q=alertname%3DHaproxyKafkaExporterDown [14:08:50] sukhe: we've a bunch of BFD not established alerts for doh/durum hosts in eqiad, does that make sense? [14:09:27] soryr actually codfw and eqiad [14:16:11] topranks: yep reboots [14:16:21] they should clear up soon [14:16:29] ok np thanks! [14:16:33] I will check again in another hour or so if they persist [14:18:34] thanks [14:23:53] fwiw they should come right back up after a reboot... but we've seen non-standard bird behaviour (some sort of a bug) with dealing with old BFD sessions before [14:24:14] bird should come back up with no knowledge of previous state, and tell the switch/cr when it sees its BFD packet with an existing session ID [14:24:31] instead it seems to echo that back, and then still try to establish a new one. I can't remember the details fully [14:24:40] but bfd should be up as soon as Bird is after reboot [14:24:48] if it's not we likely hit some issue [14:25:12] it's funny all of these are for IPv6 as well, I wonder if that is a factor [14:25:36] topranks: meeting, will read shortly! [14:32:48] FIRING: PuppetFailure: Puppet has failed on cp7001:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [14:32:49] FIRING: PuppetFailure: Puppet has failed on cp7009:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [15:12:21] 10netops, 06Infrastructure-Foundations, 10ops-drmrs, 06SRE: Move TATA IP Transit circuit 1243318 cross-connect from rack B12 to rack B13 - https://phabricator.wikimedia.org/T431605#12120098 (10cmooney) 05Open→03Declined As it turns out dc-ops report that the panel in B13 is also full. So we have o... [15:17:32] 06Traffic, 06collaboration-services, 06Data-Persistence, 06Infrastructure-Foundations, and 2 others: codfw: rack B8 maintenance - https://phabricator.wikimedia.org/T430929#12120108 (10MatthewVernon) So @Eevans has kindly agreed to manage this in my absence - apus-be2004 will need putting into maintenance m... [15:18:16] 06Traffic, 06collaboration-services, 06Data-Persistence, 06Infrastructure-Foundations, and 2 others: codfw: rack B8 maintenance - https://phabricator.wikimedia.org/T430929#12120135 (10MatthewVernon) [16:15:29] FIRING: HAProxyRestarted: HAProxy server restarted on cp4039:9100 - https://wikitech.wikimedia.org/wiki/HAProxy#HAProxy_for_edge_caching - https://grafana.wikimedia.org/d/gQblbjtnk/haproxy-drilldown?orgId=1&var-site=ulsfo%20prometheus/ops&var-instance=cp4039&viewPanel=10 - https://alerts.wikimedia.org/?q=alertname%3DHAProxyRestarted [16:15:40] yeah that's me [16:15:42] it's depooled [16:20:29] RESOLVED: HAProxyRestarted: HAProxy server restarted on cp4039:9100 - https://wikitech.wikimedia.org/wiki/HAProxy#HAProxy_for_edge_caching - https://grafana.wikimedia.org/d/gQblbjtnk/haproxy-drilldown?orgId=1&var-site=ulsfo%20prometheus/ops&var-instance=cp4039&viewPanel=10 - https://alerts.wikimedia.org/?q=alertname%3DHAProxyRestarted [16:22:25] FIRING: [3x] SystemdUnitFailed: haproxy.service on cp4039:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [16:47:25] FIRING: [3x] SystemdUnitFailed: haproxy.service on cp4039:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [16:52:25] FIRING: [3x] SystemdUnitFailed: haproxy.service on cp4039:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [16:52:48] RESOLVED: PuppetFailure: Puppet has failed on cp7001:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [17:22:25] RESOLVED: SystemdUnitFailed: haproxy.service on cp7009:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [17:22:48] RESOLVED: PuppetFailure: Puppet has failed on cp7009:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [18:42:02] 06Traffic, 06DC-Ops, 10ops-drmrs: hw troubleshooting: Memory failure for cp6008.drmrs.wmnet - https://phabricator.wikimedia.org/T431651#12121189 (10ssingh) Hi DC-Ops. Just following up on this in case it was missed, thanks! This is not urgent but will be good to know how long it may take. [18:43:14] 06Traffic, 06Data-Persistence, 13Patch-For-Review: Move thumbnail caching from upload cluster to text - https://phabricator.wikimedia.org/T427465#12121195 (10BCornwall) 05Open→03In progress p:05Triage→03High