[08:25:02] 10netops, 06DC-Ops, 06Infrastructure-Foundations: Consider removing cable IDs from interface descriptions - https://phabricator.wikimedia.org/T432689 (10ayounsi) 03NEW p:05Triage→03Low [08:52:05] 10netops, 06Infrastructure-Foundations, 06SRE, 13Patch-For-Review: Create alerting for saturation on sub-rated interfaces - https://phabricator.wikimedia.org/T374614#12140168 (10ayounsi) Manually tested in ulsfo and that seems to work as expected: {F95014173} [08:56:41] 10netops, 06DC-Ops, 06Infrastructure-Foundations: Consider removing cable IDs from interface descriptions - https://phabricator.wikimedia.org/T432689#12140182 (10ayounsi) [09:16:28] 10netops, 06cloud-services-team, 06Data-Persistence, 06Infrastructure-Foundations, and 6 others: codfw: rack B7 maintenance - Tuesday July 21st 14:00 UTC - https://phabricator.wikimedia.org/T430928#12140218 (10CWilliams-WMF) [09:55:14] 06Traffic, 06ServiceOps new: Receiving traffic from origin without sending it back to users (21st Jul) - https://phabricator.wikimedia.org/T432699#12140538 (10jijiki) [09:56:56] 06Traffic, 06ServiceOps new: Receiving traffic from origin without sending it back to users (21st Jul) - https://phabricator.wikimedia.org/T432699#12140542 (10jijiki) [11:11:17] 10netops, 06cloud-services-team, 06Data-Persistence, 06Infrastructure-Foundations, and 7 others: codfw: rack B7 maintenance - Tuesday July 21st 14:00 UTC - https://phabricator.wikimedia.org/T430928#12140945 (10cmooney) [12:02:00] FIRING: [3x] PurgedHighEventLag: High event process lag with purged on cp3071:2112 - https://wikitech.wikimedia.org/wiki/Purged#Alerts - https://alerts.wikimedia.org/?q=alertname%3DPurgedHighEventLag [12:07:00] FIRING: [5x] PurgedHighEventLag: High event process lag with purged on cp3068:2112 - https://wikitech.wikimedia.org/wiki/Purged#Alerts - https://alerts.wikimedia.org/?q=alertname%3DPurgedHighEventLag [12:07:04] hey folks, for whatever reason doh3005 seems to refuse to resolve WMCS hosted names: https://phabricator.wikimedia.org/P94981 [12:08:32] thanks taavi , can you fill a quick ticket about this so we'll not lose track? [12:13:47] 06Traffic: Wikimedia DNS somehow unable to resolve WMCS hosted domain names - https://phabricator.wikimedia.org/T432720 (10taavi) 03NEW [12:13:48] fabfur: sure, T432720 [12:13:49] T432720: Wikimedia DNS somehow unable to resolve WMCS hosted domain names - https://phabricator.wikimedia.org/T432720 [12:27:00] FIRING: [10x] PurgedHighEventLag: High event process lag with purged on cp3068:2112 - https://wikitech.wikimedia.org/wiki/Purged#Alerts - https://alerts.wikimedia.org/?q=alertname%3DPurgedHighEventLag [12:29:57] 10netops, 06DC-Ops, 06Infrastructure-Foundations: Consider removing cable IDs from interface descriptions - https://phabricator.wikimedia.org/T432689#12141213 (10cmooney) Yeah I think this makes sense. It's not so often we have to check the cable label that it's much hassle to look it up in Netbox. It migh... [12:32:00] RESOLVED: [10x] PurgedHighEventLag: High event process lag with purged on cp3068:2112 - https://wikitech.wikimedia.org/wiki/Purged#Alerts - https://alerts.wikimedia.org/?q=alertname%3DPurgedHighEventLag [12:38:05] 10netops, 06cloud-services-team, 06Data-Persistence, 06Infrastructure-Foundations, and 6 others: codfw: rack B7 maintenance - Tuesday July 21st 14:00 UTC - https://phabricator.wikimedia.org/T430928#12141236 (10cmooney) [13:33:35] 10netops, 06cloud-services-team, 06Data-Persistence, 06Infrastructure-Foundations, and 6 others: codfw: rack B7 maintenance - Tuesday July 21st 14:00 UTC - https://phabricator.wikimedia.org/T430928#12141478 (10cmooney) [13:35:55] 10netops, 06Infrastructure-Foundations, 06SRE: Create alerting for saturation on sub-rated interfaces - https://phabricator.wikimedia.org/T374614#12141492 (10ayounsi) 05Open→03Resolved All done ! [13:37:59] 10netops, 06cloud-services-team, 06Data-Persistence, 06Infrastructure-Foundations, and 6 others: codfw: rack B7 maintenance - Tuesday July 21st 14:00 UTC - https://phabricator.wikimedia.org/T430928#12141501 (10ops-monitoring-bot) Draining ganeti2032.codfw.wmnet of running VMs [13:38:21] sukhe is now a good time to merge https://gerrit.wikimedia.org/r/c/operations/puppet/+/1312517 ? looks like I will have to restart pybal on the low-traffic servers in eqiad [13:38:33] happy to schedule a time if not [13:39:34] inflatador: thanks for checking, go ahead please. start with lvs1020 (backup) and then 1019 (low-traffic) [13:39:57] sukhe ACK, merging/puppet-merging now [13:44:01] just restarted pybal on lvs2020, waiting a bit for for things to settle [13:45:19] 06Traffic: Wikimedia DNS somehow unable to resolve WMCS hosted domain names - https://phabricator.wikimedia.org/T432720#12141567 (10ssingh) Thanks for reporting. It is a bit weird but now if I test it against doh3005, it seems to have resolved. Can you confirm? [13:46:35] 06Traffic: Wikimedia DNS somehow unable to resolve WMCS hosted domain names - https://phabricator.wikimedia.org/T432720#12141571 (10ssingh) Not that we have any logs around queries but at least looking at the SERVFAILs in pdns-rec's journal, there seem to be no matches for this domain, which doesn't help much. [13:48:52] I do see some pybal LVS diff alerts but they are > 30m old, proceeding with the lvs1019 pybal restart [13:49:17] inflatador: looking [13:49:49] dig -x 10.2.2.36 @ns0.wikimedia.org +short [13:49:49] wdqs-scholarly.svc.eqiad.wmnet. [13:50:00] are these expected? [13:50:07] > Services known to PyBal but not to IPVS: set(['10.2.2.36:443']) [13:50:10] maybe that's a hangover from the depool/move-vlan stuff from yesterday? [13:50:40] did we remove any services yesterday though? [13:50:40] No, not expected but we are reimaging all the wdqs hosts ATM [13:50:57] 06Traffic: Wikimedia DNS somehow unable to resolve WMCS hosted domain names - https://phabricator.wikimedia.org/T432720#12141578 (10taavi) Indeed, works again for me. Is there something I could've done when it was broken that would've given you more info to work off of? [13:55:10] sukhe sounds like the next step is to run `ipvsadm --delete-service --tcp-service 10.2.2.71:9200`? I'm getting the IP from https://gerrit.wikimedia.org/r/c/operations/puppet/+/1312517/6/hieradata/common/service.yaml and the port from https://gerrit.wikimedia.org/r/plugins/gitiles/operations/puppet/+/refs/heads/production/hieradata/common/service.yaml#439 [13:55:40] if we need to clean up some of the wdqs stuff first LMK, I think that was caused by the VLAN moves yesterday but not sure how to clean [13:56:09] yes that sounds right. [13:56:16] > ipvsadm --delete-service --tcp-service 10.2.2.71:9200 [13:56:16] this [13:56:24] +1 [13:56:41] we will take a look at the WDQS stuff after this [13:58:01] 10netops, 06cloud-services-team, 06Data-Persistence, 06Infrastructure-Foundations, and 6 others: codfw: rack B7 maintenance - Tuesday July 21st 14:00 UTC - https://phabricator.wikimedia.org/T430928#12141588 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=da979790-60cc-4098-88a0-763d9ca7... [13:59:22] sukhe ACK, i've run the `ipvsadm` cmd on both lvs hosts [14:00:05] 10netops, 06cloud-services-team, 06Data-Persistence, 06Infrastructure-Foundations, and 6 others: codfw: rack B7 maintenance - Tuesday July 21st 14:00 UTC - https://phabricator.wikimedia.org/T430928#12141591 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=e7152add-70e3-4a71-9065-f1276f12... [14:01:06] inflatador: thanks. back to back meetings so I will take a look at the WDQS stuff in ab it [14:02:44] sukhe NP, hit me up if I can help! [14:27:27] 06Traffic, 06Data-Engineering, 06Data-Persistence: Data Lake - mediarequests - adapt HQL query to include counting mediarequests from webrequest_text - https://phabricator.wikimedia.org/T431475#12141725 (10GGoncalves-WMF) [14:28:00] 06Traffic, 06Data-Persistence, 13Patch-For-Review: Move thumbnail caching from upload cluster to text - https://phabricator.wikimedia.org/T427465#12141727 (10GGoncalves-WMF) [14:28:17] 06Traffic, 06Data-Engineering, 06Data-Persistence: Data Lake - mediarequests - adapt HQL query to include counting mediarequests from webrequest_text - https://phabricator.wikimedia.org/T431475#12141729 (10GGoncalves-WMF) [14:28:20] 06Traffic, 06Data-Persistence, 13Patch-For-Review: Move thumbnail caching from upload cluster to text - https://phabricator.wikimedia.org/T427465#12141730 (10GGoncalves-WMF) [14:29:23] 10netops, 06cloud-services-team, 06Data-Persistence, 06Infrastructure-Foundations, and 6 others: codfw: rack B7 maintenance - Tuesday July 21st 14:00 UTC - https://phabricator.wikimedia.org/T430928#12141737 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=988f9bff-f6d2-4af1-9bf4-182d9352... [14:39:21] 10netops, 06cloud-services-team, 06Data-Persistence, 06Infrastructure-Foundations, and 6 others: codfw: rack B7 maintenance - Tuesday July 21st 14:00 UTC - https://phabricator.wikimedia.org/T430928#12141790 (10cmooney) Switch is back up after upgrade and things look good at first glance. [15:38:08] 06Traffic, 06DC-Ops, 10ops-drmrs: hw troubleshooting: Memory failure for cp6008.drmrs.wmnet - https://phabricator.wikimedia.org/T431651#12142126 (10BCornwall) My understanding of the situation: * We don't have any spare DIMMs at the edge, so we're unable to install any new sticks * drmrs will be the first t... [15:51:04] s-ukhe FYI it looks like some of the WDQS hosts weren't starting blazegraph after reimage, per https://sal.toolforge.org/log/a85chZ8BffdvpiTr5mvG this should be fixed now. LMK if that doesn't make the pybal alerts go away [15:52:02] 06Traffic, 10GitLab (Project Migration): Migrate DNS repository from Gerrit to Gitlab - https://phabricator.wikimedia.org/T355906#12142181 (10BCornwall) 05Open→03Declined Heh, yeah, forgot this existed. Declining for now, we're keeping dns-related items on gerrit. Thanks! [15:52:46] inflatador: checking [15:55:41] inflatador: looks good in eqiad, and even codfw for your hosts, thanks [15:55:56] we have a pending alert from some other team in codfw but not yours [16:01:24] sukhe nice, thanks for checkin [17:16:27] 06Traffic, 13Patch-For-Review, 07Sustainability (Incident Followup): Experiment with single backend CDN nodes - https://phabricator.wikimedia.org/T288106#12142775 (10BCornwall) cp3073 (text) is now running on a single drive as of this comment. Will check back in a week to compare. [17:27:35] 10netops, 06cloud-services-team, 06Data-Persistence, 06Infrastructure-Foundations, and 6 others: codfw: rack B7 maintenance - Tuesday July 21st 14:00 UTC - https://phabricator.wikimedia.org/T430928#12142803 (10cmooney) 05Open→03Resolved All seems good, hosts have been repooled. [17:45:18] 06Traffic: Wikimedia DNS somehow unable to resolve WMCS hosted domain names - https://phabricator.wikimedia.org/T432720#12142857 (10ssingh) >>! In T432720#12141578, @taavi wrote: > Indeed, works again for me. Is there something I could've done when it was broken that would've given you more info to work off of?... [17:47:43] 06Traffic: Wikimedia DNS somehow unable to resolve WMCS hosted domain names - https://phabricator.wikimedia.org/T432720#12142862 (10ssingh) `pdns-recursor` on Wikidough hosts is running on port 53, so you can also do a local query to see if it is something with dnsdist, or pdns-rec. `dig @127.0.0.1 ` sho... [18:44:25] FIRING: SystemdUnitFailed: prometheus-varnish-exporter@frontend.service on cp3073:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [18:44:28] FIRING: VarnishPrometheusExporterDown: Varnish Exporter on instance cp3073:9331 is unreachable - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/000000304/varnish-dc-stats?viewPanel=17 - https://alerts.wikimedia.org/?q=alertname%3DVarnishPrometheusExporterDown [18:44:43] brett: ^ that's the new host on single NVME experiment? [18:45:00] 3073 :( [18:45:31] that said, it's not uncommon for this to fire [19:07:11] RESOLVED: VarnishPrometheusExporterDown: Varnish Exporter on instance cp3073:9331 is unreachable - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/000000304/varnish-dc-stats?viewPanel=17 - https://alerts.wikimedia.org/?q=alertname%3DVarnishPrometheusExporterDown [19:09:25] RESOLVED: SystemdUnitFailed: prometheus-varnish-exporter@frontend.service on cp3073:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [23:21:46] 06Traffic, 06Commons: HTTP 429 error on image requests on Commons (non-WMF native apps treated as bot-traffic) - https://phabricator.wikimedia.org/T413570#12143858 (10ssingh) >>! In T413570#12135347, @Nylki wrote: > @ssingh Sorry for another ping, but as bbub, I was wondering if this topic can be tracked somew... [23:32:23] 06Traffic: Provide a way to allow firewall to be bypassed - https://phabricator.wikimedia.org/T432510#12143899 (10ssingh) Thanks for the task. The broader request mentioned is of course more work and requires more discussion but I am hoping we can relax some of the rate-limits around this particular workflow, if... [23:51:04] 06Traffic: Provide a way to allow firewall to be bypassed - https://phabricator.wikimedia.org/T432510#12143933 (10RoySmith) Hmmm. I can not reproduce this right now, but from memory, what I did was went to lens.google.com and then in the box where it says "Past image link", I pasted something like: ` https://...