[00:13:54] PROBLEM - LDAP -writable server- on seaborgium is CRITICAL: Could not bind to the LDAP server https://wikitech.wikimedia.org/wiki/LDAP%23Troubleshooting [00:18:04] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - labweb-ssl_7443: Servers cloudweb1003.wikimedia.org are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [00:18:04] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - labweb-ssl_7443: Servers cloudweb1004.wikimedia.org are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [00:28:36] FIRING: [5x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-3/1/2 (Transport: Hurricane Electric (dc4841.dal5) {#122601}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [00:38:54] RECOVERY - Kafka Broker Server on kafka-test1007 is OK: PROCS OK: 1 process with command name java, args Kafka /etc/kafka/server.properties https://wikitech.wikimedia.org/wiki/Kafka/Administration [00:40:38] FIRING: [13x] CertAlmostExpired: gNMI TLS certificate for lsw1-c2-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [00:40:42] FIRING: [2x] JobUnavailable: Reduced availability for job atlas_exporter in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [00:41:54] RECOVERY - LDAP -writable server- on seaborgium is OK: LDAP OK - 0.009 seconds response time https://wikitech.wikimedia.org/wiki/LDAP%23Troubleshooting [00:41:54] PROBLEM - Kafka Broker Server on kafka-test1007 is CRITICAL: PROCS CRITICAL: 0 processes with command name java, args Kafka /etc/kafka/server.properties https://wikitech.wikimedia.org/wiki/Kafka/Administration [00:42:04] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [00:42:04] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [00:45:42] FIRING: [2x] JobUnavailable: Reduced availability for job atlas_exporter in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [01:07:54] RECOVERY - Kafka Broker Server on kafka-test1009 is OK: PROCS OK: 1 process with command name java, args Kafka /etc/kafka/server.properties https://wikitech.wikimedia.org/wiki/Kafka/Administration [01:10:54] PROBLEM - Kafka Broker Server on kafka-test1009 is CRITICAL: PROCS CRITICAL: 0 processes with command name java, args Kafka /etc/kafka/server.properties https://wikitech.wikimedia.org/wiki/Kafka/Administration [01:12:04] (03PS1) 10TrainBranchBot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1309831 [01:12:04] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1309831 (owner: 10TrainBranchBot) [01:20:19] (03Merged) 10jenkins-bot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1309831 (owner: 10TrainBranchBot) [01:43:04] (03PS1) 10Ladsgroup: thumbor: Allow 2 connection per worker [deployment-charts] - 10https://gerrit.wikimedia.org/r/1309832 (https://phabricator.wikimedia.org/T431767) [02:00:11] !log mwpresync@deploy2003 Started scap build-images: Publishing wmf/next image [02:01:28] !log mwpresync@deploy2003 Finished scap build-images: Publishing wmf/next image (duration: 01m 17s) [02:04:35] FIRING: HelmReleaseBadStatus: Helm release wdqs/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [02:07:19] FIRING: KafkaUnderReplicatedPartitions: Under replicated partitions for Kafka cluster test-eqiad in eqiad - https://wikitech.wikimedia.org/wiki/Kafka/Administration - https://grafana.wikimedia.org/d/000000027/kafka?orgId=1&var-datasource=eqiad%20prometheus/ops&var-kafka_cluster=test-eqiad - https://alerts.wikimedia.org/?q=alertname%3DKafkaUnderReplicatedPartitions [02:08:54] RECOVERY - Kafka Broker Server on kafka-test1007 is OK: PROCS OK: 1 process with command name java, args Kafka /etc/kafka/server.properties https://wikitech.wikimedia.org/wiki/Kafka/Administration [02:10:42] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:11:54] PROBLEM - Kafka Broker Server on kafka-test1007 is CRITICAL: PROCS CRITICAL: 0 processes with command name java, args Kafka /etc/kafka/server.properties https://wikitech.wikimedia.org/wiki/Kafka/Administration [02:15:42] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [03:39:40] FIRING: [3x] SystemdUnitFailed: cowbuilder_update_sid-amd64.service on build2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [04:05:23] FIRING: [13x] CertAlmostExpired: gNMI TLS certificate for lsw1-c2-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [04:08:54] RECOVERY - Kafka Broker Server on kafka-test1007 is OK: PROCS OK: 1 process with command name java, args Kafka /etc/kafka/server.properties https://wikitech.wikimedia.org/wiki/Kafka/Administration [04:11:54] PROBLEM - Kafka Broker Server on kafka-test1007 is CRITICAL: PROCS CRITICAL: 0 processes with command name java, args Kafka /etc/kafka/server.properties https://wikitech.wikimedia.org/wiki/Kafka/Administration [04:28:36] FIRING: [5x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-3/1/2 (Transport: Hurricane Electric (dc4841.dal5) {#122601}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [05:08:54] RECOVERY - Kafka Broker Server on kafka-test1007 is OK: PROCS OK: 1 process with command name java, args Kafka /etc/kafka/server.properties https://wikitech.wikimedia.org/wiki/Kafka/Administration [05:11:54] PROBLEM - Kafka Broker Server on kafka-test1007 is CRITICAL: PROCS CRITICAL: 0 processes with command name java, args Kafka /etc/kafka/server.properties https://wikitech.wikimedia.org/wiki/Kafka/Administration [05:45:42] FIRING: [2x] JobUnavailable: Reduced availability for job atlas_exporter in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [06:04:34] FIRING: HelmReleaseBadStatus: Helm release wdqs/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [06:07:19] FIRING: KafkaUnderReplicatedPartitions: Under replicated partitions for Kafka cluster test-eqiad in eqiad - https://wikitech.wikimedia.org/wiki/Kafka/Administration - https://grafana.wikimedia.org/d/000000027/kafka?orgId=1&var-datasource=eqiad%20prometheus/ops&var-kafka_cluster=test-eqiad - https://alerts.wikimedia.org/?q=alertname%3DKafkaUnderReplicatedPartitions [06:14:24] PROBLEM - Check unit status of httpbb_kubernetes_mw-web-next_hourly on cumin2003 is CRITICAL: CRITICAL: Status of the systemd unit httpbb_kubernetes_mw-web-next_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [06:37:54] RECOVERY - Kafka Broker Server on kafka-test1009 is OK: PROCS OK: 1 process with command name java, args Kafka /etc/kafka/server.properties https://wikitech.wikimedia.org/wiki/Kafka/Administration [06:40:54] PROBLEM - Kafka Broker Server on kafka-test1009 is CRITICAL: PROCS CRITICAL: 0 processes with command name java, args Kafka /etc/kafka/server.properties https://wikitech.wikimedia.org/wiki/Kafka/Administration [06:41:50] (03CR) 10Neriah: "recheck" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1307836 (https://phabricator.wikimedia.org/T431334) (owner: 10Danielyepezgarces) [07:00:05] Deploy window No deploys all day! See Deployments/Emergencies if things are broken. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260712T0700) [07:14:24] RECOVERY - Check unit status of httpbb_kubernetes_mw-web-next_hourly on cumin2003 is OK: OK: Status of the systemd unit httpbb_kubernetes_mw-web-next_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [07:14:24] PROBLEM - Check unit status of httpbb_kubernetes_mw-api-ext-next_hourly on cumin2003 is CRITICAL: CRITICAL: Status of the systemd unit httpbb_kubernetes_mw-api-ext-next_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [07:39:25] RESOLVED: [3x] SystemdUnitFailed: cowbuilder_update_sid-amd64.service on build2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [08:02:45] FIRING: RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95133216 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [08:05:38] FIRING: [12x] CertAlmostExpired: gNMI TLS certificate for lsw1-c2-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [08:07:44] FIRING: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95133216 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [08:14:24] RECOVERY - Check unit status of httpbb_kubernetes_mw-api-ext-next_hourly on cumin2003 is OK: OK: Status of the systemd unit httpbb_kubernetes_mw-api-ext-next_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [08:28:36] FIRING: [5x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-3/1/2 (Transport: Hurricane Electric (dc4841.dal5) {#122601}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [08:37:45] FIRING: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95133216 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [08:42:45] RESOLVED: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95133216 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [09:45:57] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [10:04:35] FIRING: HelmReleaseBadStatus: Helm release wdqs/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [10:07:19] FIRING: KafkaUnderReplicatedPartitions: Under replicated partitions for Kafka cluster test-eqiad in eqiad - https://wikitech.wikimedia.org/wiki/Kafka/Administration - https://grafana.wikimedia.org/d/000000027/kafka?orgId=1&var-datasource=eqiad%20prometheus/ops&var-kafka_cluster=test-eqiad - https://alerts.wikimedia.org/?q=alertname%3DKafkaUnderReplicatedPartitions [12:05:38] FIRING: [12x] CertAlmostExpired: gNMI TLS certificate for lsw1-c2-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [12:28:36] FIRING: [5x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-3/1/2 (Transport: Hurricane Electric (dc4841.dal5) {#122601}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [13:36:19] (03PS3) 10Hashar: ci: add buildx on production Jenkins agents [puppet] - 10https://gerrit.wikimedia.org/r/1308659 (https://phabricator.wikimedia.org/T431582) [13:37:54] RECOVERY - Kafka Broker Server on kafka-test1009 is OK: PROCS OK: 1 process with command name java, args Kafka /etc/kafka/server.properties https://wikitech.wikimedia.org/wiki/Kafka/Administration [13:38:14] (03CR) 10Hashar: "**done**, I have moved `docker-buildx` with the `docker-cli` install which matches how Debian has split Docker in multiple packages." [puppet] - 10https://gerrit.wikimedia.org/r/1308659 (https://phabricator.wikimedia.org/T431582) (owner: 10Hashar) [13:40:54] PROBLEM - Kafka Broker Server on kafka-test1009 is CRITICAL: PROCS CRITICAL: 0 processes with command name java, args Kafka /etc/kafka/server.properties https://wikitech.wikimedia.org/wiki/Kafka/Administration [13:45:57] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [14:04:35] FIRING: HelmReleaseBadStatus: Helm release wdqs/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [14:07:19] FIRING: KafkaUnderReplicatedPartitions: Under replicated partitions for Kafka cluster test-eqiad in eqiad - https://wikitech.wikimedia.org/wiki/Kafka/Administration - https://grafana.wikimedia.org/d/000000027/kafka?orgId=1&var-datasource=eqiad%20prometheus/ops&var-kafka_cluster=test-eqiad - https://alerts.wikimedia.org/?q=alertname%3DKafkaUnderReplicatedPartitions [14:56:44] PROBLEM - Druid historical on an-druid1007 is CRITICAL: PROCS CRITICAL: 0 processes with command name java, args org.apache.druid.cli.Main server historical https://wikitech.wikimedia.org/wiki/Analytics/Systems/Druid [14:56:54] PROBLEM - SSH on an-druid1006 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/SSH/monitoring [14:58:44] RECOVERY - SSH on an-druid1006 is OK: SSH OK - OpenSSH_9.2p1 Debian-2+deb12u7 (protocol 2.0) https://wikitech.wikimedia.org/wiki/SSH/monitoring [14:58:44] PROBLEM - Druid historical on an-druid1006 is CRITICAL: PROCS CRITICAL: 0 processes with command name java, args org.apache.druid.cli.Main server historical https://wikitech.wikimedia.org/wiki/Analytics/Systems/Druid [15:01:44] RECOVERY - Druid historical on an-druid1007 is OK: PROCS OK: 1 process with command name java, args org.apache.druid.cli.Main server historical https://wikitech.wikimedia.org/wiki/Analytics/Systems/Druid [15:07:19] PROBLEM - Host db2209 #page is DOWN: PING CRITICAL - Packet loss = 100% [15:08:44] RECOVERY - Druid historical on an-druid1006 is OK: PROCS OK: 1 process with command name java, args org.apache.druid.cli.Main server historical https://wikitech.wikimedia.org/wiki/Analytics/Systems/Druid [15:09:01] PROBLEM - MariaDB Replica IO: s3 #page on db2177 is CRITICAL: CRITICAL slave_io_state Slave_IO_Running: No, Errno: 2003, Errmsg: error reconnecting to master repl2024@db2209.codfw.wmnet:3306 - retry-time: 60 maximum-retries: 100000 message: Cant connect to server on db2209.codfw.wmnet (110 Connection timed out) https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [15:09:05] PROBLEM - MariaDB Replica IO: s3 #page on db2190 is CRITICAL: CRITICAL slave_io_state Slave_IO_Running: No, Errno: 2003, Errmsg: error reconnecting to master repl2024@db2209.codfw.wmnet:3306 - retry-time: 60 maximum-retries: 100000 message: Cant connect to server on db2209.codfw.wmnet (113 No route to host) https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [15:09:06] PROBLEM - MariaDB Replica IO: s3 #page on db2194 is CRITICAL: CRITICAL slave_io_state Slave_IO_Running: No, Errno: 2003, Errmsg: error reconnecting to master repl2024@db2209.codfw.wmnet:3306 - retry-time: 60 maximum-retries: 100000 message: Cant connect to server on db2209.codfw.wmnet (110 Connection timed out) https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [15:09:09] PROBLEM - MariaDB Replica IO: s3 #page on db2205 is CRITICAL: CRITICAL slave_io_state Slave_IO_Running: No, Errno: 2003, Errmsg: error reconnecting to master repl2024@db2209.codfw.wmnet:3306 - retry-time: 60 maximum-retries: 100000 message: Cant connect to server on db2209.codfw.wmnet (110 Connection timed out) https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [15:09:13] PROBLEM - MariaDB Replica IO: s3 #page on db2227 is CRITICAL: CRITICAL slave_io_state Slave_IO_Running: No, Errno: 2003, Errmsg: error reconnecting to master repl2024@db2209.codfw.wmnet:3306 - retry-time: 60 maximum-retries: 100000 message: Cant connect to server on db2209.codfw.wmnet (110 Connection timed out) https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [15:09:14] PROBLEM - MariaDB Replica IO: s3 on db2239 is CRITICAL: CRITICAL slave_io_state Slave_IO_Running: No, Errno: 2003, Errmsg: error reconnecting to master repl2024@db2209.codfw.wmnet:3306 - retry-time: 60 maximum-retries: 100000 message: Cant connect to server on db2209.codfw.wmnet (110 Connection timed out) https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [15:09:55] PROBLEM - MariaDB Replica IO: s3 #page on db2156 is CRITICAL: CRITICAL slave_io_state Slave_IO_Running: No, Errno: 2003, Errmsg: error reconnecting to master repl2024@db2209.codfw.wmnet:3306 - retry-time: 60 maximum-retries: 100000 message: Cant connect to server on db2209.codfw.wmnet (110 Connection timed out) https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [15:10:59] o/ [15:11:03] looking [15:11:11] RECOVERY - Host db2209 #page is UP: PING OK - Packet loss = 0%, RTA = 31.76 ms [15:11:29] !incidents [15:11:29] 8173 (ACKED) db2177 (paged)/MariaDB Replica IO: s3 (paged) [15:11:29] 8174 (ACKED) db2190 (paged)/MariaDB Replica IO: s3 (paged) [15:11:30] 8175 (ACKED) db2194 (paged)/MariaDB Replica IO: s3 (paged) [15:11:30] 8176 (ACKED) db2205 (paged)/MariaDB Replica IO: s3 (paged) [15:11:30] 8177 (ACKED) db2227 (paged)/MariaDB Replica IO: s3 (paged) [15:11:30] 8178 (ACKED) db2156 (paged)/MariaDB Replica IO: s3 (paged) [15:11:30] 8172 (RESOLVED) Host db2209 (paged) [15:14:08] PROBLEM - mysqld processes on db2209 is CRITICAL: PROCS CRITICAL: 0 processes with command name mysqld https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting [15:14:09] PROBLEM - MariaDB Replica IO: s3 #page on db2209 is CRITICAL: CRITICAL slave_io_state could not connect https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [15:14:09] PROBLEM - MariaDB Events s3 on db2209 is CRITICAL: CRITICAL - Failed to query events: ERROR 2002 (HY000): Cant connect to local server through socket /run/mysqld/mysqld.sock (2) https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Event_Scheduler [15:14:09] PROBLEM - pt-heartbeat-wikimedia process on db2209 is CRITICAL: PROCS CRITICAL: 0 processes with args pt-heartbeat-wikimedia https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23pt-heartbeat [15:14:10] PROBLEM - MariaDB Replica SQL: s3 #page on db2209 is CRITICAL: CRITICAL slave_sql_state could not connect https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [15:14:10] PROBLEM - MariaDB Event Scheduler s3 on db2209 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Event_Scheduler [15:14:10] PROBLEM - MariaDB read only s3 #page on db2209 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Master_comes_back_in_read_only [15:14:55] PROBLEM - MariaDB Replica Lag: s3 #page on db2156 is CRITICAL: CRITICAL slave_sql_lag Replication lag: 607.11 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [15:15:01] PROBLEM - MariaDB Replica Lag: s3 #page on db2177 is CRITICAL: CRITICAL slave_sql_lag Replication lag: 613.10 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [15:15:05] PROBLEM - MariaDB Replica Lag: s3 #page on db2190 is CRITICAL: CRITICAL slave_sql_lag Replication lag: 617.28 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [15:15:06] PROBLEM - MariaDB Replica Lag: s3 #page on db2194 is CRITICAL: CRITICAL slave_sql_lag Replication lag: 617.31 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [15:15:09] PROBLEM - MariaDB Replica Lag: s3 #page on db2205 is CRITICAL: CRITICAL slave_sql_lag Replication lag: 620.32 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [15:15:15] PROBLEM - MariaDB Replica Lag: s3 #page on db2227 is CRITICAL: CRITICAL slave_sql_lag Replication lag: 625.55 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [15:16:20] !incidents [15:16:21] 8173 (ACKED) db2177 (paged)/MariaDB Replica IO: s3 (paged) [15:16:21] 8174 (ACKED) db2190 (paged)/MariaDB Replica IO: s3 (paged) [15:16:21] 8175 (ACKED) db2194 (paged)/MariaDB Replica IO: s3 (paged) [15:16:21] 8176 (ACKED) db2205 (paged)/MariaDB Replica IO: s3 (paged) [15:16:21] 8177 (ACKED) db2227 (paged)/MariaDB Replica IO: s3 (paged) [15:16:22] 8178 (ACKED) db2156 (paged)/MariaDB Replica IO: s3 (paged) [15:16:22] 8179 (ACKED) db2209 (paged)/MariaDB Replica IO: s3 (paged) [15:16:23] 8180 (ACKED) db2209 (paged)/MariaDB Replica SQL: s3 (paged) [15:16:23] 8181 (ACKED) db2209 (paged)/MariaDB read only s3 (paged) [15:16:24] 8182 (ACKED) db2156 (paged)/MariaDB Replica Lag: s3 (paged) [15:16:24] 8183 (ACKED) db2177 (paged)/MariaDB Replica Lag: s3 (paged) [15:16:25] 8184 (ACKED) db2194 (paged)/MariaDB Replica Lag: s3 (paged) [15:16:25] 8185 (ACKED) db2190 (paged)/MariaDB Replica Lag: s3 (paged) [15:16:26] 8186 (ACKED) db2205 (paged)/MariaDB Replica Lag: s3 (paged) [15:16:26] 8187 (UNACKED) db2227 (paged)/MariaDB Replica Lag: s3 (paged) [15:16:27] 8172 (RESOLVED) Host db2209 (paged) [15:16:36] !ack 8187 [15:16:36] 8187 (ACKED) db2227 (paged)/MariaDB Replica Lag: s3 (paged) [15:20:09] PROBLEM - MariaDB Replica Lag: s3 #page on db2209 is CRITICAL: CRITICAL slave_sql_lag could not connect https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [15:31:45] FIRING: CirrusConsumerFetchErrorRate: cirrus_streaming_updater_consumer_search_codfw in codfw (k8s): fetch error rate too high - https://wikitech.wikimedia.org/wiki/Search#Streaming_Updater - https://grafana.wikimedia.org/d/K9x0c4aVk/flink-app?var-datasource=codfw+prometheus%2Fk8s&var-namespace=cirrus-streaming-updater&var-helm_release=consumer-search - https://alerts.wikimedia.org/?q=alertname%3DCirrusConsumerFetchErrorRate [15:45:42] FIRING: [2x] JobUnavailable: Reduced availability for job atlas_exporter in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [15:49:08] RECOVERY - MariaDB Events s3 on db2209 is OK: OK - All 2 events in ops database are ENABLED https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Event_Scheduler [15:49:08] RECOVERY - mysqld processes on db2209 is OK: PROCS OK: 1 process with command name mysqld https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting [15:49:09] RECOVERY - MariaDB Replica SQL: s3 #page on db2209 is OK: OK slave_sql_state Slave_SQL_Running: Yes https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [15:49:09] (03PS1) 10Gerrit maintenance bot: mariadb: Promote db2205 to s3 master [puppet] - 10https://gerrit.wikimedia.org/r/1309861 (https://phabricator.wikimedia.org/T431950) [15:49:11] RECOVERY - MariaDB read only s3 #page on db2209 is OK: Version 10.11.16-MariaDB-log, Uptime 27s, read_only: True, event_scheduler: True, 56.94 QPS, connection latency: 0.025618s, query latency: 0.000790s https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Master_comes_back_in_read_only [15:49:11] RECOVERY - MariaDB Event Scheduler s3 on db2209 is OK: Version 10.11.16-MariaDB-log, Uptime 27s, read_only: True, event_scheduler: True, 56.88 QPS, connection latency: 0.024355s, query latency: 0.001183s https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Event_Scheduler [15:50:01] RECOVERY - MariaDB Replica IO: s3 #page on db2177 is OK: OK slave_io_state Slave_IO_Running: Yes https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [15:50:05] RECOVERY - MariaDB Replica IO: s3 #page on db2190 is OK: OK slave_io_state Slave_IO_Running: Yes https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [15:50:06] RECOVERY - MariaDB Replica IO: s3 #page on db2194 is OK: OK slave_io_state Slave_IO_Running: Yes https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [15:50:09] RECOVERY - MariaDB Replica IO: s3 #page on db2205 is OK: OK slave_io_state Slave_IO_Running: Yes https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [15:50:10] RECOVERY - MariaDB Replica IO: s3 #page on db2209 is OK: OK slave_io_state Slave_IO_Running: Yes https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [15:50:13] RECOVERY - MariaDB Replica IO: s3 #page on db2227 is OK: OK slave_io_state Slave_IO_Running: Yes https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [15:50:16] RECOVERY - MariaDB Replica IO: s3 on db2239 is OK: OK slave_io_state Slave_IO_Running: Yes https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [15:50:42] FIRING: [2x] JobUnavailable: Reduced availability for job atlas_exporter in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [15:50:55] RECOVERY - MariaDB Replica IO: s3 #page on db2156 is OK: OK slave_io_state Slave_IO_Running: Yes https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [15:51:22] !log marostegui@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 T431950 [15:51:25] T431950: Switchover s3 master (db2209 -> db2205) - https://phabricator.wikimedia.org/T431950 [15:51:36] !log marostegui@cumin1003 dbctl commit (dc=all): 'Set db2205 with weight 0 T431950', diff saved to https://phabricator.wikimedia.org/P94790 and previous config saved to /var/cache/conftool/dbconfig/20260712-155135-marostegui.json [15:52:08] RECOVERY - pt-heartbeat-wikimedia process on db2209 is OK: PROCS OK: 1 process with args pt-heartbeat-wikimedia https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23pt-heartbeat [15:52:14] RECOVERY - MariaDB Replica Lag: s3 #page on db2227 is OK: OK slave_sql_lag Replication lag: 0.45 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [15:52:55] RECOVERY - MariaDB Replica Lag: s3 #page on db2156 is OK: OK slave_sql_lag Replication lag: 0.00 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [15:53:00] !incidents [15:53:00] 8183 (ACKED) db2177 (paged)/MariaDB Replica Lag: s3 (paged) [15:53:00] 8184 (ACKED) db2194 (paged)/MariaDB Replica Lag: s3 (paged) [15:53:00] 8185 (ACKED) db2190 (paged)/MariaDB Replica Lag: s3 (paged) [15:53:01] 8186 (ACKED) db2205 (paged)/MariaDB Replica Lag: s3 (paged) [15:53:01] 8188 (ACKED) db2209 (paged)/MariaDB Replica Lag: s3 (paged) [15:53:01] 8189 (ACKED) db2209 - master for s3 - crashed and rebooted in codfw, needs dba to look [15:53:01] 8182 (RESOLVED) db2156 (paged)/MariaDB Replica Lag: s3 (paged) [15:53:02] RECOVERY - MariaDB Replica Lag: s3 #page on db2177 is OK: OK slave_sql_lag Replication lag: 0.13 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [15:53:02] 8187 (RESOLVED) db2227 (paged)/MariaDB Replica Lag: s3 (paged) [15:53:02] 8178 (RESOLVED) db2156 (paged)/MariaDB Replica IO: s3 (paged) [15:53:03] 8177 (RESOLVED) db2227 (paged)/MariaDB Replica IO: s3 (paged) [15:53:03] 8179 (RESOLVED) db2209 (paged)/MariaDB Replica IO: s3 (paged) [15:53:04] 8176 (RESOLVED) db2205 (paged)/MariaDB Replica IO: s3 (paged) [15:53:04] 8175 (RESOLVED) db2194 (paged)/MariaDB Replica IO: s3 (paged) [15:53:05] 8174 (RESOLVED) db2190 (paged)/MariaDB Replica IO: s3 (paged) [15:53:05] 8173 (RESOLVED) db2177 (paged)/MariaDB Replica IO: s3 (paged) [15:53:05] RECOVERY - MariaDB Replica Lag: s3 #page on db2194 is OK: OK slave_sql_lag Replication lag: 0.11 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [15:53:06] RECOVERY - MariaDB Replica Lag: s3 #page on db2190 is OK: OK slave_sql_lag Replication lag: 0.19 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [15:53:06] 8181 (RESOLVED) db2209 (paged)/MariaDB read only s3 (paged) [15:53:06] 8180 (RESOLVED) db2209 (paged)/MariaDB Replica SQL: s3 (paged) [15:53:07] 8172 (RESOLVED) Host db2209 (paged) [15:53:09] RECOVERY - MariaDB Replica Lag: s3 #page on db2205 is OK: OK slave_sql_lag Replication lag: 0.39 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [15:54:09] RECOVERY - MariaDB Replica Lag: s3 #page on db2209 is OK: OK slave_sql_lag Replication lag: 0.19 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [15:54:52] !incidents [15:54:52] 8189 (ACKED) db2209 - master for s3 - crashed and rebooted in codfw, needs dba to look [15:54:52] 8188 (RESOLVED) db2209 (paged)/MariaDB Replica Lag: s3 (paged) [15:54:52] 8186 (RESOLVED) db2205 (paged)/MariaDB Replica Lag: s3 (paged) [15:54:53] 8185 (RESOLVED) db2190 (paged)/MariaDB Replica Lag: s3 (paged) [15:54:53] 8184 (RESOLVED) db2194 (paged)/MariaDB Replica Lag: s3 (paged) [15:54:53] 8183 (RESOLVED) db2177 (paged)/MariaDB Replica Lag: s3 (paged) [15:54:53] 8182 (RESOLVED) db2156 (paged)/MariaDB Replica Lag: s3 (paged) [15:54:54] 8187 (RESOLVED) db2227 (paged)/MariaDB Replica Lag: s3 (paged) [15:54:54] 8178 (RESOLVED) db2156 (paged)/MariaDB Replica IO: s3 (paged) [15:54:55] 8177 (RESOLVED) db2227 (paged)/MariaDB Replica IO: s3 (paged) [15:54:55] 8179 (RESOLVED) db2209 (paged)/MariaDB Replica IO: s3 (paged) [15:54:56] 8176 (RESOLVED) db2205 (paged)/MariaDB Replica IO: s3 (paged) [15:54:56] 8175 (RESOLVED) db2194 (paged)/MariaDB Replica IO: s3 (paged) [15:54:57] 8174 (RESOLVED) db2190 (paged)/MariaDB Replica IO: s3 (paged) [15:54:57] 8173 (RESOLVED) db2177 (paged)/MariaDB Replica IO: s3 (paged) [15:54:58] 8181 (RESOLVED) db2209 (paged)/MariaDB read only s3 (paged) [15:54:58] 8180 (RESOLVED) db2209 (paged)/MariaDB Replica SQL: s3 (paged) [15:54:59] 8172 (RESOLVED) Host db2209 (paged) [15:57:53] (03CR) 10Marostegui: [C:03+2] mariadb: Promote db2205 to s3 master [puppet] - 10https://gerrit.wikimedia.org/r/1309861 (https://phabricator.wikimedia.org/T431950) (owner: 10Gerrit maintenance bot) [15:58:24] !log Starting s3 codfw emergency failover from db2209 to db2205 - T431950 [15:58:27] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [15:58:27] T431950: Switchover s3 master (db2209 -> db2205) - https://phabricator.wikimedia.org/T431950 [15:58:56] !log marostegui@cumin1003 dbctl commit (dc=all): 'Promote db2205 to s3 primary T431950', diff saved to https://phabricator.wikimedia.org/P94791 and previous config saved to /var/cache/conftool/dbconfig/20260712-155853-marostegui.json [16:01:25] !log marostegui@cumin1003 dbctl commit (dc=all): 'Depool db2209 T431950', diff saved to https://phabricator.wikimedia.org/P94792 and previous config saved to /var/cache/conftool/dbconfig/20260712-160124-marostegui.json [16:01:45] RESOLVED: CirrusConsumerFetchErrorRate: cirrus_streaming_updater_consumer_search_codfw in codfw (k8s): fetch error rate too high - https://wikitech.wikimedia.org/wiki/Search#Streaming_Updater - https://grafana.wikimedia.org/d/K9x0c4aVk/flink-app?var-datasource=codfw+prometheus%2Fk8s&var-namespace=cirrus-streaming-updater&var-helm_release=consumer-search - https://alerts.wikimedia.org/?q=alertname%3DCirrusConsumerFetchErrorRate [16:04:32] 10ops-codfw, 06DBA, 06DC-Ops: db2209 firmware upgrade after a crash - https://phabricator.wikimedia.org/T431952 (10Marostegui) 03NEW [16:05:38] FIRING: [12x] CertAlmostExpired: gNMI TLS certificate for lsw1-c2-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [16:05:38] (03PS1) 10Marostegui: db2209: Disable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1309863 (https://phabricator.wikimedia.org/T431952) [16:07:55] RECOVERY - Kafka Broker Server on kafka-test1009 is OK: PROCS OK: 1 process with command name java, args Kafka /etc/kafka/server.properties https://wikitech.wikimedia.org/wiki/Kafka/Administration [16:10:42] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:10:55] PROBLEM - Kafka Broker Server on kafka-test1009 is CRITICAL: PROCS CRITICAL: 0 processes with command name java, args Kafka /etc/kafka/server.properties https://wikitech.wikimedia.org/wiki/Kafka/Administration [16:11:50] (03CR) 10Marostegui: [C:03+2] db2209: Disable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1309863 (https://phabricator.wikimedia.org/T431952) (owner: 10Marostegui) [16:15:42] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:28:36] FIRING: [5x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-3/1/2 (Transport: Hurricane Electric (dc4841.dal5) {#122601}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [16:37:37] (03PS1) 10Priyankar22: thwiki: Change to Wikipedia 25 logo [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309865 (https://phabricator.wikimedia.org/T431094) [16:42:31] (03CR) 10Priyankar22: "Requesting review for the thwiki Wikipedia 25 logo change per T431094." [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309865 (https://phabricator.wikimedia.org/T431094) (owner: 10Priyankar22) [17:44:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 13.19% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [17:50:42] FIRING: [2x] JobUnavailable: Reduced availability for job atlas_exporter in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [17:59:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.04% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [18:04:35] FIRING: HelmReleaseBadStatus: Helm release wdqs/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [18:05:28] (03CR) 10Aklapper: "Please do not add me as a reviewer for no obvious reasons - thanks!" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309865 (https://phabricator.wikimedia.org/T431094) (owner: 10Priyankar22) [18:07:19] FIRING: KafkaUnderReplicatedPartitions: Under replicated partitions for Kafka cluster test-eqiad in eqiad - https://wikitech.wikimedia.org/wiki/Kafka/Administration - https://grafana.wikimedia.org/d/000000027/kafka?orgId=1&var-datasource=eqiad%20prometheus/ops&var-kafka_cluster=test-eqiad - https://alerts.wikimedia.org/?q=alertname%3DKafkaUnderReplicatedPartitions [19:03:51] PROBLEM - Postgres Replication Lag on puppetdb2003 is CRITICAL: POSTGRES_HOT_STANDBY_DELAY CRITICAL: DB puppetdb (host:localhost) 445857432 and 27 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [19:16:51] RECOVERY - Postgres Replication Lag on puppetdb2003 is OK: POSTGRES_HOT_STANDBY_DELAY OK: DB puppetdb (host:localhost) 25328 and 0 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [19:35:51] PROBLEM - Postgres Replication Lag on puppetdb2003 is CRITICAL: POSTGRES_HOT_STANDBY_DELAY CRITICAL: DB puppetdb (host:localhost) 426700536 and 31 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [19:37:51] RECOVERY - Postgres Replication Lag on puppetdb2003 is OK: POSTGRES_HOT_STANDBY_DELAY OK: DB puppetdb (host:localhost) 107976 and 1 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [19:55:51] PROBLEM - Postgres Replication Lag on puppetdb2003 is CRITICAL: POSTGRES_HOT_STANDBY_DELAY CRITICAL: DB puppetdb (host:localhost) 571455816 and 44 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [20:02:51] RECOVERY - Postgres Replication Lag on puppetdb2003 is OK: POSTGRES_HOT_STANDBY_DELAY OK: DB puppetdb (host:localhost) 75816 and 0 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [20:05:38] FIRING: [12x] CertAlmostExpired: gNMI TLS certificate for lsw1-c2-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [20:07:55] RECOVERY - Kafka Broker Server on kafka-test1009 is OK: PROCS OK: 1 process with command name java, args Kafka /etc/kafka/server.properties https://wikitech.wikimedia.org/wiki/Kafka/Administration [20:10:55] PROBLEM - Kafka Broker Server on kafka-test1009 is CRITICAL: PROCS CRITICAL: 0 processes with command name java, args Kafka /etc/kafka/server.properties https://wikitech.wikimedia.org/wiki/Kafka/Administration [20:18:51] PROBLEM - Postgres Replication Lag on puppetdb2003 is CRITICAL: POSTGRES_HOT_STANDBY_DELAY CRITICAL: DB puppetdb (host:localhost) 319720120 and 20 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [20:28:36] FIRING: [5x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-3/1/2 (Transport: Hurricane Electric (dc4841.dal5) {#122601}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [20:30:51] RECOVERY - Postgres Replication Lag on puppetdb2003 is OK: POSTGRES_HOT_STANDBY_DELAY OK: DB puppetdb (host:localhost) 77784 and 1 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [20:34:51] PROBLEM - Postgres Replication Lag on puppetdb2003 is CRITICAL: POSTGRES_HOT_STANDBY_DELAY CRITICAL: DB puppetdb (host:localhost) 264768888 and 5 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [20:35:51] RECOVERY - Postgres Replication Lag on puppetdb2003 is OK: POSTGRES_HOT_STANDBY_DELAY OK: DB puppetdb (host:localhost) 0 and 1 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [20:37:55] RECOVERY - Kafka Broker Server on kafka-test1009 is OK: PROCS OK: 1 process with command name java, args Kafka /etc/kafka/server.properties https://wikitech.wikimedia.org/wiki/Kafka/Administration [20:40:55] PROBLEM - Kafka Broker Server on kafka-test1009 is CRITICAL: PROCS CRITICAL: 0 processes with command name java, args Kafka /etc/kafka/server.properties https://wikitech.wikimedia.org/wiki/Kafka/Administration [20:43:51] PROBLEM - Postgres Replication Lag on puppetdb2003 is CRITICAL: POSTGRES_HOT_STANDBY_DELAY CRITICAL: DB puppetdb (host:localhost) 408486872 and 16 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [20:54:51] RECOVERY - Postgres Replication Lag on puppetdb2003 is OK: POSTGRES_HOT_STANDBY_DELAY OK: DB puppetdb (host:localhost) 93464 and 1 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [21:09:20] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, July 13 UTC morning backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-it" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1309302 (https://phabricator.wikimedia.org/T431765) (owner: 10Danielyepezgarces) [21:10:29] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, July 13 UTC morning backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-it" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1307836 (https://phabricator.wikimedia.org/T431334) (owner: 10Danielyepezgarces) [21:15:51] PROBLEM - Postgres Replication Lag on puppetdb2003 is CRITICAL: POSTGRES_HOT_STANDBY_DELAY CRITICAL: DB puppetdb (host:localhost) 279030808 and 5 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [21:16:51] RECOVERY - Postgres Replication Lag on puppetdb2003 is OK: POSTGRES_HOT_STANDBY_DELAY OK: DB puppetdb (host:localhost) 37256 and 0 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [21:50:57] FIRING: [2x] JobUnavailable: Reduced availability for job atlas_exporter in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [22:04:35] FIRING: HelmReleaseBadStatus: Helm release wdqs/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [22:07:19] FIRING: KafkaUnderReplicatedPartitions: Under replicated partitions for Kafka cluster test-eqiad in eqiad - https://wikitech.wikimedia.org/wiki/Kafka/Administration - https://grafana.wikimedia.org/d/000000027/kafka?orgId=1&var-datasource=eqiad%20prometheus/ops&var-kafka_cluster=test-eqiad - https://alerts.wikimedia.org/?q=alertname%3DKafkaUnderReplicatedPartitions [23:24:36] (03PS3) 10Cathal Mooney: Adjust order BGP_outfilter is applied to routes announced [homer/public] - 10https://gerrit.wikimedia.org/r/1309707 (https://phabricator.wikimedia.org/T431849) [23:42:01] (03PS1) 10TrainBranchBot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1309885 [23:42:01] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1309885 (owner: 10TrainBranchBot) [23:49:59] (03Merged) 10jenkins-bot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1309885 (owner: 10TrainBranchBot)