[06:38:15] 10netops, 06Traffic, 10Cloud-VPS, 06Data-Platform-SRE, and 4 others: Plan to make clouddumps more resilient and easier to operate - https://phabricator.wikimedia.org/T411248#12122473 (10fgiunchedi) A side effect of the toolsbeta deployment from yesterday was {T432115} i.e. pods getting stuck while starting... [11:05:40] 10netops, 06Infrastructure-Foundations, 06SRE, 13Patch-For-Review: Don't announce OSPF routes in unicast BGP on Nokia SR-Linux - https://phabricator.wikimedia.org/T423430#12123447 (10cmooney) Ok all merged and looking good. We now have no hidden routes in the table on the codfw spines ` cmooney@ssw1-e... [11:05:44] 10netops, 06Infrastructure-Foundations, 06SRE, 13Patch-For-Review: Don't announce OSPF routes in unicast BGP on Nokia SR-Linux - https://phabricator.wikimedia.org/T423430#12123450 (10cmooney) 05Open→03Resolved [11:10:45] 10netops, 06Infrastructure-Foundations, 06SRE, 13Patch-For-Review: Don't announce OSPF routes in unicast BGP on Nokia SR-Linux - https://phabricator.wikimedia.org/T423430#12123477 (10cmooney) 05Resolved→03Open [11:38:52] ello - we'd like to migrate some of the outstanding TCP checks away from icinga and into alertmanager. ncredir needs the same treatment - migration is quite simple though if someone has a second to review https://gerrit.wikimedia.org/r/c/operations/puppet/+/1307182 [11:39:58] er wrong review, oops https://gerrit.wikimedia.org/r/c/operations/puppet/+/1307419 [11:52:05] 06Traffic, 06cloud-services-team, 10Infrastructure Security, 06Infrastructure-Foundations, and 5 others: July 2026 Bookworm/Trixie reboots: Data Platform SRE - https://phabricator.wikimedia.org/T431826#12123705 (10BTullis) [13:02:10] 10netops, 06Traffic, 10Cloud-VPS, 06collaboration-services, and 8 others: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918#12124086 (10cmooney) 05Open→03Resolved [13:17:54] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-eqiad, 06SRE: Install new MPC10E-10C line cards on cr1-eqiad and cr2-eqiad slot 0. - https://phabricator.wikimedia.org/T426343#12124189 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=121d96b4-7ab1-4cc6-9b24-75a50e15a132) set by... [13:19:32] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-eqiad, 06SRE: Install new MPC10E-10C line cards on cr1-eqiad and cr2-eqiad slot 0. - https://phabricator.wikimedia.org/T426343#12124196 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=84179922-a401-44a7-95de-38b063d38fb1) set by... [13:24:20] FIRING: DnsboxServiceMismatch: Service ntp-a state mismatch on dns1004:9100 - https://wikitech.wikimedia.org/wiki/DNS#DnsboxServiceMismatch - https://grafana.wikimedia.org/d/96fb573c-0f3c-456a-886c-e50c29f3ed48/dns-box-service-state?var-site=eqiad&var-instance=dns1004:9100 - https://alerts.wikimedia.org/?q=alertname%3DDnsboxServiceMismatch [13:25:01] reboots, not a worry [13:29:20] RESOLVED: DnsboxServiceMismatch: Service ntp-a state mismatch on dns1004:9100 - https://wikitech.wikimedia.org/wiki/DNS#DnsboxServiceMismatch - https://grafana.wikimedia.org/d/96fb573c-0f3c-456a-886c-e50c29f3ed48/dns-box-service-state?var-site=eqiad&var-instance=dns1004:9100 - https://alerts.wikimedia.org/?q=alertname%3DDnsboxServiceMismatch [14:00:20] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-eqiad, 06SRE: Install new MPC10E-10C line cards on cr1-eqiad and cr2-eqiad slot 0. - https://phabricator.wikimedia.org/T426343#12124411 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=7bebb0cd-fd97-4261-8062-22ac9765f376) set by... [14:16:50] FIRING: [4x] DnsboxServiceMismatch: Service authdns-ns0 state mismatch on dns1006:9100 - https://wikitech.wikimedia.org/wiki/DNS#DnsboxServiceMismatch - https://grafana.wikimedia.org/d/96fb573c-0f3c-456a-886c-e50c29f3ed48/dns-box-service-state?var-site=eqiad&var-instance=dns1006:9100 - https://alerts.wikimedia.org/?q=alertname%3DDnsboxServiceMismatch [14:21:50] RESOLVED: [4x] DnsboxServiceMismatch: Service authdns-ns0 state mismatch on dns1006:9100 - https://wikitech.wikimedia.org/wiki/DNS#DnsboxServiceMismatch - https://grafana.wikimedia.org/d/96fb573c-0f3c-456a-886c-e50c29f3ed48/dns-box-service-state?var-site=eqiad&var-instance=dns1006:9100 - https://alerts.wikimedia.org/?q=alertname%3DDnsboxServiceMismatch [14:24:18] sukhe: two traffic servers are not sending prefixes over BGP to the network - https://grafana-rw.wikimedia.org/goto/efs6wahi9vv28a?orgId=default durum6002 (for v6 t oasw1-b13-drmrs) and dns7002 (v4 and v6) expected? [14:26:42] 06Traffic, 06cloud-services-team, 10Infrastructure Security, 06Infrastructure-Foundations, and 5 others: July 2026 Bookworm/Trixie reboots: Data Platform SRE - https://phabricator.wikimedia.org/T431826#12124569 (10brouberol) [14:31:08] 06Traffic, 06cloud-services-team, 06collaboration-services, 10Infrastructure Security, and 6 others: Reboot cirrussearch in CODFW - https://phabricator.wikimedia.org/T432119#12124591 (10hnowlan) [14:31:52] XioNoX: meeting. durur6002 not expected but dns7002 expected (trixie reimage) [14:31:56] will look in a bit [14:32:09] 06Traffic, 06cloud-services-team, 06collaboration-services, 10Infrastructure Security, and 7 others: Reboot cirrussearch in CODFW - https://phabricator.wikimedia.org/T432119#12124605 (10hnowlan) Is there anything observability can help with on testing this? [14:43:48] 06Traffic, 06SRE, 06SRE Observability, 10SRE-SLO: Page on ATS backend errors relative to traffic - https://phabricator.wikimedia.org/T400675#12124674 (10hnowlan) [14:49:52] 06Traffic, 06cloud-services-team, 10Infrastructure Security, 06Infrastructure-Foundations, and 5 others: July 2026 Bookworm/Trixie reboots: Data Platform SRE - https://phabricator.wikimedia.org/T431826#12124698 (10bking) [14:50:27] 06Traffic, 06cloud-services-team, 10Infrastructure Security, 06Infrastructure-Foundations, and 5 others: July 2026 Bookworm/Trixie reboots: Data Platform SRE - https://phabricator.wikimedia.org/T431826#12124703 (10brouberol) [14:51:36] 06Traffic, 06SRE, 06SRE Observability, 10SRE-SLO: Page on ATS backend errors relative to traffic - https://phabricator.wikimedia.org/T400675#12124708 (10hnowlan) For the most part I think @fgiunchedi's original proposal of a ratio based on requests makes the most sense here. We already do [[ https://gerrit... [14:57:53] 10netops, 06Infrastructure-Foundations: eqiad: upgrade routers (2026) - https://phabricator.wikimedia.org/T417873#12124790 (10cmooney) [15:00:08] Hello. I have a patch here which adds a new service to the catalog. Would anyone have time to review it, please? https://gerrit.wikimedia.org/r/c/operations/puppet/+/1311036 [15:02:14] We are adding a port to the istio ingress gateway on the dse-k8s clusters, which allows us to route traffic into our postgresql clusters from outside of the k8s environment. It's for T432104 [15:02:14] T432104: Create a new PostgreSQL cluster database with read-only from Superset and write access from Spark - https://phabricator.wikimedia.org/T432104 [15:02:20] many thanks. [15:03:01] btullis: yeah happy to, we will comment on the CR [15:03:21] thx <3 [15:08:28] 06Traffic, 06cloud-services-team, 10Infrastructure Security, 06Infrastructure-Foundations, and 5 others: July 2026 Bookworm/Trixie reboots: Data Platform SRE - https://phabricator.wikimedia.org/T431826#12124837 (10klausman) [16:00:13] 06Traffic, 06SRE, 06SRE Observability, 10SRE-SLO: Page on ATS backend errors relative to traffic - https://phabricator.wikimedia.org/T400675#12125007 (10fgiunchedi) Thank you for the feedback @hnowlan @ssingh ! re: exclusion list I'm also expecting that moving to ratio-based paging will carry more signal (... [17:34:25] FIRING: [2x] SystemdUnitFailed: haproxy.service on cp4051:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [17:37:14] FIRING: [2x] HaproxyKafkaExporterDown: HaproxyKafka on cp4051 is down - https://wikitech.wikimedia.org/wiki/HAProxyKafka#HaproxyKafkaExporterDown - https://alerts.wikimedia.org/?q=alertname%3DHaproxyKafkaExporterDown [17:39:25] FIRING: [4x] SystemdUnitFailed: haproxy.service on cp4051:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [17:42:14] RESOLVED: [2x] HaproxyKafkaExporterDown: HaproxyKafka on cp4051 is down - https://wikitech.wikimedia.org/wiki/HAProxyKafka#HaproxyKafkaExporterDown - https://alerts.wikimedia.org/?q=alertname%3DHaproxyKafkaExporterDown [17:43:33] weird [17:43:43] happy again [17:43:44] oh well [17:44:25] RESOLVED: [4x] SystemdUnitFailed: haproxy.service on cp4051:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [17:47:32] 06Traffic: Upgrade Traffic hosts to trixie - https://phabricator.wikimedia.org/T401832#12125529 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by cdobbins@cumin2003 for host dns7002.wikimedia.org with OS trixie [18:23:38] 10netops, 06Infrastructure-Foundations, 06SRE: GSHUT (and other?) community matching/actions not working on SR-Linux - https://phabricator.wikimedia.org/T430810#12125726 (10cmooney) Going to close this one, with 'graceful-shutdown sender' configured on cr1-eqiad I can see that the community is matches on ssw... [18:23:49] 10netops, 06Infrastructure-Foundations, 06SRE: GSHUT (and other?) community matching/actions not working on SR-Linux - https://phabricator.wikimedia.org/T430810#12125727 (10cmooney) 05Open→03Resolved a:03cmooney [18:27:42] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-eqiad, 06SRE: Install new MPC10E-10C line cards on cr1-eqiad and cr2-eqiad slot 0. - https://phabricator.wikimedia.org/T426343#12125735 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=b3a498bf-e4dc-461c-9cbe-4c3255e0019c) set by... [18:29:26] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-eqiad, 06SRE: Install new MPC10E-10C line cards on cr1-eqiad and cr2-eqiad slot 0. - https://phabricator.wikimedia.org/T426343#12125749 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=944ebf34-ac03-4b88-9838-2666b665680e) set by... [19:00:04] 06Traffic: Upgrade Traffic hosts to trixie - https://phabricator.wikimedia.org/T401832#12125852 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by cdobbins@cumin2003 for host dns7002.wikimedia.org with OS trixie completed: - dns7002 (**WARN**) - Downtimed on Icinga/Alertmanager - Disabled... [19:21:40] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-eqiad, 06SRE: Install new MPC10E-10C line cards on cr1-eqiad and cr2-eqiad slot 0. - https://phabricator.wikimedia.org/T426343#12125906 (10cmooney) We have successfully installed the new switch-control boards and and MPC10E in cr2-eqiad. All is wor... [19:52:08] btw folks we still see bfd down for all these hosts off cr2-codfw. I did a session reset but made no diff [19:52:12] doh2001.wikimedia.org. [19:52:13] doh2002.wikimedia.org. [19:52:13] dns2004.wikimedia.org. [19:52:13] dns2005.wikimedia.org. [19:52:13] dns2006.wikimedia.org. [19:52:13] durum2001.codfw.wmnet. [19:52:13] durum2002.codfw.wmnet. [19:52:17] is it expected? [19:52:44] really? [19:52:49] all of dns in codfw? [19:53:03] no [19:53:07] not expected [19:54:17] just IPv6 [19:54:17] i am away from.keyboard. be back in ten [19:54:44] BGP is up to all of them so no panic [19:55:27] BFD only tears BGP down if it has been UP, and goes DOWN. In this case it seems the BFD session never came up (after I presume the reimages that have been done) [19:55:49] did they become multihop during that transition or something? [19:56:28] or become not-multihop [19:56:45] because it looks like all our birds default to multihop [19:56:47] they're configured on our side as mutlihop, so that's forced I think [19:56:54] ok [19:57:11] I think that's needed given the peering is to the CR loopback - and thus not an on-link IP [19:57:50] I said yesterday I thought it could be a BIRD bug we seen before - whereby it wasn't tearing down the CR session after a restart [19:58:05] but I think in that situation a force reset from the CR side fixed it, which it doesn't seem to have done here [20:00:11] birdc thinks bfd is up [20:00:14] bfd1 BFD --- up 15:32:15.534 [20:00:14] bgp1 BGP --- up 15:32:55.114 Established [20:00:14] bgp2 BGP --- up 15:32:24.320 Established [20:00:14] bgp3 BGP --- up 15:32:22.803 Established [20:00:15] bgp4 BGP --- up 15:32:27.582 Established [20:00:50] yeah [20:01:07] a "restart bfd1" sorted it on dns2005 [20:03:06] ok [20:03:38] maybe some race (puppet-dependency-configuration-hell) when initially setting up and bringing up bird post-reimage [20:04:02] e.g. maybe some interface config on the host wasn't actually ready for it when it first started [20:04:40] can just restart bfd1 on them all via cumin I guess [20:04:55] yeah [20:05:10] seems like bird isn't sending any BFD packets or responding to the CR's... [20:05:12] https://phabricator.wikimedia.org/P94865 [20:05:32] but yeah a "birdc restart bfd1" on them using cumin should fix it [20:05:36] interesting [20:05:38] how did this happen though [20:05:50] but there was no reimage in codfw [20:05:55] oh, a reboot potentially then yeah [20:06:06] it's different to the bug we seen before which was a session ID mis-match thing, and bird not tearing down it's status based on the mis-match (as it should) [20:06:44] sukhe: second half is starting can I leave you to do the restart/ [20:07:28] 2026-07-15T15:32:15.201236+00:00 dns2006 bird: bfd1: Socket error: bind: Cannot assign requested address [20:07:38] ah [20:07:39] ^ that, is probably some race in initial puppetization [20:07:46] I wonder is it after service restart [20:07:52] topranks: yes please [20:07:55] yeah, old pid is holding on to the socket or something [20:08:28] interesting that it was IPv6 in all cases, and IPv4 was ok [20:08:53] or a race in the systemd deps rather than puppet perhaps [20:08:58] I know it was just BFD and not the BGP session itself [20:09:07] or a race between systemd and IPv6 RA taking effect? who knows [20:09:13] but it all again comes back to the fact that we need to split how do we do advertisements for the nsXes [20:09:15] yeah it was only BFD [20:09:16] everything on bird is a disaster [20:09:26] blblack: indeed as the destination is off-net it needs the default route [20:09:50] when eqiad rows a/b are refreshed with L3 switches we can review the setup a little [20:09:55] as everything can peer with it's top-of-rack [20:09:56] we are one bird bug or one bad bird change away from all the nsXes not being announed :) [20:10:03] I wonder if there's a magic dependency in systemd to wait for the default route on IPv6 before starting a service? [20:11:03] "Cannot assign requested address" doesn't make sense though, we manually configure the IPv6 address [20:11:09] the RA is only needed for the default route [20:11:36] so I'd expect Bird to generate the packet and get a destination unreachable back from the system or something [20:11:52] jhathaway: do we know why puppet on dnsboxes is disabled because of Gerrit:1308760? [20:11:59] > Puppet is disabled. Gerrit:1308760 [20:12:07] yes that is me [20:12:20] ok so in systemd terms: bird.service has After=/BindsTo= on just anycast-healthchecker.service [20:12:28] just trying to be extra careful with this kafka patch [20:12:31] jhathaway: ok [20:12:48] anycast-healthchecker.service has After=/Requires= on "network.target" [20:13:06] should be finished in 10mins or so, but if I am blocking you folks we can do something else? [20:13:09] there's some explanation here: https://systemd.io/NETWORK_ONLINE/ [20:13:18] jhathaway: no sorry just wanted to be sure. please carry on. [20:13:25] but basically, we probably want that to be network-online.target at least (still might not be enough, but at least closer) [20:13:48] sukhe: great, should have put my name on the disable message, sorry [20:13:55] blblack: that sounds sensible yes [20:15:14] yeah we will patch it [20:15:22] makes sense, that's what we have done in oher places [20:16:00] topranks: it should all be resolved now [20:16:22] with the England goal ? [20:16:36] ha BFD on the affected hosts above [20:16:48] don't @ me but I don't watch football :) [20:16:51] great strike, didn't expect it'd fix the bfd problem too :) [20:17:29] haha no stress I know [20:17:51] the only real winners in soccer are those that avoid watching it /ducks [20:20:37] ducks love football, I'd say all the English ducks are going mad right now :P [20:21:59] btw, the systemd.io article of course eventually berates software authors to avoid needing an accurate "network-online" dependency anyways (because it's truly impossible to get it perfect for all scenarios) [20:22:32] one of the methods they discuss to make things more reliably is to set IP_FREEBIND sockopt (which will let bind() succeed even before an address is available or working) [20:22:40] to be fair to the bird guys (and I guess the BFD authors) this does fail in the right way [20:23:05] and bird does have a config option "bgp-free-bind" to set that... [20:23:16] but it's not clear to me (yet) if this can be set for or affects the BFD socket [20:23:21] hmm yeah that might not be bad idea [20:23:30] though seems in our case we don't bind to a specific IP? [20:23:35] cmooney@durum2002:~$ ss -tulpna | grep 4784 [20:23:35] udp UNCONN 0 0 0.0.0.0:4784 0.0.0.0:* [20:23:35] udp UNCONN 0 0 [::]:4784 [::]:* [20:23:45] hmmmm [20:25:11] my thinking is Bird should be able to generate the packets even if there is no default route yet (they just don't get transmitted) [20:25:26] but I do know Bird in its internal structures has a concept of what interface a protocol is bound to [20:25:45] that for instance can affect bfd and bgp multihop or not if you don't configure it [20:26:07] I wonder if bird starts before the RA does it fail to determine what interface the peer is reachable on, and fail? [20:27:26] though that said BGP starts ok.. though checking a few of these hosts the connection was initiated by the CR [20:28:21] at the end of the day this is rare, we get alerted about it, and it fails in a way that leaves BGP up and traffic working [20:28:43] yeah but it bugs me [20:28:45] if it's a default route / RA issue we can address that when we move to peering with the top-of-raack [20:28:51] blblack: oh for sure [20:28:58] I'm trying to even reconstruct what happened in this supposed "reboot" event on 2006 [20:29:05] because the syslog looks fishy [20:30:11] uptime is 5 hours on it, which matches the bfd session failures [20:30:22] there were reboots for kernel patches afaik [20:30:36] yeah maybe it's just some unrelated confusion [20:30:58] but the /var/log/syslog record shows it shutting everything down in systemd like a reboot, then 3 seconds later it bringing everything back on like a fresh boot [20:31:23] oh 3 minutes, brain fart [20:31:28] 16:30:22 < topranks> there were reboots for kernel patches afaik [20:31:29] yes [20:31:37] and the timing matches yep [20:32:11] 2026-07-15T15:32:14.817718+00:00 dns2006 systemd[1]: Starting bird.service [20:32:22] there is also of course another consideration [20:32:29] bird isn't advertising anything right away [20:32:41] since anycast-hc is depooled for things post reboot [20:32:47] 2026-07-15T15:32:15.192401+00:00 dns2006 systemd[1]: Started bird.service - BIRD Internet Routing Daemon (BIRD2: IPv4 and IPv6). [20:32:47] it is only pooled after the reboot and everything else is done [20:32:48] I don't think that should affect the protocols though [20:32:50] 2026-07-15T15:32:15.201064+00:00 dns2006 bird: Started [20:32:52] 2026-07-15T15:32:15.201236+00:00 dns2006 bird: bfd1: Socket error: bind: Cannot assign requested address [20:33:40] topranks: I am thinking if this ties into bird not finding that dummy IP and failing in that way somehow. it's a weak link yes [20:34:28] Bird is configured to bind for BGP to the specific interface IP [20:34:31] "local 2620:0:860:4:208:80:153:107 as 64605;" [20:35:08] but that should be added with the interface [20:35:14] I don't see anything explicitly out-of-order with the systemd steps on startup. And it only affected IPv6. So I still think somehow RA delay vs very early bird daemon start is the driver. [20:35:15] cmooney@dns2006:~$ grep 2620:0:860:4:208:80:153:107 /etc/network/interfaces [20:35:15] up ip addr add 2620:0:860:4:208:80:153:107/64 dev eno8303 [20:36:10] https://bird.network.cz/pipermail/bird-users/2018-August/012620.html [20:36:16] >This seems to be a common pattern for services that are started when [20:36:19] network is supposedly ready, but it really isn't (see many discussions [20:36:22] around network.target vs. network-online.target). [20:36:24] exact same error, not sure if the exact same cause [20:36:42] if that is the cause, the hacky way to work around it would be to put something like: ExecStartPre=/bin/sleep N [20:36:49] where N > than our RA interval [20:37:35] or as a more-general solution, rather than hacking that into bird.service... [20:37:52] blback: yeah I overall think that is likely, though it may be something within bird maybe [20:37:58] create a type=oneshot or whatever service, which sleeps for > RA interval, and which network-online.target depends on [20:38:13] maybe that would solve it more-generally? [20:38:42] afaik the system should send an RS when the interface comes up, resulting in an out-of-interval RA being sent [20:39:03] i should try to confirm [20:39:07] [or we could revive some years-old tickets and actually fix our ipv6 configuration mess to not rely on RA] [20:40:11] yeah that is going to happen I think just need to get a little time to work on it