[02:33:20] 10netops, 06Traffic, 06Data-Persistence, 06Data-Platform-SRE, and 5 others: codfw: rack B2 maintenance 2026-07-01 11:00 am CT - https://phabricator.wikimedia.org/T429861#12078721 (10Papaul) 05Open→03Resolved Complete [02:34:32] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-codfw, 06SRE: codfw: pod AB switches upgrade (2026) - https://phabricator.wikimedia.org/T426197#12078723 (10Papaul) [02:36:00] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-codfw, 06SRE: codfw: rack A8 maintenance 2026-07-01 10:00 am CT - https://phabricator.wikimedia.org/T429856#12078724 (10Papaul) 05Open→03Resolved Complete [02:36:30] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-codfw, 06SRE: codfw: pod AB switches upgrade (2026) - https://phabricator.wikimedia.org/T426197#12078726 (10Papaul) [06:02:28] 06Traffic, 06collaboration-services, 10Continuous-Integration-Infrastructure, 10Gerrit, and 2 others: Get traces/logs from ATS on Gerrit CI errors - https://phabricator.wikimedia.org/T430901 (10ABran-WMF) 03NEW [06:33:21] 06Traffic, 06collaboration-services, 10Continuous-Integration-Infrastructure, 10Gerrit, and 2 others: Get traces/logs from ATS on Gerrit CI errors - https://phabricator.wikimedia.org/T430901#12078972 (10ABran-WMF) **(Analysis below written by Claude - with Fable 5 xhigh effort)** I analyzed the full-data... [07:30:35] 06Traffic, 06collaboration-services, 10Continuous-Integration-Infrastructure, 10Gerrit, and 2 others: Get traces/logs from ATS on Gerrit CI errors - https://phabricator.wikimedia.org/T430901#12079033 (10ABran-WMF) p:05Triage→03High [08:08:01] 10netops, 06Traffic, 06Data-Persistence, 06Infrastructure-Foundations, 06ServiceOps new: codfw: rack B3 maintenance - https://phabricator.wikimedia.org/T430909 (10ayounsi) 03NEW p:05Triage→03Medium [08:08:33] 10netops, 06Traffic, 06Data-Persistence, 06Infrastructure-Foundations, 06ServiceOps new: codfw: rack B3 maintenance - https://phabricator.wikimedia.org/T430909#12079216 (10ayounsi) [08:08:37] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-codfw, 06SRE: codfw: pod AB switches upgrade (2026) - https://phabricator.wikimedia.org/T426197#12079217 (10ayounsi) [08:13:39] 10netops, 06Traffic, 06Data-Persistence, 06Infrastructure-Foundations, 06ServiceOps new: codfw: rack B3 maintenance - https://phabricator.wikimedia.org/T430909#12079236 (10ayounsi) [08:25:47] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-codfw, 06SRE: codfw: pod AB switches upgrade (2026) - https://phabricator.wikimedia.org/T426197#12079304 (10ayounsi) [08:26:21] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-codfw, 06SRE: codfw: pod AB switches upgrade (2026) - https://phabricator.wikimedia.org/T426197#12079312 (10ayounsi) [08:53:59] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-codfw, 06SRE: codfw: pod AB switches upgrade (2026) - https://phabricator.wikimedia.org/T426197#12079442 (10ayounsi) [09:12:34] 10netops, 06Traffic, 10Cloud-Services, 06collaboration-services, and 6 others: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918 (10ayounsi) 03NEW p:05Triage→03Medium The #Cloud-Services project tag is not intended to have any tasks. Please check the list on https://phabricator.w... [09:22:24] 10netops, 06Traffic, 10Cloud-Services, 06collaboration-services, and 6 others: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918#12079602 (10ayounsi) @cmooney @Papaul can one of you take that upgrade ? And update the date/time to anything that suits you in that week. [09:25:03] 10netops, 06Traffic, 10Cloud-Services, 06collaboration-services, and 6 others: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918#12079609 (10CWilliams-WMF) [09:25:31] 10netops, 06Traffic, 10Cloud-Services, 06collaboration-services, and 6 others: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918#12079614 (10CWilliams-WMF) [09:27:30] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-codfw, 06SRE: codfw: pod AB switches upgrade (2026) - https://phabricator.wikimedia.org/T426197#12079630 (10ayounsi) [09:27:41] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-codfw, 06SRE: codfw: pod AB switches upgrade (2026) - https://phabricator.wikimedia.org/T426197#12079635 (10ayounsi) [09:27:59] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-codfw, 06SRE: codfw: pod AB switches upgrade (2026) - https://phabricator.wikimedia.org/T426197#12079637 (10ayounsi) [09:28:07] 10netops, 06Traffic, 10Cloud-Services, 06collaboration-services, and 6 others: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918#12079636 (10ayounsi) [09:30:45] 10netops, 06Traffic, 10Cloud-Services, 06collaboration-services, and 6 others: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918#12079648 (10CWilliams-WMF) [09:31:26] 06Traffic, 06collaboration-services, 10Continuous-Integration-Infrastructure, 10Gerrit, and 2 others: Get traces/logs from ATS on Gerrit CI errors - https://phabricator.wikimedia.org/T430901#12079653 (10Joe) (**Analysis below written by Giuseppe 1.0 on moderate effort**) :P There is at least one bad assu... [09:42:22] 10netops, 06Traffic, 10Cloud-Services, 06collaboration-services, and 6 others: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918#12079752 (10Marostegui) db2160 needs just a downtime. dbproxy will fail during the maintenance (just irc notification) and they'll recover on their own once... [09:49:01] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-codfw, 06SRE: codfw: pod AB switches upgrade (2026) - https://phabricator.wikimedia.org/T426197#12079790 (10ayounsi) [09:55:38] 06Traffic, 06collaboration-services, 06Data-Persistence, 06Data-Platform-SRE, and 2 others: codfw: rack B8 maintenance - https://phabricator.wikimedia.org/T430929 (10ayounsi) 03NEW [09:56:48] 06Traffic, 06collaboration-services, 06Data-Persistence, 06Data-Platform-SRE, and 2 others: codfw: rack B8 maintenance - https://phabricator.wikimedia.org/T430929#12079819 (10ayounsi) [09:56:49] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-codfw, 06SRE: codfw: pod AB switches upgrade (2026) - https://phabricator.wikimedia.org/T426197#12079820 (10ayounsi) [09:58:24] 06Traffic, 06collaboration-services, 06Data-Persistence, 06Data-Platform-SRE, and 2 others: codfw: rack B8 maintenance - https://phabricator.wikimedia.org/T430929#12079823 (10ayounsi) [10:00:26] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-codfw, 06SRE: codfw: pod AB switches upgrade (2026) - https://phabricator.wikimedia.org/T426197#12079859 (10ayounsi) [10:00:43] 10netops, 06Traffic, 06Data-Persistence, 06Infrastructure-Foundations, 06ServiceOps new: codfw: rack B3 maintenance - https://phabricator.wikimedia.org/T430909#12079866 (10ayounsi) a:03ayounsi [10:04:14] 10netops, 06Traffic, 10Cloud-Services, 06collaboration-services, and 6 others: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918#12079891 (10cmooney) a:03cmooney [10:04:57] 06Traffic, 06collaboration-services, 06Data-Persistence, 06Data-Platform-SRE, and 2 others: codfw: rack B8 maintenance - https://phabricator.wikimedia.org/T430929#12079896 (10cmooney) a:03cmooney [10:06:13] 10netops, 06Traffic, 10Cloud-Services, 06collaboration-services, and 6 others: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918#12079906 (10cmooney) [10:13:04] 06Traffic, 06collaboration-services, 06Data-Persistence, 06Data-Platform-SRE, and 2 others: codfw: rack B8 maintenance - https://phabricator.wikimedia.org/T430929#12079928 (10cmooney) [10:18:26] 10netops, 06Traffic, 06Data-Persistence, 06Infrastructure-Foundations, 06ServiceOps new: codfw: rack B3 maintenance - https://phabricator.wikimedia.org/T430909#12079937 (10jcrespo) backup2014: should be ok without depooling for a short period ms-backup2004 will require manual stopping of backup processes... [10:19:01] 10netops, 06Traffic, 10Cloud-VPS, 06Data-Platform-SRE, and 3 others: Plan to make clouddumps more resilient and easier to operate - https://phabricator.wikimedia.org/T411248#12079944 (10fgiunchedi) Ok so I spent some time getting a better understanding of NFS and how Linux clients reacts to a server failov... [11:47:24] 06Traffic, 06collaboration-services, 06Data-Persistence, 06Data-Platform-SRE, and 2 others: codfw: rack B8 maintenance - https://phabricator.wikimedia.org/T430929#12080169 (10Jelto) [11:48:50] 06Traffic, 06collaboration-services, 06Data-Persistence, 06Data-Platform-SRE, and 2 others: codfw: rack B8 maintenance - https://phabricator.wikimedia.org/T430929#12080173 (10Jelto) [12:18:27] seems like wikidough was having some issues for a couple of minutes starting at ~12:05Z? got a 502 when manually testing over DoT https://phabricator.wikimedia.org/P94701, can't tell exactly what my forwarding+caching resolver got over DoT but it was just serving SERVFAILs as well :( [12:23:52] 10netops, 06Infrastructure-Foundations: eqiad: upgrade routers (2026) - https://phabricator.wikimedia.org/T417873#12080468 (10cmooney) [12:27:05] 10netops, 06Traffic, 10Cloud-Services, 06collaboration-services, and 7 others: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918#12080487 (10cmooney) [12:37:45] hey folks [12:37:56] ok if I disable puppet on all cp hosts to rollout https://gerrit.wikimedia.org/r/c/operations/puppet/+/1306545 ? [12:37:58] cc: fabfur [12:39:24] ack [12:41:50] fabfur: for the rollout - should I depool - run-puppet - repool for every node? Or just for the first ones, and then rollout carefully but without repooling? [12:43:02] I would do a text node and an upload one on magru, if everything is ok, repool the two nodes, have a look at eventual haproxy metrics and if all is good rollout everywhere (re-enable puppet) [12:43:32] okk [12:55:15] fabfur: https://gerrit.wikimedia.org/r/c/operations/puppet/+/1307142 :D [12:55:18] if you have a second [12:55:35] ah missed that [12:55:44] +1 [12:59:28] Could someone help me with deployment of this upload.wm.o varnish patch today or tomorrow perhaps? This is to fixup the with corrupt URLs in the iOS app and unblocks re-enabling media provenance on Wikipedia. https://gerrit.wikimedia.org/r/c/operations/puppet/+/1306230 [13:01:49] Krinkle: how do you want to proceed with rollout? [13:02:14] one host on verify with curl once more, then everywhere? [13:02:52] that's what I've done in the past, happy to follow other approach as well [13:03:47] ok which host are you hitting currently from your computer? [13:04:58] elukey: knock when rollout is complete [13:06:25] 10netops, 06Infrastructure-Foundations, 06SRE: cr2-esams rpd failure after enabling bgp 'graceful-shutdown' (June 2026) - https://phabricator.wikimedia.org/T429386#12080641 (10cmooney) 05Resolved→03Open Probably a little premature to close. I've opened ticket 2026-0616-761841 to confirm exactly what Jun... [13:17:22] fabfur: testing in magru went fine, puppet re-enabled! [13:25:27] 10netops, 06Traffic, 10Cloud-VPS, 06Data-Platform-SRE, and 3 others: Plan to make clouddumps more resilient and easier to operate - https://phabricator.wikimedia.org/T411248#12080727 (10ayounsi) >>! In T411248#12076020, @fgiunchedi wrote: >>>! In T411248#12074721, @ayounsi wrote: >> @fgiunchedi maybe silly... [13:26:51] taavi: thanks for sharing. that's interesting; there wasn't any planned maintenance or anything else at that time [13:26:56] I am assuming you connect to esams? [13:27:03] will take a look in a bit [13:30:31] sukhe: at least at the moment, yeah. my main forwarded doesn't keep detailed logs, and it recovered by the time I managed to switch to an alternative upstream, so the only things I have is that one kdig in the history and when the number of SERVFAILs started going up [13:32:48] taavi: ok thanks. I will see if there was something at ~12UTC [13:33:32] there is one "famous" case in which SERVFAIL is expected for a relatively popular domain but that should be persisent [13:33:41] "expected" is better, see https://phabricator.wikimedia.org/T420514 [13:35:23] fabfur: cp3081 [13:35:36] sorry for distracted because it differed between browser and curl and forgot what I was doing [13:36:20] anyway that was because one was over http, seems we route that differently [13:36:24] Krinkle: that's interesting. it should not differ between browser and curl? [13:36:27] or at least it's part of the hash I guess [13:36:37] I did curl without https:// [13:36:48] with it, it's the same [13:37:12] elukey: ack [13:37:29] Krinkle: ah ok, fair [13:37:37] port 80 vs 443 [13:39:08] so I would 1) disable puppet on A:cp 2) enable puppet on cp3081 only 3) merge the CR 4) apply puppet on cp3081 && test [13:39:16] if all is good, enable puppet everywhere [13:39:39] SGTM [14:06:45] 10netops, 06Infrastructure-Foundations, 06SRE: Blackbox probe for TLS cert expriy failing on multiple eqiad SR-Linux nodes - https://phabricator.wikimedia.org/T429242#12081032 (10cmooney) Nokia came back to advise this is a known bug in 24.10.4. ` Hello Cathal, The gRPC server stops responding on SR-Linux?... [14:07:25] 06Traffic, 06Data-Engineering (Q4 FS25/26 April 1st - June 30st): Tune refine webrequest data loss threshold to avoid noisy irrelevant alerts. - https://phabricator.wikimedia.org/T429809#12081034 (10Snwachukwu) I made the following changes to the threshold: **data_loss_warning** - data_loss_warning: prev... [14:40:48] 06Traffic, 06collaboration-services, 10Continuous-Integration-Infrastructure, 10Gerrit, and 2 others: Get traces/logs from ATS on Gerrit CI errors - https://phabricator.wikimedia.org/T430901#12081230 (10ABran-WMF) Thanks for taking a look! For the sake of thoroughness, I asked Claude to fetch the relevant... [14:48:21] 06Traffic, 06collaboration-services, 06Data-Persistence, 06Infrastructure-Foundations, and 2 others: codfw: rack B8 maintenance - https://phabricator.wikimedia.org/T430929#12081265 (10Gehel) [14:48:44] 10netops, 06Traffic, 06Infrastructure-Foundations, 06cloud-services-team (FY2025/2026-Q3-Q4): LOAD_BALANCER_HEALTH_CHECKS firewall set and asymmetry lvs1018 / lvs1020 - https://phabricator.wikimedia.org/T430651#12081268 (10cmooney) I merged the patch to add the lvs public IPs to the healtcheck list now, an... [14:54:04] 10netops, 06Traffic, 06Data-Persistence, 06DC-Ops, and 5 others: codfw: rack B2 maintenance 2026-07-01 11:00 am CT - https://phabricator.wikimedia.org/T429861#12081319 (10Gehel) [14:56:13] Krinkle: maybe I wasn't clear, are you confident with this rollout or do you want someone to do that? [14:57:19] 10netops, 06Traffic, 10Cloud-Services, 06collaboration-services, and 6 others: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918#12081351 (10Gehel) [15:00:09] fabfur: I'm not an SRE / don't have access. [15:00:54] Looking for an SRE to deploy it (and review it from a Puppet/Varnish pov beyond intent/output that I've implemented) [15:01:06] ah sorry, ok, I'll try right after this meeting [15:48:22] fabfur: bbl in 2h, might ask bre.tt instead by then depending on work hours :) [15:48:46] yep, I think it's a bin late for me... :) [15:48:53] fabfur: I've got it, take off [15:49:03] brett: thanks! [16:03:27] Krinkle: Is this the only way to move forward? Generally we don't like handling app-level bug workarounds at the edge [16:05:27] hi traffic! resurfacing my question from yesterday, we're about to undeploy the ipoid service including the LVS bits -- I can handle that if everything goes smoothly, but if I get into LVS trouble, will anyone be around? [16:05:43] rzl: Apologies for missing that yesterday. I'm here [16:05:50] awesome, thanks :) [16:34:59] brett: yeah, the bug has dormant in the apps for years, image urls never had query strings so the bad URL parsing code they had (for resizing images without an API call; one half happened to handled query strings the other half not, the offsets go in the middle of the string), blew up when we rolled out image provenance [16:35:36] The bug is fixed in the latest version but it'll take months before that's safe to reroll due to the nature of app stores and local updates [16:35:55] So we found a way to make those corrupted URLs work. [16:38:05] The alternative would be to vary pageview and API responses by user agent and configure MW to never emit these for pageviews / API calls from the app. Made more difficult by PCS in the middle serving both apps and web, and layers of caches etc. This seemed more viable. [16:39:00] Ie carefully prevent the app from seeing something that it crashes on, or let it happen, detect, and fixup those requests. [16:39:43] rzl: brett: just a word of caution I guess, with my ever-cautious-but-does-no-real-work hat on [16:39:59] removing services can lead to unintended side effects and it's Thursday afternoon, with a long weekend [16:40:05] do factor that in but both of your call :) [16:40:09] Basically it turns Foo/100px-Foo.jpg?query into Foo?query/100px-Foo.jpg?query [16:42:14] sukhe: true enough! fwiw though, the ipoid backend is already gone [16:42:21] and sorry rzl for missing your question yesterday [16:42:50] rzl: yeah, the tricky bit is the PyBal restarts and the subsequent ipvsadm delete command [16:43:00] as in the kubernetes deployments are already deleted -- currently we have the lvs service pointing nowhere, so it's a config change but if it was going to break anything it'd already be broken [16:43:02] oh hm, that's fair [16:43:11] but like I said, it's brett's call and I don't mean to override it [16:43:27] I guess it's easy to be over cautious like me when you don't do real work :P [16:43:29] no for sure, that's good input though [16:43:33] but just pointing it out for future [16:43:34] haha [16:44:51] hmm, I've puppet-merged as far as production->lvs_setup, and finishing that step out should be pretty low-risk, right? then we could postpone lvs_setup->service_setup, and everything after that, until Monday [16:44:57] wdyt both? [16:45:34] rzl: I think +1 since you started already and just be a little extra careful on the ipvsadm delete thing [16:45:54] rzl: I'd say keep going [16:46:11] sukhe: sure -- ipvsadm would come after the service_setup change, if I understand correctly? so we could do that monday [16:46:36] rzl: deferring to brett, if he is comfortable, my +1 stands :D [16:48:12] thanks both - I'll finish https://wikitech.wikimedia.org/wiki/LVS#Remove_network_probes_/_monitoring (run puppet on authdns) and then stop before https://wikitech.wikimedia.org/wiki/LVS#Remove_the_service_from_the_load-balancers_and_the_backend_servers [16:48:23] ack [16:55:46] Krinkle: Considering this long weekend and this is EOW, what's the risk associated with this change? [16:57:46] brett: we're in lvs_setup, so probes are gone, I see on alertmanager that the ProbeDown alert is cleared -- I'm extending my silences into next week just for good measure, but otherwise I'm hands-off for today [16:58:01] understood. Thanks for your flexibility [16:59:48] 06Traffic: Bump purged go version - https://phabricator.wikimedia.org/T401687#12082199 (10ssingh) +1 from me on marking this as resolved, at least. [17:01:55] 06Traffic, 06SRE: Traffic cache daemon restart scripts need some rework - https://phabricator.wikimedia.org/T346640#12082207 (10ssingh) This will be greatly simplified when we move drmrs to single backend, so there is just going to be the `cdn` service which we will need to care about since depooling a host fo... [17:02:42] 06Traffic: Expose the list of CPUs handling NIC queues to prometheus - https://phabricator.wikimedia.org/T397303#12082209 (10ssingh) If you were looking for a vote to closing this, submitting mine for a yes since I also don't see what else needs to be done. [17:04:56] for planning, we'll wrap up ipoid on Tuesday 16:00 in what's ostensibly the puppet window again :) brett does that work for you? [17:05:26] rzl: Yep, thanks! Marked my calendar [17:05:35] much appreciated [17:06:23] 06Traffic, 06SRE: Traffic cache daemon restart scripts need some rework - https://phabricator.wikimedia.org/T346640#12082218 (10BCornwall) I think the larger question is whether we want pooling logic on the hosts themselves. I argue against it and think that such logic/scripts should be removed - the node shou... [17:06:57] brett: I'd say low risk because the condition is not reachable except by the app and except for damaged URLs. As of now, there are now damaged URls because we rolled back the media provenance feature.. Next week I'll reroll the media provenance and, thanks to having this in place, will not cause widespread crashes a second time. [17:07:10] no damaged* [17:07:57] so it should essentially be a no-op [17:09:31] Yep. There may be a small number of cached pages that still have some media provenance in them in which case we'd fix those instead of the iOS app crashing. However since the rollback we've not heard new reports of crashes so it's assumed to be a no-op. But yeah in theory also some immediate relief of otherwise continued iOS crashes [17:15:32] not to delay this but I do have a question about the task itself [17:15:33] > However, as is the nature with apps, users control if and when they update, so it may take months for "enough" users to have the update. [17:15:48] that surprises me -- on iOS, apps don't update themselves? [17:16:00] I don't use iOS so I know nothing about it but that seems to be a security nightmare? [17:16:21] this is from https://phabricator.wikimedia.org/T427623 and on why we are doing the VCL change above [17:20:10] 8.2.0 which has the fixes, is the vast majority of WikipediaApp user agents [17:20:30] brett: oh good point, that too [17:20:33] (on the usage) [17:20:51] but yeah, it hurts users that aren't on the latest version, which seems counter to what we should do [17:21:13] So we either punish the MW devs with pain, delay the media provenance, or take the debt on [17:21:14] brett: got a Turnilo link handy? I am curious to see the percentages [17:21:36] https://w.wiki/RyhW [17:21:40] thanks [17:23:20] App team said they want to support old versions for 3-6 months depending on usage. [17:24:02] Krinkle: ok. [17:24:10] Well, then they need to not make things crash for 3-6 months :P [17:24:25] It's been about 1 month since the bug fix. The most recent other server-related support issue (OAuth login issue) was last year where we kept it for 6 months I think. [17:24:36] I jest, I jest [17:24:36] that's interesting [17:26:04] Yeah, once could talk about why someone would implement a URL parser from scratch or why someone would reverse engineer a thumb URL from an original non-thumb URL (eg Foo.pdf -> thumb/Foo.pdf/200px-Foo.pdf.jpg). But those are past mistakes outside MW control. [17:26:19] It's pretty flawed to think that they can depend on any single version that they would release as supported for 3-6 months [17:26:34] Or are they meaning supporting specific release branches? [17:26:50] Surely a bad release with a segfault on startup can't be supported [17:26:53] There's not really the concept of semver in app stores. It's linear. [17:27:02] I presume it is because some older iOS devices don't even get updates or something? [17:27:24] Yeah, but with said segfault-on-startup release, they'd be forced to tell the user "you need to update". I don't see how this is any different [17:27:31] Automatic updates are user controlled, and dependent on WiFi due to being large binaries [17:27:45] I'm concerned that they're treating the edge as their personal backporting mechanism [17:28:35] You can bypass and do it on mobile data, but unless on 4G unlimited you probably wouldn't want to. [17:28:52] app bloat is another issue :P [17:29:09] Krinkle: ok, that's fair. then perhaps I am guessing they want to keep supporting this for countries where data is expensive and users don't have access to wifi [17:29:33] Ack, as is reputation. Native apps have a bad reputation for being too large even if ours isn't. [17:29:36] 87 MiB on fdroid - I expected it to be hundreds, so I'm pleasantly surprised :) [17:30:50] Pulling it back, though - network-level workarounds seem to be the mechanism by which they'll fix things when they push bad versions with no recourse to updates? [17:31:03] I don't like this from a planning perspective, but that's beyond my pay-grade [17:31:15] In any case, I +1'ed the patch [17:31:16] no I think your concern is fair [17:32:41] the thing that helps contextualize this is that some people will probably not update for a while [17:33:01] and so I guess we fix what we can but also not let this be the standard solution for fixing this stuff in a way [17:33:15] The iOS app in question is 209M [17:34:02] ew [17:34:02] Yeah normally id say wait six months and absorb the delay [17:34:16] But in this case it's something "we" want. [17:34:23] yeah [17:34:43] it'd be nicer if it were their own feature, then we could just say "nah, it's your problem" and make them wait [17:34:52] So it circles back to how badly we want to get media provenance going. [17:37:29] I've also considered absorbing it in PCS, by stripping the query params there (and thus preventing exercising the bug) and living without provenance for app clients for now while we have it elsewhere. But found not all image URLs come through there. The apps calls the MW API directly as well. [17:40:43] Krinkle: merging now [17:42:16] er, rather, running tests [17:42:24] 10netops, 10Infrastructure Security, 06Infrastructure-Foundations, 06SRE: Lumen transport eqiad codfw down July 2026 - https://phabricator.wikimedia.org/T430874#12082394 (10VRiley-WMF) Swapped out the fiber and it looks like it came back up [17:42:31] Ack. I'll be ready for testing in ~15min. [17:43:42] 10netops, 10Infrastructure Security, 06Infrastructure-Foundations, 06SRE: Lumen transport eqiad codfw down July 2026 - https://phabricator.wikimedia.org/T430874#12082397 (10cmooney) 05Open→03Resolved a:03cmooney @VRiley-WMF thanks! Yep traffic back on the link looks good now cheers. ` cmooney@re... [17:52:52] 06Traffic, 06SRE: Traffic cache daemon restart scripts need some rework - https://phabricator.wikimedia.org/T346640#12082434 (10ssingh) >>! In T346640#12082218, @BCornwall wrote: > I think the larger question is whether we want pooling logic on the hosts themselves. I argue against it and think that such logic... [17:59:31] 06Traffic, 10DNS, 06SRE: Wikimedia DNS over Port 53 support - https://phabricator.wikimedia.org/T429650#12082472 (10ssingh) Thanks for filing this task. I agree that in the absence of discovery mechanisms (for which there are some draft RFCs that are looking into this problem), configuring encrypted resolver... [18:30:26] hi folks o/ mind if I ask for a pair of reviews on these perchance? [18:30:26] [0] - https://gerrit.wikimedia.org/r/c/operations/dns/+/1307216 [18:30:26] [1] - https://gerrit.wikimedia.org/r/c/operations/dns/+/1307215 [18:30:28] sukhe: have to say I agree, we don’t need to run public plaintext resolvers. [18:33:56] topranks: ah you saw that. yeah, that's a decision Brandon and I made at that time [18:34:25] our reasoning was basically what we wrote in the ticket which I am guessing is yours as well IIRC since I remember you and I discussed it too [18:34:28] in a separate task [18:34:45] jasmine_: looking [18:34:51] yeah. it would probably give a false sense of security if we did [18:35:43] people can run their own recursor if they don’t mind plaintext transport and at the same time don’t trust the other public resolver operators [18:36:02] that's also fair yeah [18:36:48] sukhe:tyty! [19:02:36] 06Traffic: Expose the list of CPUs handling NIC queues to prometheus - https://phabricator.wikimedia.org/T397303#12082747 (10BCornwall) 05Open→03Resolved a:03BCornwall [19:21:51] Krinkle: ready for a deploy? [19:22:10] i.e. do you have a way to easily test the mobile app changes? [19:22:17] cos I sure don't :) [19:24:26] brett: ack. I'll be using cURL same as the vtc basically. I have the app as a null test, but that's not particularly insightful right now since the bug isn't live. I'll test both normal URLs and bugged URLs in cURL. [19:24:58] I'm on cp3081 [19:25:55] okay, will deploy on that [19:33:07] Krinkle: Deployed on cp3081 [19:34:55] dumb question. Google is not helping :( obj.hits in vcl is the number of hits after the server has started or it gets restarted each day? if the ttl of the object is a day, does it mean it resets after a day? Sorry, I looked at docs and couldn't find anything [19:36:55] Amir1: I believe that obj.hits does not get reset even after cache ttl expiration [19:37:52] ah okay, thanks [19:39:59] The object is just replaced once the TTL is hit (i.e. a miss) IIRC [19:41:28] brett: bugged URLs work as expected for WikipediaApp UA but continue to fail as normal for everyone else. Regular URLs still work normal for both it and normal UAs. so... LGTM! [19:41:41] the picture on today's main page might be an interesting example: [19:41:42] https://upload.wikimedia.org/wikipedia/commons/thumb/e/e2/LCF-24-_MRF_Search_%26_Rescue_Efforts_in_Venezuela_%289781661%29.jpg/330px-LCF-24-_MRF_Search_%26_Rescue_Efforts_in_Venezuela_%289781661%29.jpg [19:42:00] Age: 62848 which is 17h [19:42:27] I guess if an object is popular enough and it gets succesfully renewed during the grace/keep period, both age and hits might just keep going? [19:42:58] I would think so, yes, but I'm also having a hard time getting that answer via docs [19:43:46] Aye, and I'm struggling to find anything with an Age over 24h even stable entries on the main page such as the Commons logo that hasn't changed in contnet or size/url in years [19:43:56] I guess that's just regular deployments for you [19:44:53] although config don't wipe cache afaik. and then there's ATS backfill where I think Age propagates? [19:45:12] so maybe it does reset? I don't know. [19:46:40] Amir1: Don't quote me on that response - I can't remember the specific behavior, but docs only talk about it in respect to how to use it, not how it behaves (i.e. obj.hits being 0 by the time vcl_deliver is reached == cache miss) [19:48:13] But yeah, I think I remember that it keeps the hits there until the miss occurs, and *then* it's set to 0 [19:49:17] Krinkle: Rolling out fleet wide [19:51:19] right, and for the purposes of this, an internal "If-Modified-Since" request that gets an HTTP 304 response and thus renews the object as-is without fetching it from the origin, is a "miss"? I wonder how this gets reflected on the X-Cache-Status header on that first request after ttl expires, whether they'd see it as "miss" or as "hit". Well, I suppose they'd see neither because grace/keep is basically stale-while-revalidate so the [19:51:19] end-user would get the stale object which counts as hit. [19:54:23] still quite useful [19:55:05] thanks [19:57:41] Krinkle: deployed fleet-wide [20:11:19] ack [22:46:22] 06Traffic, 06Data-Engineering (Q4 FS25/26 April 1st - June 30st): Tune refine webrequest data loss threshold to avoid noisy irrelevant alerts. - https://phabricator.wikimedia.org/T429809#12083707 (10Ahoelzl) 05Open→03Resolved