[00:03:14] 10Toolforge, 06tools-platform-team, 10Toolhub, 07Documentation: Set up a dedicate and official space for documenation of toolforge tools - https://phabricator.wikimedia.org/T433325#12161381 (10JJMC89) > My preference would be having something like a wiki page with subpages (e.g. Toolforge:Coverme in wikite... [01:14:47] FIRING: NodeDown: Node cloudcephosd1044 has been down for long. - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Runbooks/NodeDown - https://grafana.wikimedia.org/d/000000377/host-overview?orgId=1&var-server=cloudcephosd1044 - https://alerts.wikimedia.org/?q=alertname%3DNodeDown [01:15:25] 06cloud-services-team, 10Cloud-VPS, 06tools-infrastructure-team, 10Ceph: cloudcephosd1044 boot issues - https://phabricator.wikimedia.org/T429267#12161476 (10Andrew) ` Enumerating Boot options... Enumerating Boot options... Done iDRAC Settings: HWC8010: The System Configuration Check operation resulted... [01:30:37] 06cloud-services-team, 10Cloud-VPS, 06tools-infrastructure-team, 10Ceph: cloudcephosd1044 boot issues - https://phabricator.wikimedia.org/T429267#12161482 (10Andrew) The boot manager concurs: ` Driver Health Menu: iDRAC Settings... [01:34:47] RESOLVED: NodeDown: Node cloudcephosd1044 has been down for long. - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Runbooks/NodeDown - https://grafana.wikimedia.org/d/000000377/host-overview?orgId=1&var-server=cloudcephosd1044 - https://alerts.wikimedia.org/?q=alertname%3DNodeDown [01:55:54] RESOLVED: CephClusterInWarning: Ceph cluster in eqiad is in warning status - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Runbooks/CephClusterInWarning - https://grafana.wikimedia.org/d/P1tFnn3Mk/wmcs-ceph-eqiad-health?orgId=1&search=open&tag=ceph&tag=health&tag=WMCS - https://alerts.wikimedia.org/?q=alertname%3DCephClusterInWarning [01:57:22] !log andrew@cloudcumin1001 admin START - Cookbook wmcs.ceph.roll_reboot_osds [02:02:32] !log andrew@cloudcumin1001 admin END (FAIL) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=99) [02:02:55] RESOLVED: SystemdUnitFailed: backup_cinder_volumes.service on cloudbackup2003:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [02:04:09] FIRING: CephClusterInWarning: Ceph cluster in eqiad is in warning status - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Runbooks/CephClusterInWarning - https://grafana.wikimedia.org/d/P1tFnn3Mk/wmcs-ceph-eqiad-health?orgId=1&search=open&tag=ceph&tag=health&tag=WMCS - https://alerts.wikimedia.org/?q=alertname%3DCephClusterInWarning [02:06:04] (03PS14) 10Andrew Bogott: roll_reboot_osds: set maintenance per-osd rather than for the whole run [cloud/wmcs-cookbooks] - 10https://gerrit.wikimedia.org/r/1293792 (https://phabricator.wikimedia.org/T427295) [02:07:37] !log andrew@cloudcumin1001 admin START - Cookbook wmcs.ceph.roll_reboot_osds [02:09:09] RESOLVED: CephClusterInWarning: Ceph cluster in eqiad is in warning status - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Runbooks/CephClusterInWarning - https://grafana.wikimedia.org/d/P1tFnn3Mk/wmcs-ceph-eqiad-health?orgId=1&search=open&tag=ceph&tag=health&tag=WMCS - https://alerts.wikimedia.org/?q=alertname%3DCephClusterInWarning [02:15:09] FIRING: CephClusterInWarning: Ceph cluster in eqiad is in warning status - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Runbooks/CephClusterInWarning - https://grafana.wikimedia.org/d/P1tFnn3Mk/wmcs-ceph-eqiad-health?orgId=1&search=open&tag=ceph&tag=health&tag=WMCS - https://alerts.wikimedia.org/?q=alertname%3DCephClusterInWarning [02:30:09] RESOLVED: CephClusterInWarning: Ceph cluster in eqiad is in warning status - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Runbooks/CephClusterInWarning - https://grafana.wikimedia.org/d/P1tFnn3Mk/wmcs-ceph-eqiad-health?orgId=1&search=open&tag=ceph&tag=health&tag=WMCS - https://alerts.wikimedia.org/?q=alertname%3DCephClusterInWarning [02:35:09] FIRING: CephClusterInWarning: Ceph cluster in eqiad is in warning status - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Runbooks/CephClusterInWarning - https://grafana.wikimedia.org/d/P1tFnn3Mk/wmcs-ceph-eqiad-health?orgId=1&search=open&tag=ceph&tag=health&tag=WMCS - https://alerts.wikimedia.org/?q=alertname%3DCephClusterInWarning [02:40:11] !log andrew@cloudcumin1001 admin END (PASS) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=0) [02:45:09] RESOLVED: CephClusterInWarning: Ceph cluster in eqiad is in warning status - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Runbooks/CephClusterInWarning - https://grafana.wikimedia.org/d/P1tFnn3Mk/wmcs-ceph-eqiad-health?orgId=1&search=open&tag=ceph&tag=health&tag=WMCS - https://alerts.wikimedia.org/?q=alertname%3DCephClusterInWarning [03:13:04] !log andrew@cloudcumin1001 admin START - Cookbook wmcs.ceph.osd.drain_node [03:19:08] 06cloud-services-team, 10Cloud-VPS, 06tools-infrastructure-team, 10Ceph, and 2 others: cloudcephosd1044 boot issues - https://phabricator.wikimedia.org/T429267#12161522 (10Andrew) Dc-ops folks: Please have a look at the backplane issue on this server. It is drained and out of service so you can power it d... [04:00:43] 06cloud-services-team, 10Toolforge, 06tools-platform-team, 13Patch-For-Review: Toolforge web egress with Istio/Envoy seems slow - https://phabricator.wikimedia.org/T425172#12161550 (10Earwig) Hey @taavi, thanks for your help debugging this earlier. It still looks like it's still an issue, unfortunately. Do... [04:56:53] 10Cloud-VPS (Debian Bullseye Deprecation), 06tools-platform-team: Send follow up messages to Cloud VPS project maintainers about Debian Bullseye deprecation - https://phabricator.wikimedia.org/T433334 (10komla) 03NEW [04:57:48] 10Cloud-VPS (Debian Bullseye Deprecation), 06tools-platform-team: Send follow up messages to Cloud VPS project maintainers about Debian Bullseye deprecation - https://phabricator.wikimedia.org/T433334#12161603 (10komla) [05:46:24] 10Striker, 06tools-platform-team, 10Toolhub, 07Documentation: Set up a dedicate and official space for documenation of toolforge tools - https://phabricator.wikimedia.org/T433325#12161615 (10Bugreporter) But this should be linked from Striker. [06:42:55] FIRING: [2x] SystemdUnitFailed: check-private-data.service on clouddb1026:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [06:55:49] !log filippo@cloudcumin1001 tools START - Cookbook wmcs.toolforge.k8s.reboot for tools-k8s-worker-nfs-19, tools-k8s-worker-nfs-32, tools-k8s-worker-nfs-21, tools-k8s-worker-nfs-26, tools-k8s-worker-nfs-13, tools-k8s-worker-nfs-23, tools-k8s-worker-nfs-12, tools-k8s-worker-nfs-22, tools-k8s-worker-nfs-11, tools-k8s-worker-nfs-33, tools-k8s-worker-nfs-17, tools-k8s-worker-nfs-1, tools-k8s-worker-nfs-34, tools-k8s-worker-nfs [06:55:49] -10, tools-k8s-worker-nfs-14, tools-k8s-worker-nfs-16, tools-k8s-worker-nfs-27, tools-k8s-worker-nfs-3, tools-k8s-worker-nfs-24, tools-k8s-worker-nfs-2 (T432325) [06:55:53] T432325: Make sure dumps-nfs mount/umount is propagated inside Toolforge containers - https://phabricator.wikimedia.org/T432325 [07:35:51] 10Cloud-VPS, 06tools-infrastructure-team: mediawiki2latex instance crashed in project collection-alt-renderer - https://phabricator.wikimedia.org/T433244#12161759 (10dhunniger) Just for your information. I reduced the memory usage of the high load process and now everything worked fine. So I really was an... [07:51:16] 10Tool-wmf-openapi-linter, 06Content-Platform-Team, 06MediaWiki-API-Platform-Team (MediaWiki-Interfaces-Sprints), 07OKR-Work, 13Patch-For-Review: Add parameter examples to TransformHandler.php - https://phabricator.wikimedia.org/T432228#12161788 (10daniel) > All operations under transform paths (/v1/tran... [07:52:36] 10Tool-wmf-openapi-linter, 06Content-Platform-Team, 06MediaWiki-API-Platform-Team (MediaWiki-Interfaces-Sprints), 07OKR-Work, 13Patch-For-Review: Add parameter examples to TransformHandler.php - https://phabricator.wikimedia.org/T432228#12161791 (10daniel) Moving to "blocked", we should get input from #c... [08:20:52] (03open) 10countcount: Replace the OAuth login with an optional SOCKS5 proxy to Toolforge [toolforge-repos/flaggedrevspromotioncheck] - 10https://gitlab.wikimedia.org/toolforge-repos/flaggedrevspromotioncheck/-/merge_requests/25 [08:25:42] !log filippo@cloudcumin1001 tools END (PASS) - Cookbook wmcs.toolforge.k8s.reboot (exit_code=0) for tools-k8s-worker-nfs-19, tools-k8s-worker-nfs-32, tools-k8s-worker-nfs-21, tools-k8s-worker-nfs-26, tools-k8s-worker-nfs-13, tools-k8s-worker-nfs-23, tools-k8s-worker-nfs-12, tools-k8s-worker-nfs-22, tools-k8s-worker-nfs-11, tools-k8s-worker-nfs-33, tools-k8s-worker-nfs-17, tools-k8s-worker-nfs-1, tools-k8s-worker-nfs-34, t [08:25:42] ools-k8s-worker-nfs-10, tools-k8s-worker-nfs-14, tools-k8s-worker-nfs-16, tools-k8s-worker-nfs-27, tools-k8s-worker-nfs-3, tools-k8s-worker-nfs-24, tools-k8s-worker-nfs-2 (T432325) [08:25:47] T432325: Make sure dumps-nfs mount/umount is propagated inside Toolforge containers - https://phabricator.wikimedia.org/T432325 [08:31:03] (03update) 10dcaro: kubernetes.py: detect conflict with httproutes created by jobs-api [repos/cloud/toolforge/webservice-cli] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/webservice-cli/-/merge_requests/114 (https://phabricator.wikimedia.org/T431417) (owner: 10raymond-ndibe) [08:34:45] (03update) 10dcaro: global: add publish option to expose job to the internet [repos/cloud/toolforge/jobs-api] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-api/-/merge_requests/262 (https://phabricator.wikimedia.org/T423408) (owner: 10raymond-ndibe) [08:35:09] (03update) 10dcaro: kubernetes.py: detect conflict with httproutes created by jobs-api [repos/cloud/toolforge/webservice-cli] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/webservice-cli/-/merge_requests/114 (https://phabricator.wikimedia.org/T431417) (owner: 10raymond-ndibe) [08:39:57] (03CR) 10David Caro: ceph unset_maintenance: don't block if 'noout' or 'norebalance' are set. (031 comment) [cloud/wmcs-cookbooks] - 10https://gerrit.wikimedia.org/r/1318200 (https://phabricator.wikimedia.org/T427295) (owner: 10Andrew Bogott) [08:40:17] (03update) 10dcaro: global: add publish option to expose job to the internet [repos/cloud/toolforge/jobs-api] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-api/-/merge_requests/262 (https://phabricator.wikimedia.org/T423408) (owner: 10raymond-ndibe) [09:10:20] 10Toolforge (Quota-requests): Request increased quota for statanalyser Toolforge tool - https://phabricator.wikimedia.org/T433184#12162042 (10dcaro) So let me know if you want to increase the quota, or if you want try to make it fit in the current one. [09:23:06] 10Striker, 06tools-platform-team, 10Toolhub, 07Documentation: Set up a dedicate and official space for documenation of toolforge tools - https://phabricator.wikimedia.org/T433325#12162086 (10Ladsgroup) >>! In T433325#12161381, @JJMC89 wrote: >> My preference would be having something like a wiki page with... [09:35:29] 10Cloud-VPS (Project-requests): Request creation of PanoCommons VPS project - https://phabricator.wikimedia.org/T433055#12162176 (10Seb35) In addition of @PanierAvide’s comment, I would say this projet is more about the **promotion of Commons’ images**, in order to share them with the OSM/Panoramax community. Th... [09:40:56] 10Cloud-VPS (Project-requests), 10Tool-Pageviews: Request creation of pageviews VPS project - https://phabricator.wikimedia.org/T433305#12162199 (10fgiunchedi) +1 LGTM, in the future would be good IMHO to explore exactly what went wrong with Toolforge, hosting code there would definitely require less maintenance [09:41:51] 10Cloud-VPS (Project-requests), 10Tool-Pageviews: Request creation of pageviews VPS project - https://phabricator.wikimedia.org/T433305#12162200 (10dcaro) a:03dcaro [09:42:01] 10Cloud-VPS (Project-requests), 10Tool-Pageviews: Request creation of pageviews VPS project - https://phabricator.wikimedia.org/T433305#12162201 (10dcaro) 05Open→03In progress [09:53:05] 10Toolforge (Push-to-Deploy), 06tools-platform-team: [components-api] add job publish support - https://phabricator.wikimedia.org/T433355 (10dcaro) 03NEW [09:54:36] 10Toolforge (Push-to-Deploy), 06tools-platform-team: [components-api] add job publish support - https://phabricator.wikimedia.org/T433355#12162229 (10dcaro) [09:54:38] 10Toolforge (Push-to-Deploy), 06tools-platform-team: [jobs-api] support exposing continuous jobs to the internet - https://phabricator.wikimedia.org/T423408#12162231 (10dcaro) [09:54:40] 10Toolforge (Push-to-Deploy), 06tools-platform-team: [jobs-cli] support publishing continuous job to the internet - https://phabricator.wikimedia.org/T423410#12162230 (10dcaro) [10:05:48] (03update) 10dcaro: support publishing continuous jobs to the internet [repos/cloud/toolforge/jobs-cli] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-cli/-/merge_requests/142 (https://phabricator.wikimedia.org/T423410) (owner: 10raymond-ndibe) [10:17:50] 10Data-Services, 06tools-infrastructure-team, 06tools-platform-team, 13Patch-For-Review: Investigate systemdunitfailed alerts for clouddb hosts - https://phabricator.wikimedia.org/T432745#12162298 (10Marostegui) >>! In T432745#12158668, @fnegri wrote: >> I can be deleted, clouddb1026 only holds s1. > > It... [10:38:12] (03update) 10tlepage: Paginate JobsLogsResponse [repos/cloud/toolforge/logs-api] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/logs-api/-/merge_requests/28 [10:38:37] (03update) 10tlepage: Paginate JobsLogsResponse [repos/cloud/toolforge/logs-api] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/logs-api/-/merge_requests/28 [10:40:01] (03update) 10tlepage: Paginate JobsLogsResponse [repos/cloud/toolforge/logs-api] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/logs-api/-/merge_requests/28 [10:40:36] (03update) 10tlepage: Support logs-api pagination [repos/cloud/toolforge/jobs-cli] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-cli/-/merge_requests/163 (https://phabricator.wikimedia.org/T428253) [10:43:10] FIRING: [2x] SystemdUnitFailed: check-private-data.service on clouddb1026:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [10:49:23] (03update) 10tlepage: Support logs-api pagination [repos/cloud/toolforge/jobs-cli] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-cli/-/merge_requests/163 (https://phabricator.wikimedia.org/T428253) [10:50:02] (03update) 10tlepage: Support logs-api pagination [repos/cloud/toolforge/jobs-cli] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-cli/-/merge_requests/163 (https://phabricator.wikimedia.org/T428253) [10:50:29] (03update) 10tlepage: Support logs-api pagination [repos/cloud/toolforge/jobs-cli] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-cli/-/merge_requests/163 (https://phabricator.wikimedia.org/T428253) [10:50:51] (03update) 10tlepage: Support logs-api pagination [repos/cloud/toolforge/jobs-cli] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-cli/-/merge_requests/163 (https://phabricator.wikimedia.org/T428253) [10:54:15] (03update) 10tlepage: Support logs-api pagination [repos/cloud/toolforge/jobs-cli] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-cli/-/merge_requests/163 (https://phabricator.wikimedia.org/T428253) [10:54:30] (03update) 10tlepage: Support logs-api pagination [repos/cloud/toolforge/jobs-cli] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-cli/-/merge_requests/163 (https://phabricator.wikimedia.org/T428253) [11:47:02] (03update) 10dcaro: global: add publish option to expose job to the internet [repos/cloud/toolforge/jobs-api] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-api/-/merge_requests/262 (https://phabricator.wikimedia.org/T423408) (owner: 10raymond-ndibe) [11:48:18] (03update) 10dcaro: support --webservice option [repos/cloud/toolforge/jobs-cli] (publish_continuous_job_to_internet) - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-cli/-/merge_requests/143 (https://phabricator.wikimedia.org/T428898) (owner: 10raymond-ndibe) [11:48:54] (03update) 10dcaro: test continuous job webservice [repos/cloud/toolforge/toolforge-deploy] (test_continuous_job_publish) - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1131 (https://phabricator.wikimedia.org/T348755) (owner: 10raymond-ndibe) [11:55:16] !log dcaro@acme admin START - Cookbook wmcs.openstack.get_project_for_proxy [11:55:20] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [11:55:47] !log dcaro@acme admin END (FAIL) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=99) [11:55:50] !log dcaro@acme admin START - Cookbook wmcs.openstack.get_project_for_proxy [11:55:51] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [11:55:53] !log dcaro@acme admin END (FAIL) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=99) [11:55:54] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [11:55:58] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [11:57:23] !log dcaro@acme admin START - Cookbook wmcs.openstack.get_project_for_proxy [11:57:27] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [12:00:14] !log dcaro@acme admin END (ERROR) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=1) [12:00:17] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [12:05:57] !log dcaro@acme admin START - Cookbook wmcs.openstack.get_project_for_proxy [12:06:02] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [12:06:08] !log dcaro@acme admin END (FAIL) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=99) [12:06:12] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [12:08:03] !log dcaro@acme admin START - Cookbook wmcs.openstack.get_project_for_proxy [12:08:06] !log dcaro@acme admin END (FAIL) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=99) [12:08:07] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [12:08:11] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [12:12:28] !log dcaro@acme admin START - Cookbook wmcs.openstack.get_project_for_proxy [12:12:31] !log dcaro@acme admin END (FAIL) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=99) [12:12:33] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [12:12:36] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [12:22:55] RESOLVED: [2x] SystemdUnitFailed: check-private-data.service on clouddb1026:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [12:27:38] 10Data-Services, 06tools-infrastructure-team, 06tools-platform-team: Investigate systemdunitfailed alerts for clouddb hosts - https://phabricator.wikimedia.org/T432745#12162613 (10fnegri) 05In progress→03Resolved Thanks, I restarted the unit with `systemctl start check-private-data.service` and the a... [12:54:11] !log dcaro@acme admin START - Cookbook wmcs.openstack.get_project_for_proxy [12:54:17] !log dcaro@acme admin END (FAIL) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=99) [12:54:17] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [12:54:20] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [12:54:25] !log dcaro@acme admin START - Cookbook wmcs.openstack.get_project_for_proxy [12:54:28] !log dcaro@acme admin END (FAIL) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=99) [12:54:30] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [12:54:34] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [12:55:14] !log dcaro@acme admin START - Cookbook wmcs.openstack.get_project_for_proxy [12:55:17] !log dcaro@acme admin END (FAIL) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=99) [12:55:18] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [12:55:21] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [12:59:20] !log dcaro@acme admin START - Cookbook wmcs.openstack.get_project_for_proxy [12:59:25] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [12:59:26] !log dcaro@acme admin END (FAIL) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=99) [12:59:29] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [12:59:57] !log dcaro@acme admin START - Cookbook wmcs.openstack.get_project_for_proxy [13:00:01] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [13:00:07] !log dcaro@acme admin END (PASS) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=0) [13:00:11] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [13:00:48] !log dcaro@acme admin START - Cookbook wmcs.openstack.get_project_for_proxy [13:00:51] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [13:00:55] !log dcaro@acme admin END (PASS) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=0) [13:00:58] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [13:02:28] !log dcaro@acme admin START - Cookbook wmcs.openstack.get_project_for_proxy [13:02:31] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [13:02:34] !log dcaro@acme admin END (PASS) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=0) [13:02:39] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [13:03:23] !log dcaro@acme admin START - Cookbook wmcs.openstack.get_project_for_proxy [13:03:26] (03approved) 10tlepage: clinic_duty: add pinging script with dcaro only [repos/cloud/wmcs/utils] - 10https://gitlab.wikimedia.org/repos/cloud/wmcs/utils/-/merge_requests/6 (owner: 10dcaro) [13:03:26] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [13:03:32] !log dcaro@acme admin END (FAIL) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=99) [13:03:37] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [13:03:49] (03approved) 10tlepage: dotfiles: remove old dotfiles [repos/cloud/toolforge/lima-kilo] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/lima-kilo/-/merge_requests/331 (owner: 10dcaro) [13:03:52] !log dcaro@acme admin START - Cookbook wmcs.openstack.get_project_for_proxy [13:03:57] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [13:04:10] !log dcaro@acme admin END (ERROR) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=1) [13:04:14] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [13:04:18] !log dcaro@acme admin START - Cookbook wmcs.openstack.get_project_for_proxy [13:04:23] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [13:04:29] !log dcaro@acme admin END (PASS) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=0) [13:04:34] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [13:04:46] !log dcaro@acme admin START - Cookbook wmcs.openstack.get_project_for_proxy [13:04:51] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [13:04:57] !log dcaro@acme admin END (PASS) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=0) [13:05:00] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [13:05:54] !log dcaro@acme admin START - Cookbook wmcs.openstack.get_project_for_proxy [13:05:57] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [13:06:08] !log dcaro@acme admin END (PASS) - Cookbook wmcs.openstack.get_project_for_proxy (exit_code=0) [13:06:11] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [13:11:43] !log andrew@cloudcumin1001 admin START - Cookbook wmcs.ceph.roll_reboot_mons (T431659) [13:13:33] (03PS1) 10David Caro: get_project_for_proxy: cookbook to get a project given a web proxy [cloud/wmcs-cookbooks] - 10https://gerrit.wikimedia.org/r/1318716 [13:14:00] !log andrew@cloudcumin1001 admin END (FAIL) - Cookbook wmcs.ceph.osd.drain_node (exit_code=99) [13:14:29] PROBLEM - Host cloudcephmon1004 is DOWN: PING CRITICAL - Packet loss = 100% [13:15:25] (03PS2) 10David Caro: get_project_for_proxy: cookbook to get a project given a web proxy [cloud/wmcs-cookbooks] - 10https://gerrit.wikimedia.org/r/1318716 [13:16:02] (03CR) 10David Caro: "Tested locally, passin run:" [cloud/wmcs-cookbooks] - 10https://gerrit.wikimedia.org/r/1318716 (owner: 10David Caro) [13:16:26] (03Abandoned) 10Andrew Bogott: ceph unset_maintenance: don't block if 'noout' or 'norebalance' are set. [cloud/wmcs-cookbooks] - 10https://gerrit.wikimedia.org/r/1318200 (https://phabricator.wikimedia.org/T427295) (owner: 10Andrew Bogott) [13:17:30] RECOVERY - Host cloudcephmon1004 is UP: PING OK - Packet loss = 0%, RTA = 0.34 ms [13:18:34] PROBLEM - Host cloudcephmon1005 is DOWN: PING CRITICAL - Packet loss = 100% [13:19:10] FIRING: CephClusterInWarning: Ceph cluster in eqiad is in warning status - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Runbooks/CephClusterInWarning - https://grafana.wikimedia.org/d/P1tFnn3Mk/wmcs-ceph-eqiad-health?orgId=1&search=open&tag=ceph&tag=health&tag=WMCS - https://alerts.wikimedia.org/?q=alertname%3DCephClusterInWarning [13:20:04] RECOVERY - Host cloudcephmon1005 is UP: PING OK - Packet loss = 0%, RTA = 0.40 ms [13:22:13] PROBLEM - Host cloudcephmon1006 is DOWN: PING CRITICAL - Packet loss = 100% [13:23:30] 10Tool-translatetagger, 06Indic-TechCom, 10MediaWiki-extensions-Translate, 07Community-collaboration, and 4 others: Integrate TranslateTagger functions/workflow into Special:PagePreparation of Translate Extension - https://phabricator.wikimedia.org/T430257#12162894 (10abi_) We can't test this on production... [13:24:03] RECOVERY - Host cloudcephmon1006 is UP: PING OK - Packet loss = 0%, RTA = 0.37 ms [13:24:41] !log andrew@cloudcumin1001 admin END (PASS) - Cookbook wmcs.ceph.roll_reboot_mons (exit_code=0) (T431659) [13:27:07] (03update) 10raymond-ndibe: global: add publish option to expose job to the internet [repos/cloud/toolforge/jobs-api] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-api/-/merge_requests/262 (https://phabricator.wikimedia.org/T423408) [13:29:09] RESOLVED: CephClusterInWarning: Ceph cluster in eqiad is in warning status - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Runbooks/CephClusterInWarning - https://grafana.wikimedia.org/d/P1tFnn3Mk/wmcs-ceph-eqiad-health?orgId=1&search=open&tag=ceph&tag=health&tag=WMCS - https://alerts.wikimedia.org/?q=alertname%3DCephClusterInWarning [13:29:39] (03update) 10raymond-ndibe: jobs-api: test continuous job publish [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1116 (https://phabricator.wikimedia.org/T388092) [13:30:23] (03CR) 10Andrew Bogott: "This now removes noout/norebalance a few seconds /before/ the OSDs come online after a reboot. I was worried this would cause thrashing bu" [cloud/wmcs-cookbooks] - 10https://gerrit.wikimedia.org/r/1293792 (https://phabricator.wikimedia.org/T427295) (owner: 10Andrew Bogott) [13:37:14] (03update) 10arian-ar: Set Secure flag on session cookie [toolforge-repos/sentinel] - 10https://gitlab.wikimedia.org/toolforge-repos/sentinel/-/merge_requests/13 (owner: 10priyankar22) [13:44:43] (03CR) 10David Caro: "This should probably use the web proxy API exposed as openstack api already (either directly or using the wmcs-webproxy cli, preferably th" [cloud/wmcs-cookbooks] - 10https://gerrit.wikimedia.org/r/1318716 (owner: 10David Caro) [13:45:31] (03CR) 10David Caro: "Note that this means adding a new endpoint to be able to either extract all proxies in all projects, or to search a project with the given" [cloud/wmcs-cookbooks] - 10https://gerrit.wikimedia.org/r/1318716 (owner: 10David Caro) [13:45:51] (03update) 10dcaro: dotfiles: remove old dotfiles [repos/cloud/toolforge/lima-kilo] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/lima-kilo/-/merge_requests/331 [13:45:57] (03merge) 10dcaro: dotfiles: remove old dotfiles [repos/cloud/toolforge/lima-kilo] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/lima-kilo/-/merge_requests/331 [13:49:39] 10Toolforge, 06tools-platform-team, 07OKR-Work: [loki,alloy,logs-api] reconfigure the labels and add metadata - https://phabricator.wikimedia.org/T432969#12163010 (10TLepage-WMF) 05Open→03In progress a:05dcaro→03TLepage-WMF [13:50:42] (03update) 10raymond-ndibe: jobs-api: test continuous job publish [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1116 (https://phabricator.wikimedia.org/T388092) [13:50:49] (03update) 10raymond-ndibe: jobs-api: test continuous job publish [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1116 (https://phabricator.wikimedia.org/T388092) [14:23:32] (03update) 10arian-ar: Remove unused SECRET_KEY setting [toolforge-repos/sentinel] - 10https://gitlab.wikimedia.org/toolforge-repos/sentinel/-/merge_requests/15 (owner: 10ebenezerrao) [14:23:33] (03update) 10arian-ar: Remove unused SECRET_KEY setting [toolforge-repos/sentinel] - 10https://gitlab.wikimedia.org/toolforge-repos/sentinel/-/merge_requests/15 (owner: 10ebenezerrao) [14:25:35] (03update) 10arian-ar: Remove unused SECRET_KEY setting [toolforge-repos/sentinel] - 10https://gitlab.wikimedia.org/toolforge-repos/sentinel/-/merge_requests/15 (owner: 10ebenezerrao) [14:25:56] (03merge) 10arian-ar: Remove unused SECRET_KEY setting [toolforge-repos/sentinel] - 10https://gitlab.wikimedia.org/toolforge-repos/sentinel/-/merge_requests/15 (owner: 10ebenezerrao) [14:26:15] 10Tool-sentinel: Remove or wire unused SECRET_KEY setting - https://phabricator.wikimedia.org/T431905#12163221 (10Arian_Ar) 05Open→03Resolved [14:31:56] 10Tool-sentinel: Remove redundant local import json in mediawiki.py - https://phabricator.wikimedia.org/T431907#12163274 (10Arian_Ar) 05Open→03Resolved Hello, @EbenezerRao Confirmed. this is already fixed on main. mediawiki.py only has the module-level`import json`; the local imports were removed with t... [14:33:58] FIRING: NeutronAgentDownForLong: Neutron neutron-openvswitch-agent on cloudvirt1048 has been down for more than 2h - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Troubleshooting#Networking_failures - https://grafana.wikimedia.org/d/wKnDJf97z/wmcs-neutron-eqiad1 - https://alerts.wikimedia.org/?q=alertname%3DNeutronAgentDownForLong [14:33:58] FIRING: NeutronAgentDown: Neutron neutron-openvswitch-agent on cloudvirt1048 is down - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Troubleshooting#Networking_failures - https://grafana.wikimedia.org/d/wKnDJf97z/wmcs-neutron-eqiad1 - https://alerts.wikimedia.org/?q=alertname%3DNeutronAgentDown [14:43:37] 10Tool-sentinel: Remove redundant local import json in mediawiki.py - https://phabricator.wikimedia.org/T431907#12163351 (10EbenezerRao) Hey Arian_Ar Yep thanks for acknowledging as well [14:48:11] !log andrew@cloudcumin1001 testlabs START - Cookbook wmcs.nfs.add_server [14:58:15] !log andrew@cloudcumin1001 testlabs END (PASS) - Cookbook wmcs.nfs.add_server (exit_code=0) [15:16:39] (03approved) 10raymond-ndibe: global: add publish option to expose job to the internet [repos/cloud/toolforge/jobs-api] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-api/-/merge_requests/262 (https://phabricator.wikimedia.org/T423408) [15:16:43] (03update) 10raymond-ndibe: global: add publish option to expose job to the internet [repos/cloud/toolforge/jobs-api] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-api/-/merge_requests/262 (https://phabricator.wikimedia.org/T423408) [15:30:42] 10Tool-wmf-openapi-linter, 06MediaWiki-API-Platform-Team, 07OKR-Work: Deploy the OpenAPI linter in CI - https://phabricator.wikimedia.org/T422920#12163619 (10OWresch-WMF) [15:34:07] 10Toolforge, 06tools-infrastructure-team, 06tools-platform-team: [infra] Allow some opt-in scraping solution for tools - https://phabricator.wikimedia.org/T433395 (10dcaro) 03NEW [15:38:20] 10Toolforge, 06tools-platform-team: [components-api] Allow setting opt-in flags for special scraping control measures - https://phabricator.wikimedia.org/T433397 (10dcaro) 03NEW [15:39:16] 10Toolforge, 06tools-infrastructure-team, 06tools-platform-team: [infra] Allow some opt-in scraping solution for tools - https://phabricator.wikimedia.org/T433395#12163742 (10dcaro) @taavi How when would you want to access the information about the tools that did opt-in? (to design the endpoints or the syste... [15:40:41] 06cloud-services-team, 10Toolforge, 06tools-platform-team: [toolsdb] Transaction History Length growing too much - https://phabricator.wikimedia.org/T428139#12163756 (10fnegri) Improved a little but still very high: {F96301434} Long transactions are still mostly from `mixnmatch`: `lang=mysql MariaDB [(non... [15:43:29] (03update) 10dcaro: Support logs-api pagination [repos/cloud/toolforge/jobs-cli] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-cli/-/merge_requests/163 (https://phabricator.wikimedia.org/T428253) (owner: 10tlepage) [15:43:38] (03update) 10dcaro: Support logs-api pagination [repos/cloud/toolforge/jobs-cli] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-cli/-/merge_requests/163 (https://phabricator.wikimedia.org/T428253) (owner: 10tlepage) [15:54:22] 10Toolforge, 06tools-platform-team: [jobs-cli] broken pipe error when using `toolforge jobs logs | head` - https://phabricator.wikimedia.org/T433405 (10dcaro) 03NEW [15:54:24] (03update) 10dcaro: Support logs-api pagination [repos/cloud/toolforge/jobs-cli] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-cli/-/merge_requests/163 (https://phabricator.wikimedia.org/T428253) (owner: 10tlepage) [15:56:13] (03update) 10kineticpelagic: Proposal: UI changes to linting results [toolforge-repos/wmf-openapi-linter] - 10https://gitlab.wikimedia.org/toolforge-repos/wmf-openapi-linter/-/merge_requests/27 (owner: 10kbach) [15:56:28] (03approved) 10kineticpelagic: Proposal: UI changes to linting results [toolforge-repos/wmf-openapi-linter] - 10https://gitlab.wikimedia.org/toolforge-repos/wmf-openapi-linter/-/merge_requests/27 (owner: 10kbach) [15:59:55] (03approved) 10aghirelli: Proposal: UI changes to linting results [toolforge-repos/wmf-openapi-linter] - 10https://gitlab.wikimedia.org/toolforge-repos/wmf-openapi-linter/-/merge_requests/27 (owner: 10kbach) [16:00:59] (03merge) 10kineticpelagic: Proposal: UI changes to linting results [toolforge-repos/wmf-openapi-linter] - 10https://gitlab.wikimedia.org/toolforge-repos/wmf-openapi-linter/-/merge_requests/27 (owner: 10kbach) [16:07:42] 10Toolforge, 06tools-platform-team, 07Epic, 07OKR-Work: [hypothesis] ST5.4.1 Extend logging capabilities - https://phabricator.wikimedia.org/T432564#12163937 (10aputhin) 05Open→03In progress [16:08:01] 10Toolforge, 06tools-platform-team, 07Epic, 07OKR-Work: ST5.4.2 Toolforge Alerting System - https://phabricator.wikimedia.org/T432860#12163950 (10aputhin) p:05Triage→03High [16:08:15] 10Toolforge, 06tools-platform-team, 07OKR-Work: Verify if we can define alert conditions in Prometheus - https://phabricator.wikimedia.org/T432863#12163951 (10aputhin) p:05Triage→03High [16:11:31] 10Toolforge, 06tools-infrastructure-team, 13Patch-For-Review: Toolforge web egress with Istio/Envoy seems slow - https://phabricator.wikimedia.org/T425172#12163960 (10aputhin) [16:11:49] 10Toolforge, 06tools-platform-team, 07OKR-Work: Provide a way for users to enable and disable alerts - https://phabricator.wikimedia.org/T432864#12163962 (10fnegri) Note: we can probably use a similar interface for this and for {T433397}. [16:38:53] (03update) 10bd808: Add visualization [toolforge-repos/toolviews] - 10https://gitlab.wikimedia.org/toolforge-repos/toolviews/-/merge_requests/13 (owner: 10lokal-profil) [16:39:06] (03update) 10bd808: Add visualization [toolforge-repos/toolviews] - 10https://gitlab.wikimedia.org/toolforge-repos/toolviews/-/merge_requests/13 (owner: 10lokal-profil) [16:40:06] (03update) 10dcaro: Paginate JobsLogsResponse [repos/cloud/toolforge/logs-api] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/logs-api/-/merge_requests/28 (owner: 10tlepage) [16:40:30] (03update) 10dcaro: Support logs-api pagination [repos/cloud/toolforge/jobs-cli] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-cli/-/merge_requests/163 (https://phabricator.wikimedia.org/T428253) (owner: 10tlepage) [16:50:09] FIRING: CephClusterInWarning: Ceph cluster in eqiad is in warning status - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Runbooks/CephClusterInWarning - https://grafana.wikimedia.org/d/P1tFnn3Mk/wmcs-ceph-eqiad-health?orgId=1&search=open&tag=ceph&tag=health&tag=WMCS - https://alerts.wikimedia.org/?q=alertname%3DCephClusterInWarning [16:51:37] 10Tool-curator: Curator: force _ for illegal characters - https://phabricator.wikimedia.org/T433415 (10PantheraLeo1359531) 03NEW [17:20:09] RESOLVED: CephClusterInWarning: Ceph cluster in eqiad is in warning status - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Runbooks/CephClusterInWarning - https://grafana.wikimedia.org/d/P1tFnn3Mk/wmcs-ceph-eqiad-health?orgId=1&search=open&tag=ceph&tag=health&tag=WMCS - https://alerts.wikimedia.org/?q=alertname%3DCephClusterInWarning [17:20:25] !log andrew@cloudcumin1001 testlabs START - Cookbook wmcs.nfs.migrate_service [17:21:00] !log andrew@cloudcumin1001 testlabs END (FAIL) - Cookbook wmcs.nfs.migrate_service (exit_code=99) [17:29:57] 10Tool-wikimedia-attribution, 06MediaWiki-API-Platform-Team, 10MediaWiki-REST-API, 07OKR-Work: Include contributor_count in Attribution API responses - https://phabricator.wikimedia.org/T429834#12164433 (10pmiazga) [17:39:05] !log andrew@cloudcumin1001 testlabs START - Cookbook wmcs.nfs.migrate_service [17:39:15] !log andrew@cloudcumin1001 testlabs END (FAIL) - Cookbook wmcs.nfs.migrate_service (exit_code=99) [17:40:41] !log andrew@cloudcumin1001 testlabs START - Cookbook wmcs.nfs.migrate_service [17:42:46] !log andrew@cloudcumin1001 testlabs END (PASS) - Cookbook wmcs.nfs.migrate_service (exit_code=0) [17:43:51] 06cloud-services-team, 10Data-Services, 10Toolforge, 06tools-platform-team, 07Documentation: Restructure and improve content for: https://wikitech.wikimedia.org/wiki/Help:Toolforge/Database - https://phabricator.wikimedia.org/T232404#12164490 (10fnegri) I did some further tweaks to https://wikitech.wikim... [17:53:36] 10Tool-wikinewsie: Algorithmic feed mix for Wikinewsie - https://phabricator.wikimedia.org/T433425 (10Pharos) 03NEW [17:57:30] 10Toolforge, 06tools-infrastructure-team, 06tools-platform-team: [infra] Allow some opt-in scraping solution for tools - https://phabricator.wikimedia.org/T433395#12164572 (10taavi) In the end it'll need to be a HAProxy map file on the proxy hosts, with an entry per domain name that's opted in to this. So I... [17:57:52] 10Cloud-VPS (Debian Bullseye Deprecation): Migrate WMCS-managed NFS servers off of Bullseye - https://phabricator.wikimedia.org/T401812#12164573 (10Andrew) [18:00:23] 10Cloud-VPS (Debian Bullseye Deprecation): Migrate WMCS-managed NFS servers off of Bullseye - https://phabricator.wikimedia.org/T401812#12164575 (10Andrew) [18:00:53] 10Cloud-VPS (Debian Bullseye Deprecation), 06tools-infrastructure-team: Replace scratch-1.cloudinfra-nfs with a trixie VM - https://phabricator.wikimedia.org/T433426 (10Andrew) 03NEW [18:00:59] 10Cloud-VPS (Debian Bullseye Deprecation), 06tools-infrastructure-team: Replace scratch-1.cloudinfra-nfs with a trixie VM - https://phabricator.wikimedia.org/T433426#12164592 (10Andrew) a:03Andrew [18:01:59] 10Cloud-VPS (Debian Bullseye Deprecation), 06tools-infrastructure-team: Replace scratch-1.cloudinfra-nfs with a trixie VM - https://phabricator.wikimedia.org/T433426#12164595 (10Andrew) Tests suggest that we can do this without rebooting clients, but access will be weird ('Permission denied') for a few minutes... [18:02:38] 10Toolforge (Quota-requests): Request increased quota for statanalyser Toolforge tool - https://phabricator.wikimedia.org/T433184#12164596 (10Leaderboard) Let's try increasing it to 8 - 10 GB for now. If issues happen we'll see. [18:04:21] !log andrew@cloudcumin1001 cloudinfra-nfs START - Cookbook wmcs.nfs.add_server [18:16:16] !log andrew@cloudcumin1001 cloudinfra-nfs END (PASS) - Cookbook wmcs.nfs.add_server (exit_code=0) [18:17:59] (03update) 10renovatebot: Lock file maintenance [toolforge-repos/wiki-mail-verify] - 10https://gitlab.wikimedia.org/toolforge-repos/wiki-mail-verify/-/merge_requests/18 [18:18:01] (03update) 10renovatebot: Lock file maintenance [toolforge-repos/wiki-mail-verify] - 10https://gitlab.wikimedia.org/toolforge-repos/wiki-mail-verify/-/merge_requests/18 [18:37:46] 10Cloud-VPS (Debian Bullseye Deprecation), 06tools-infrastructure-team: Replace scratch-1.cloudinfra-nfs with a trixie VM - https://phabricator.wikimedia.org/T433426#12164658 (10Andrew) This migration will look like: ` sudo cookbook wmcs.nfs.migrate_service --from-host-id 2fd8eb82-33ec-4060-91c6-cc0a90de899... [20:51:19] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12165090 (10VRiley-WMF) Sure thing, I'm planning on this tomorrow. Thank you! [21:26:04] FIRING: PuppetCertificateAboutToExpire: Puppet CA certificate gitlab-runners-puppetmaster-01.gitlab-runners.eqiad1.wikimedia.cloud is about to expire in 26d 23h 58m 16s - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Runbooks/PuppetCertificateAboutToExpire - https://prometheus-alerts.wmcloud.org/?q=alertname%3DPuppetCertificateAboutToExpire [21:26:17] FIRING: JobUnavailable: Reduced availability for job maintain_dbusers_eqiad in cloud@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [21:31:04] FIRING: PuppetCertificateAboutToExpire: Puppet CA certificate Puppet CA: gitlab-runners-puppetmaster-01.gitlab-runners.eqiad1.wikimedia.cloud is about to expire in 26d 23h 58m 21s - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Runbooks/PuppetCertificateAboutToExpire - https://prometheus-alerts.wmcloud.org/?q=alertname%3DPuppetCertificateAboutToExpire [21:31:17] RESOLVED: JobUnavailable: Reduced availability for job maintain_dbusers_eqiad in cloud@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [22:13:45] 10Cloud-VPS, 06tools-infrastructure-team, 13Patch-For-Review: Network unavailable on a few VMs - https://phabricator.wikimedia.org/T432426#12165313 (10BLiviero-WMF) @Andrew has logging-logstash-04 remained stable or do we have logging data of a new failure? additionally were there any more instances of thi... [22:17:36] (03update) 10renovatebot: Lock file maintenance [toolforge-repos/wiki-mail-verify] - 10https://gitlab.wikimedia.org/toolforge-repos/wiki-mail-verify/-/merge_requests/18 [22:36:45] (03update) 10raymond-ndibe: support publishing continuous jobs to the internet [repos/cloud/toolforge/jobs-cli] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-cli/-/merge_requests/142 (https://phabricator.wikimedia.org/T423410) [22:40:52] (03open) 10raymond-ndibe: toolforge-common.yaml.j2: add publish_root_domain and publish_proto [repos/cloud/toolforge/lima-kilo] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/lima-kilo/-/merge_requests/332 (https://phabricator.wikimedia.org/T423410) [22:41:44] (03update) 10raymond-ndibe: toolforge-common.yaml.j2: add publish_root_domain and publish_proto [repos/cloud/toolforge/lima-kilo] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/lima-kilo/-/merge_requests/332 (https://phabricator.wikimedia.org/T423410) [22:42:20] (03update) 10raymond-ndibe: support publishing continuous jobs to the internet [repos/cloud/toolforge/jobs-cli] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-cli/-/merge_requests/142 (https://phabricator.wikimedia.org/T423410) [22:42:55] (03update) 10raymond-ndibe: support publishing continuous jobs to the internet [repos/cloud/toolforge/jobs-cli] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-cli/-/merge_requests/142 (https://phabricator.wikimedia.org/T423410) [22:43:00] (03update) 10raymond-ndibe: support publishing continuous jobs to the internet [repos/cloud/toolforge/jobs-cli] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-cli/-/merge_requests/142 (https://phabricator.wikimedia.org/T423410) [23:13:03] (03update) 10raymond-ndibe: kubernetes.py: detect conflict with httproutes created by jobs-api [repos/cloud/toolforge/webservice-cli] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/webservice-cli/-/merge_requests/114 (https://phabricator.wikimedia.org/T431417) [23:29:36] (03update) 10raymond-ndibe: kubernetes.py: detect conflict with httproutes created by jobs-api [repos/cloud/toolforge/webservice-cli] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/webservice-cli/-/merge_requests/114 (https://phabricator.wikimedia.org/T431417) [23:30:06] (03update) 10raymond-ndibe: kubernetes.py: detect conflict with httproutes created by jobs-api [repos/cloud/toolforge/webservice-cli] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/webservice-cli/-/merge_requests/114 (https://phabricator.wikimedia.org/T431417) [23:42:29] 10Tool-wikinewsie: Newsletter mode for Wikinewsie - https://phabricator.wikimedia.org/T432501#12165504 (10Pharos) A proposal along similar lines was made by @Aced at the same "Unpopular" lightning talks session at Wikimania 2026, where the Wikinewsie tool was also first shared: [[ https://docs.google.com/present... [23:52:20] 10Tool-wikinewsie, 03Wikimania-Hackathon-2026: Portal:Current events mode for Wikinewsie - https://phabricator.wikimedia.org/T432971#12165511 (10Pharos) Could be integrated withe the Internet Archive's [[ https://wayback-labs.sf.archive.org/ | Wayback Machine Labs ]], and particularly [[ http://208.70.27.132/p...