Monday, 2026-06-15

-@gerrit:opendev.org- Steven van der Schoot proposed: [opendev/git-review] 993318: Unquote percent-encoded usernames parsed from remote URLs https://review.opendev.org/c/opendev/git-review/+/99331809:58
@priteau:matrix.orgWe are seeing many post failures in CI right now. Anything broken?11:14
@priteau:matrix.orgIt's also been mentioned by someone else on IRC in #openstack-infra11:15
@mnasiadka:matrix.orgIt seems rax_iad Swift provider has some problems11:33
@mnasiadka:matrix.orgOr maybe not, because another job failed on ovh_bhs11:36
@mnasiadka:matrix.orgAh, majority of the failures is ovh_gra and ovh_bhs11:38
@mnasiadka:matrix.orgSeems there is something going on with OVH Object Storage: https://public-cloud.status-ovhcloud.com/incidents/7jjsc5hq8tnn11:38
-@gerrit:opendev.org- Michal Nasiadka proposed: [opendev/base-jobs] 993335: Disable job log uploads to ovh swift https://review.opendev.org/c/opendev/base-jobs/+/99333511:40
@priteau:matrix.orgStart time : 13/05/2026 00:00 UTC though…11:43
@priteau:matrix.orgIt might be a new incident not yet on their status page11:43
@mnasiadka:matrix.orgwell, that one is not closed ;-)11:46
@mnasiadka:matrix.orgAlthough I agree it might be some old one11:46
@sean-k-mooney:matrix.orgit might be a good idea to send out a status message to tell peopel to avoid rechecks until that is resolved12:00
@sean-k-mooney:matrix.orgi know its does nto alwasy fail right now but its failing enought htat any one build set is unlikely to pass 12:00
-@gerrit:opendev.org- Steven van der Schoot proposed: [opendev/git-review] 993339: Escape/unescape usernames with special characters in remote URLs https://review.opendev.org/c/opendev/git-review/+/99333912:01
@mnasiadka:matrix.orgLet me force merge that patch and send a notice12:09
-@gerrit:opendev.org- mnasiadka.admin merged on behalf of Michal Nasiadka: [opendev/base-jobs] 993335: Disable job log uploads to ovh swift https://review.opendev.org/c/opendev/base-jobs/+/99333512:34
@mnasiadka:matrix.org#status notice Recent POST_FAILURE job results with no logs were due to upload errors in one of our providers, which has been temporarily disabled now so rechecking those should be safe12:41
@status:opendev.org@mnasiadka:matrix.org: sending notice12:41
@priteau:matrix.orgThanks for the prompt fix!12:45
-@status:opendev.org- NOTICE: Recent POST_FAILURE job results with no logs were due to upload errors in one of our providers, which has been temporarily disabled now so rechecking those should be safe12:45
@status:opendev.org@mnasiadka:matrix.org: finished sending notice12:45
@mnasiadka:matrix.orgphew, all done12:46
@fungicide:matrix.orgthanks mnasiadka!13:10
@fungicide:matrix.orglooks like i showed up too late to catch all the fun13:10
@mnasiadka:matrix.orgfungi: No problem! That was a good degree of learning how to force merge a patch ;)13:10
@fungicide:matrix.orghopefully the documentation was straightforward, though it's a lot of copy/paste of gerrit ssh api commands13:18
@fungicide:matrix.orgmostly as a means of being able to use separate admin accounts that don't have openid logins, for added safety13:19
@mnasiadka:matrix.orgyeah, I already had an admin account, so that wasn't a problem - but yes, I have a tendency of trying to understand what I'm copy pasting ;-)13:50
@fungicide:matrix.orgas you should, of course, i'm just hoping the text around those copy-paste examples was sufficient explanation13:53
@fungicide:matrix.orgif not, we should improve it13:53
@fungicide:matrix.orgi'm going to disappear for an early lunch, but everyone check that topics you want covered are in the agenda and i'll send it out to service-discuss later today: https://wiki.openstack.org/wiki/Meetings/InfraTeamMeeting#Agenda_for_next_meeting14:35
@fungicide:matrix.orgstatic03 is getting beaten up pretty hard at the moment, load average over 2k for at least the past 15 minutes, looks like. server-status indicates most workers are busy writing responses to clients14:37
@fungicide:matrix.orgif this persists, it's worth seeing whether there are new crawler signatures we can add or if we need to consider other mitigations like bringing a larger replacement similar to static04 back (would need to create it from scratch again since i did eventually delete the old one)14:40
@fungicide:matrix.orgaccording to top, the cpus are still fairly idle, but are spending the bulk of their ticks on "system" tasks14:42
@fungicide:matrix.orgthat may be indicative of interfacing with the openafs driver14:43
@priteau:matrix.orgYes, https://tarballs.openstack.org is down from here 😢14:59
-@gerrit:opendev.org- Monty Taylor https://matrix.to/#/@mordred:inaugust.com proposed: [openstack/project-config] 993214: Add repos for drizzle website and sysadmin automation https://review.opendev.org/c/openstack/project-config/+/99321415:35
@fungicide:matrix.orgPierre Riteau: should be back up again now. some crawler bot army was crushing it while i was at lunch, load average on the server was up over 2.5k that i saw16:38
@fungicide:matrix.orgseems like they've moved on (for now anyway)16:38
@jim:acmegating.comit looks like the weekend zuul restart is stuck mid-upgrade16:56
@jim:acmegating.comit's because there is a buildset with a paused job waiting on other jobs which are waiting on arm nodes17:09
@jim:acmegating.comhttps://zuul.opendev.org/t/openstack/status?project=openstack%2Fopenstack-helm-images17:09
@jim:acmegating.commy understanding is that our arm provider is currently offline indefinitely17:10
@fungicide:matrix.orgyeah, seems like the situation at osuosl hasn't changed yet17:10
@jim:acmegating.comi think at this point we may want to consider having zuul fail jobs rather than queueing indefinitely17:11
@fungicide:matrix.orglast week was finals so we were hesitant to prod them, but it's definitely time to check in17:11
-@gerrit:opendev.org- James E. Blair https://matrix.to/#/@jim:acmegating.com proposed: [opendev/zuul-providers] 993417: Remove all labels from osuosl https://review.opendev.org/c/opendev/zuul-providers/+/99341717:12
@jim:acmegating.comfungi: mnasiadka ^ 17:13
@fungicide:matrix.orgi've prodded ramereth in our old irc channel, since that was the only place i spotted his client17:13
@jim:acmegating.comthanks!17:13
@fungicide:matrix.orgi can follow up by e-mail if that doesn't pan out17:13
@jim:acmegating.comsince the change to remove the osuosl labels isn't going to affect already submitted requests, i'm going to manually dequeue that item to get the restart moving agin17:15
@fungicide:matrix.orgthanks, that sounds like the most pragmatic move17:15
@jim:acmegating.comthat was on ze02, it's now rebooting, but that means the playbook didn't get very far over the weekend.  this may take a while.17:17
@fungicide:matrix.orgjust another manic monday17:17
@fungicide:matrix.org(i don't like mondays)17:17
@jim:acmegating.comheh :)17:18
@jim:acmegating.comi wonder why we're not hitting timeouts in zuul-launcher17:18
@fungicide:matrix.orgohai!17:19
@ramereth:osuosl.org> Ramereth: seems like you may still be lingering here after we moved to matrix, but are you aware of stuck nova api requests in the arm cloud at osuosl? it started late last week but we didn't want to pester you since it seemed like everyone was busy with finals week there17:20
It's likely related to me changing to an HA setup. I thought I fixed the issue late last week however
@fungicide:matrix.orgi saw there were about a dozen server instances there that we didn't have record of in zuul and were a few days old as of ~saturday so i tried to delete them manually via openstackclient, but they never transition from active to deleting17:21
@fungicide:matrix.organd not even any error state reported via `openstack server show ...`17:22
@fungicide:matrix.orgso it's like the api requests are just not being processed17:22
@ramereth:osuosl.orgLet me look again then17:22
@fungicide:matrix.orgi suspect the same is happening to our server create requests17:22
@fungicide:matrix.orgthanks! appreciated as always, of course17:22
@jim:acmegating.comah, i think zuul sees those instances, plus around 3 instances that were managed by zuul (that it thinks are leaked), and believes (probably correctly) that there is insufficient quota to handle requests, so that's why they're all just sitting there.17:23
@jim:acmegating.comthe few instances that zuul thinks have leaked -- it's continually trying to delete those17:23
@fungicide:matrix.orgyeah, that was my initial hope, that we just had zuul seeing no available quota and if i deleted the "leaked" unused nodes it would go back to normal17:24
@jim:acmegating.comi think we can infer that the problem (api calls are ineffective) is still there17:24
@fungicide:matrix.orgwhen i last looked (on saturday) there was also one server instance in that tenant stuck in an error state17:25
@jim:acmegating.com(since zuul is continually testing that for us)17:25
@fungicide:matrix.orgbut all the others were "active"17:25
@fungicide:matrix.orgso i was less concerned about the error server and more about the active ones that weren't deletable17:25
@jim:acmegating.comi think we can hold off on merging https://review.opendev.org/993417 for a bit; it'll be easier to see progress if Ramereth is able to fix it17:25
@fungicide:matrix.orgbut getting the error one cleaned up as well could help free a smidge of quota too17:26
@ramereth:osuosl.orgWhich ones were you trying to delete?17:27
@fungicide:matrix.orgthese:17:28
1fb87003-2dd9-471c-bb18-0ab2a075957e
30162831-381c-4944-89c7-abbcd9a4f736
b623f7bc-ab4b-4cbc-b289-b90526eb7089
cf742b6d-0361-4d79-99db-998fba80ad3d
28d87aa9-f789-465b-86a6-918571482738
2f7580b8-4db9-4e77-8ff1-b102717189ae
e986317a-1d8c-4ff0-a32c-954924c5254b
3951310a-b033-4af8-8a8c-ce0a12349c96
31a1eb10-2b5b-49b3-9386-db4dab5c285e
028cdd7b-d5bf-4a9c-8043-bdbd1172afd5
@jim:acmegating.comzuul is continuously trying to delete those too17:29
@fungicide:matrix.orgthe error one is 2fbf2d88-c119-4704-818f-32ea089bdb4e but the rest are in active state just unused17:29
@jim:acmegating.comit tries about once/minute17:30
@fungicide:matrix.orgso 11 in total (1 error, 10 active)17:30
@ramereth:osuosl.orgI got 1fb87003-2dd9-471c-bb18-0ab2a075957e to delete, I need to restart the compute services on the nodes to fix this it seems17:35
@ramereth:osuosl.orgIn a meeting, will do this here in a bit17:35
@fungicide:matrix.orgthanks for digging deep there, and apologies it's going to take some effort17:36
@fungicide:matrix.orgwe're in no rush17:36
@fungicide:matrix.orgwe saw it around thursday/friday but didn't want to eat into anybody's weekend17:36
-@gerrit:opendev.org- Jeremy Stanley https://matrix.to/#/@fungicide:matrix.org proposed: [opendev/system-config] 993428: Add another user agent to our apache blocklist https://review.opendev.org/c/opendev/system-config/+/99342817:49
@fungicide:matrix.orgit's started again ^ load average on static03 was back up around 2k briefly17:50
@ramereth:osuosl.orgfungi: I've restarted all of the hypervisor services. should hopefully be better now18:06
@fungicide:matrix.orgthanks Ramereth!!! i'll take a look18:07
@fungicide:matrix.orgcool, i see a few new server instances already, and the old ones are gone18:08
@fungicide:matrix.organd more building now18:09
@fungicide:matrix.orghttps://grafana.opendev.org/d/2c6f499090/zuul-launcher3a-osuosl is looking more healthy, but i'll keep an eye on things18:11
@fungicide:matrix.orgin continued news, i'm seeing e-mailed events now of buildsets waiting on arm jobs completing, so it seems like the backlog is burning down18:54
@fungicide:matrix.orglast call for adding agenda items to https://wiki.openstack.org/wiki/Meetings/InfraTeamMeeting#Agenda_for_next_meeting otherwise i'm sending what's up there to the service-discuss ml shortly in preparation for tomorrow's meeting19:03
@fungicide:matrix.orgtomorrow's meeting agenda has been distributed to the mailing list now20:47

Generated by irclog2html.py 4.1.0 by Marius Gedminas - find it at https://mg.pov.lt/irclog2html/!