| -@gerrit:opendev.org- Steven van der Schoot proposed: [opendev/git-review] 993318: Unquote percent-encoded usernames parsed from remote URLs https://review.opendev.org/c/opendev/git-review/+/993318 | 09:58 | |
| @priteau:matrix.org | We are seeing many post failures in CI right now. Anything broken? | 11:14 |
|---|---|---|
| @priteau:matrix.org | It's also been mentioned by someone else on IRC in #openstack-infra | 11:15 |
| @mnasiadka:matrix.org | It seems rax_iad Swift provider has some problems | 11:33 |
| @mnasiadka:matrix.org | Or maybe not, because another job failed on ovh_bhs | 11:36 |
| @mnasiadka:matrix.org | Ah, majority of the failures is ovh_gra and ovh_bhs | 11:38 |
| @mnasiadka:matrix.org | Seems there is something going on with OVH Object Storage: https://public-cloud.status-ovhcloud.com/incidents/7jjsc5hq8tnn | 11:38 |
| -@gerrit:opendev.org- Michal Nasiadka proposed: [opendev/base-jobs] 993335: Disable job log uploads to ovh swift https://review.opendev.org/c/opendev/base-jobs/+/993335 | 11:40 | |
| @priteau:matrix.org | Start time : 13/05/2026 00:00 UTC though… | 11:43 |
| @priteau:matrix.org | It might be a new incident not yet on their status page | 11:43 |
| @mnasiadka:matrix.org | well, that one is not closed ;-) | 11:46 |
| @mnasiadka:matrix.org | Although I agree it might be some old one | 11:46 |
| @sean-k-mooney:matrix.org | it might be a good idea to send out a status message to tell peopel to avoid rechecks until that is resolved | 12:00 |
| @sean-k-mooney:matrix.org | i know its does nto alwasy fail right now but its failing enought htat any one build set is unlikely to pass | 12:00 |
| -@gerrit:opendev.org- Steven van der Schoot proposed: [opendev/git-review] 993339: Escape/unescape usernames with special characters in remote URLs https://review.opendev.org/c/opendev/git-review/+/993339 | 12:01 | |
| @mnasiadka:matrix.org | Let me force merge that patch and send a notice | 12:09 |
| -@gerrit:opendev.org- mnasiadka.admin merged on behalf of Michal Nasiadka: [opendev/base-jobs] 993335: Disable job log uploads to ovh swift https://review.opendev.org/c/opendev/base-jobs/+/993335 | 12:34 | |
| @mnasiadka:matrix.org | #status notice Recent POST_FAILURE job results with no logs were due to upload errors in one of our providers, which has been temporarily disabled now so rechecking those should be safe | 12:41 |
| @status:opendev.org | @mnasiadka:matrix.org: sending notice | 12:41 |
| @priteau:matrix.org | Thanks for the prompt fix! | 12:45 |
| -@status:opendev.org- NOTICE: Recent POST_FAILURE job results with no logs were due to upload errors in one of our providers, which has been temporarily disabled now so rechecking those should be safe | 12:45 | |
| @status:opendev.org | @mnasiadka:matrix.org: finished sending notice | 12:45 |
| @mnasiadka:matrix.org | phew, all done | 12:46 |
| @fungicide:matrix.org | thanks mnasiadka! | 13:10 |
| @fungicide:matrix.org | looks like i showed up too late to catch all the fun | 13:10 |
| @mnasiadka:matrix.org | fungi: No problem! That was a good degree of learning how to force merge a patch ;) | 13:10 |
| @fungicide:matrix.org | hopefully the documentation was straightforward, though it's a lot of copy/paste of gerrit ssh api commands | 13:18 |
| @fungicide:matrix.org | mostly as a means of being able to use separate admin accounts that don't have openid logins, for added safety | 13:19 |
| @mnasiadka:matrix.org | yeah, I already had an admin account, so that wasn't a problem - but yes, I have a tendency of trying to understand what I'm copy pasting ;-) | 13:50 |
| @fungicide:matrix.org | as you should, of course, i'm just hoping the text around those copy-paste examples was sufficient explanation | 13:53 |
| @fungicide:matrix.org | if not, we should improve it | 13:53 |
| @fungicide:matrix.org | i'm going to disappear for an early lunch, but everyone check that topics you want covered are in the agenda and i'll send it out to service-discuss later today: https://wiki.openstack.org/wiki/Meetings/InfraTeamMeeting#Agenda_for_next_meeting | 14:35 |
| @fungicide:matrix.org | static03 is getting beaten up pretty hard at the moment, load average over 2k for at least the past 15 minutes, looks like. server-status indicates most workers are busy writing responses to clients | 14:37 |
| @fungicide:matrix.org | if this persists, it's worth seeing whether there are new crawler signatures we can add or if we need to consider other mitigations like bringing a larger replacement similar to static04 back (would need to create it from scratch again since i did eventually delete the old one) | 14:40 |
| @fungicide:matrix.org | according to top, the cpus are still fairly idle, but are spending the bulk of their ticks on "system" tasks | 14:42 |
| @fungicide:matrix.org | that may be indicative of interfacing with the openafs driver | 14:43 |
| @priteau:matrix.org | Yes, https://tarballs.openstack.org is down from here 😢 | 14:59 |
| -@gerrit:opendev.org- Monty Taylor https://matrix.to/#/@mordred:inaugust.com proposed: [openstack/project-config] 993214: Add repos for drizzle website and sysadmin automation https://review.opendev.org/c/openstack/project-config/+/993214 | 15:35 | |
| @fungicide:matrix.org | Pierre Riteau: should be back up again now. some crawler bot army was crushing it while i was at lunch, load average on the server was up over 2.5k that i saw | 16:38 |
| @fungicide:matrix.org | seems like they've moved on (for now anyway) | 16:38 |
| @jim:acmegating.com | it looks like the weekend zuul restart is stuck mid-upgrade | 16:56 |
| @jim:acmegating.com | it's because there is a buildset with a paused job waiting on other jobs which are waiting on arm nodes | 17:09 |
| @jim:acmegating.com | https://zuul.opendev.org/t/openstack/status?project=openstack%2Fopenstack-helm-images | 17:09 |
| @jim:acmegating.com | my understanding is that our arm provider is currently offline indefinitely | 17:10 |
| @fungicide:matrix.org | yeah, seems like the situation at osuosl hasn't changed yet | 17:10 |
| @jim:acmegating.com | i think at this point we may want to consider having zuul fail jobs rather than queueing indefinitely | 17:11 |
| @fungicide:matrix.org | last week was finals so we were hesitant to prod them, but it's definitely time to check in | 17:11 |
| -@gerrit:opendev.org- James E. Blair https://matrix.to/#/@jim:acmegating.com proposed: [opendev/zuul-providers] 993417: Remove all labels from osuosl https://review.opendev.org/c/opendev/zuul-providers/+/993417 | 17:12 | |
| @jim:acmegating.com | fungi: mnasiadka ^ | 17:13 |
| @fungicide:matrix.org | i've prodded ramereth in our old irc channel, since that was the only place i spotted his client | 17:13 |
| @jim:acmegating.com | thanks! | 17:13 |
| @fungicide:matrix.org | i can follow up by e-mail if that doesn't pan out | 17:13 |
| @jim:acmegating.com | since the change to remove the osuosl labels isn't going to affect already submitted requests, i'm going to manually dequeue that item to get the restart moving agin | 17:15 |
| @fungicide:matrix.org | thanks, that sounds like the most pragmatic move | 17:15 |
| @jim:acmegating.com | that was on ze02, it's now rebooting, but that means the playbook didn't get very far over the weekend. this may take a while. | 17:17 |
| @fungicide:matrix.org | just another manic monday | 17:17 |
| @fungicide:matrix.org | (i don't like mondays) | 17:17 |
| @jim:acmegating.com | heh :) | 17:18 |
| @jim:acmegating.com | i wonder why we're not hitting timeouts in zuul-launcher | 17:18 |
| @fungicide:matrix.org | ohai! | 17:19 |
| @ramereth:osuosl.org | > Ramereth: seems like you may still be lingering here after we moved to matrix, but are you aware of stuck nova api requests in the arm cloud at osuosl? it started late last week but we didn't want to pester you since it seemed like everyone was busy with finals week there | 17:20 |
| It's likely related to me changing to an HA setup. I thought I fixed the issue late last week however | ||
| @fungicide:matrix.org | i saw there were about a dozen server instances there that we didn't have record of in zuul and were a few days old as of ~saturday so i tried to delete them manually via openstackclient, but they never transition from active to deleting | 17:21 |
| @fungicide:matrix.org | and not even any error state reported via `openstack server show ...` | 17:22 |
| @fungicide:matrix.org | so it's like the api requests are just not being processed | 17:22 |
| @ramereth:osuosl.org | Let me look again then | 17:22 |
| @fungicide:matrix.org | i suspect the same is happening to our server create requests | 17:22 |
| @fungicide:matrix.org | thanks! appreciated as always, of course | 17:22 |
| @jim:acmegating.com | ah, i think zuul sees those instances, plus around 3 instances that were managed by zuul (that it thinks are leaked), and believes (probably correctly) that there is insufficient quota to handle requests, so that's why they're all just sitting there. | 17:23 |
| @jim:acmegating.com | the few instances that zuul thinks have leaked -- it's continually trying to delete those | 17:23 |
| @fungicide:matrix.org | yeah, that was my initial hope, that we just had zuul seeing no available quota and if i deleted the "leaked" unused nodes it would go back to normal | 17:24 |
| @jim:acmegating.com | i think we can infer that the problem (api calls are ineffective) is still there | 17:24 |
| @fungicide:matrix.org | when i last looked (on saturday) there was also one server instance in that tenant stuck in an error state | 17:25 |
| @jim:acmegating.com | (since zuul is continually testing that for us) | 17:25 |
| @fungicide:matrix.org | but all the others were "active" | 17:25 |
| @fungicide:matrix.org | so i was less concerned about the error server and more about the active ones that weren't deletable | 17:25 |
| @jim:acmegating.com | i think we can hold off on merging https://review.opendev.org/993417 for a bit; it'll be easier to see progress if Ramereth is able to fix it | 17:25 |
| @fungicide:matrix.org | but getting the error one cleaned up as well could help free a smidge of quota too | 17:26 |
| @ramereth:osuosl.org | Which ones were you trying to delete? | 17:27 |
| @fungicide:matrix.org | these: | 17:28 |
| 1fb87003-2dd9-471c-bb18-0ab2a075957e | ||
| 30162831-381c-4944-89c7-abbcd9a4f736 | ||
| b623f7bc-ab4b-4cbc-b289-b90526eb7089 | ||
| cf742b6d-0361-4d79-99db-998fba80ad3d | ||
| 28d87aa9-f789-465b-86a6-918571482738 | ||
| 2f7580b8-4db9-4e77-8ff1-b102717189ae | ||
| e986317a-1d8c-4ff0-a32c-954924c5254b | ||
| 3951310a-b033-4af8-8a8c-ce0a12349c96 | ||
| 31a1eb10-2b5b-49b3-9386-db4dab5c285e | ||
| 028cdd7b-d5bf-4a9c-8043-bdbd1172afd5 | ||
| @jim:acmegating.com | zuul is continuously trying to delete those too | 17:29 |
| @fungicide:matrix.org | the error one is 2fbf2d88-c119-4704-818f-32ea089bdb4e but the rest are in active state just unused | 17:29 |
| @jim:acmegating.com | it tries about once/minute | 17:30 |
| @fungicide:matrix.org | so 11 in total (1 error, 10 active) | 17:30 |
| @ramereth:osuosl.org | I got 1fb87003-2dd9-471c-bb18-0ab2a075957e to delete, I need to restart the compute services on the nodes to fix this it seems | 17:35 |
| @ramereth:osuosl.org | In a meeting, will do this here in a bit | 17:35 |
| @fungicide:matrix.org | thanks for digging deep there, and apologies it's going to take some effort | 17:36 |
| @fungicide:matrix.org | we're in no rush | 17:36 |
| @fungicide:matrix.org | we saw it around thursday/friday but didn't want to eat into anybody's weekend | 17:36 |
| -@gerrit:opendev.org- Jeremy Stanley https://matrix.to/#/@fungicide:matrix.org proposed: [opendev/system-config] 993428: Add another user agent to our apache blocklist https://review.opendev.org/c/opendev/system-config/+/993428 | 17:49 | |
| @fungicide:matrix.org | it's started again ^ load average on static03 was back up around 2k briefly | 17:50 |
| @ramereth:osuosl.org | fungi: I've restarted all of the hypervisor services. should hopefully be better now | 18:06 |
| @fungicide:matrix.org | thanks Ramereth!!! i'll take a look | 18:07 |
| @fungicide:matrix.org | cool, i see a few new server instances already, and the old ones are gone | 18:08 |
| @fungicide:matrix.org | and more building now | 18:09 |
| @fungicide:matrix.org | https://grafana.opendev.org/d/2c6f499090/zuul-launcher3a-osuosl is looking more healthy, but i'll keep an eye on things | 18:11 |
| @fungicide:matrix.org | in continued news, i'm seeing e-mailed events now of buildsets waiting on arm jobs completing, so it seems like the backlog is burning down | 18:54 |
| @fungicide:matrix.org | last call for adding agenda items to https://wiki.openstack.org/wiki/Meetings/InfraTeamMeeting#Agenda_for_next_meeting otherwise i'm sending what's up there to the service-discuss ml shortly in preparation for tomorrow's meeting | 19:03 |
| @fungicide:matrix.org | tomorrow's meeting agenda has been distributed to the mailing list now | 20:47 |
Generated by irclog2html.py 4.1.0 by Marius Gedminas - find it at https://mg.pov.lt/irclog2html/!