| -@gerrit:opendev.org- Anil Shashikumar Belur proposed on behalf of Chaemin Lim: [openstack/project-config] 961499: i18n: Add new Zuul job for Weblate migration testing https://review.opendev.org/c/openstack/project-config/+/961499 | 01:23 | |
| @abelur:matrix.org | re: 1005814 - The image build ENOSPC fix merged last week - ubuntu-jammy is 6/6 since, including 3/3 on raxflex-iad3 which had never passed before. Remaining POST_FAILUREs on noble are a different thing: swift 500s on segment upload, not disk. | 01:27 |
|---|---|---|
| @mnasiadka:matrix.org | It seems dfw3 mirror has broken afs cache | 07:54 |
| @mnasiadka:matrix.org | #status log restarted mirror01.dfw3.raxflex.opendev.org due to afs cache and i/o problems | 08:10 |
| @status:opendev.org | @mnasiadka:matrix.org: finished logging | 08:10 |
| @harbott.osism.tech:regio.chat | as noted in #openstack-infra, it seems mirror03.dfw.rax.opendev.org is being sad, connection refused to both mirror requests and SSH login. trying to reboot now | 11:14 |
| @harbott.osism.tech:regio.chat | ok, seems to be working again fine now. there was an unscheduled (or so I think) reboot earlier, which seems to have gotten stuck somehow: | 11:25 |
| ``` | ||
| $ sudo journalctl --list-boots | ||
| IDX BOOT ID FIRST ENTRY LAST ENTRY | ||
| -3 ee9f5fd576d5434fa4a8047e0955ebd4 Thu 2025-09-18 04:17:01 UTC Thu 2026-08-06 20:00:41 UTC | ||
| -2 042cbbc808084c4eb6c9d03aa7a6f3d6 Thu 2026-08-06 20:01:27 UTC Wed 2026-09-23 08:00:26 UTC | ||
| -1 86003fb016604aa38e974a0fdffd5d72 Wed 2026-09-23 08:02:06 UTC Wed 2026-09-23 08:03:20 UTC | ||
| 0 afbf2b0bc81f4913939e6d43aabd97ae Wed 2026-09-23 11:15:46 UTC Wed 2026-09-23 11:21:28 UTC | ||
| ``` | ||
| @noonedeadpunk:matrix.org | fwiw, zuul<->gerrit connection looks a bit slow or smth. It tooks up to couple of minuites for zull to pick-up a job since yesterday evenining. No idea if that's related or was an accident | 12:25 |
| @noonedeadpunk:matrix.org | To a point that today I tried to add extra +W 2 times, then I thought there's some conflict as the gate job wasn't scheduled, rebased the patch, and only in a minute after rebase the gate job showed up... And got cancelled in next 5 minutes, due to rebase... | 12:33 |
| @jim:acmegating.com | Dmitriy Rabotyagov: reports like that are more helpful with a specific change or buildset that we can look at | 13:15 |
| @noonedeadpunk:matrix.org | It was https://review.opendev.org/c/openstack/openstack-ansible-os_ceilometer/+/1003639 I believe | 13:16 |
| @noonedeadpunk:matrix.org | Also gate finished before check was as you can see, so it wasn't even aborted (I was wroing about that) | 13:17 |
| @jim:acmegating.com | Dmitriy Rabotyagov: thanks. 2 things happened: | 13:56 |
| 1) when zuul got the original W+1, it performed gerrit http queries which encountered network errors as you suspected. that may be related to the ongoing ipv6 issues. due to the errors, it discarded the event. | ||
| 2) the later events all happened at almost the exact time that one of the periodic pipelines started. zuul didn't miss any events or encounter any errors, but due to the event backlog, the change wasn't enqueued in gate by the time the rebased patchset was uploaded, so zuul ignored it. that's an intentional optimization since the result isn't an error, just an inefficiency, and it's not usually a problem. | ||
| @jim:acmegating.com | it does seem likely we're seeing the ongoing review.o.o ipv6 issues start to affect zuul (via the http requests) | 13:57 |
| @noonedeadpunk:matrix.org | Well, there were at least 2 minutes between events. I'd say zuul used to pick-up events at least to show up on status page at all faster then that? | 13:57 |
| @noonedeadpunk:matrix.org | But yes, sure, ipv6 issues indeed could mess things up | 13:58 |
| @jim:acmegating.com | for the events around 0900 utc zuul received the events at the exact correct time, they just ended up behind several 10s of thousands of other events. it takes a minute. | 13:59 |
| @noonedeadpunk:matrix.org | And then it was 11 minutes until it picked the +W up and didn't cancel it with rebase... So I am not saying it was doing smth wrong, but it feels like either a queue or some kind of underlying issue | 13:59 |
| @jim:acmegating.com | i literally said it was a queue | 14:00 |
| @noonedeadpunk:matrix.org | should not it have cancelled the gate job though? when it realized the rebase event? | 14:01 |
| @noonedeadpunk:matrix.org | (just wondering) | 14:01 |
| @jim:acmegating.com | there are multiple queues involved, and they are processed asynchronously. everything is processed in order, eventually, but we don't always forward every event to every pipeline. the rebase event was not forwarded to the gate pipeline because the change was not in the gate pipeline at that time (due to the periodic-induced backlog). that's the optimization i was referring to. it's expensive for a pipeline to process an event, and we consider canceling running builds "best effort". the gate pipeline wouldn't (and didn't) miss any actual triggering events. incidentally, if openstack didn't have the clean-check requirement, then the last W+1 would have immediately moved it to gate and canceled it then. we always forward events that match the trigger conditions. | 14:06 |
| @noonedeadpunk:matrix.org | ok, gotcha, makes total sense now. thanks for explanation! | 14:13 |
| @harbott.osism.tech:regio.chat | hmm, not sure what rax did this morning, seems mirror.dfw3.raxflex.opendev.org also got rebooted five minutes after the dfw mirror. noticed that while checking the failure here https://zuul.opendev.org/t/openstack/build/2fe17ce4fd7e47b2bdf1c9efac150f23 | 14:21 |
| @fungicide:matrix.org | Jens Harbott: different clouds entirely, but likely in the same building. maybe something related to the facility? Dan With might know | 14:26 |
| -@gerrit:opendev.org- Harald Jensås proposed: [openstack/diskimage-builder] 1006958: DNM - Test CentOS 10-stream build in CI https://review.opendev.org/c/openstack/diskimage-builder/+/1006958 | 14:37 | |
| @fungicide:matrix.org | i've got an appointment and errands to run before the storm reaches us, back in a couple of hours | 15:01 |
| -@gerrit:opendev.org- Clark Boylan proposed: [opendev/system-config] 1006976: Convert review03 to static ipv6 config https://review.opendev.org/c/opendev/system-config/+/1006976 | 16:00 | |
| @clarkb:matrix.org | Jens Harbott: ^ | 16:00 |
| @clarkb:matrix.org | Note I'm not sure how to verify the static routes in there. They are carried over from the review02 config | 16:00 |
| @clarkb:matrix.org | but I think I got the ipv6 address and the mac address correctly updated for review03 | 16:00 |
| @clarkb:matrix.org | Yup I think your fix is allowing us to discover other issues that were masked. The 500 errors are theoretically something on the cloud side that Dan With may be able to help with. But not something we can debug directly. | 16:52 |
| -@gerrit:opendev.org- Clark Boylan proposed: | 17:54 | |
| - [opendev/system-config] 1006976: Convert review03 to static ipv6 config https://review.opendev.org/c/opendev/system-config/+/1006976 | ||
| - [opendev/system-config] 1007008: Unquote Registered Users in Gerrit config https://review.opendev.org/c/opendev/system-config/+/1007008 | ||
| @clarkb:matrix.org | that is a fun issue caused by direct enquing the upgrade change to the gate yesterday. Basically the 3.13 -> 3.14 upgrade test has a complaint and we only run that in check | 17:54 |
| @clarkb:matrix.org | it doesn't appear to be anything that would impact the running 3.13. This is just about being prepped for an eventual upgrade and we're tripping over that in a way that fails our checks | 17:55 |
| @fungicide:matrix.org | aha, i didn't realize we weren't including the upgrade job in the gate | 17:58 |
| -@gerrit:opendev.org- Harald Jensås proposed: [openstack/diskimage-builder] 1007010: Fix centos element BASE_IMAGE_FILE https://review.opendev.org/c/openstack/diskimage-builder/+/1007010 | 18:01 | |
| @jim:acmegating.com | i think that's okay | 18:01 |
| @clarkb:matrix.org | ya I think this situation illustrates it only catches forward looking problems anyway | 18:01 |
| @clarkb:matrix.org | which we can probably deal with in the rare cases they occur | 18:01 |
| @fungicide:matrix.org | sure, we're not developing gerrit such that we want to make sure other people can upgrade it, we just want to know that we have a viable upgrade path from our current gerrit to whatever we're looking at running down the road | 18:04 |
| @fungicide:matrix.org | and we don't want temporary lack of upgradability to the next minor release to block us from upgrading to a patch version anyway | 18:05 |
| @clarkb:matrix.org | fungi: maybe you can go over the netplan details in that change with a fine toothed comb? I noted that I'm not quite sure if those routes are correct for the current host | 18:09 |
| @fungicide:matrix.org | yep, have a few minutes to take a look at it now | 18:39 |
| @fungicide:matrix.org | i guess we would plan to perform a controlled reboot of the server after that deploys? | 18:42 |
| @clarkb:matrix.org | fungi: I think at the very least we would do a netplan try either after manually applying that or after deploying the change | 18:42 |
| @clarkb:matrix.org | but we can apply the netplan config properly without a reboot if that goes well. Then its a decision on whether or not a reboot is necessary | 18:42 |
| @fungicide:matrix.org | i wonder if it would also make sense to temporarily remove the aaaa records for it first and wait a bit until most traffic is using ipv4 | 18:43 |
| @fungicide:matrix.org | i think i would still want an eventual reboot either way since the change is overwriting cloud-init generated configuration and we want to make sure it doesn't get rewritten | 18:45 |
| @clarkb:matrix.org | ya that is a good point | 18:46 |
| @clarkb:matrix.org | it worked with review02 but that was an older platform and things may have changed behavior since | 18:46 |
| @fungicide:matrix.org | cloud-init hasn't touched the file since last year so it probably won't, but for all i know it could be comparing a hash of content and deciding it doesn't need to | 18:46 |
| @fungicide:matrix.org | the current v6 default route on review03 ens3 is fe80::21c:73ff:fe00:2000 which has the same neighbor mac as 2604:e100:1::1 (00:1c:73:00:20:00) | 18:56 |
| @fungicide:matrix.org | for whatever reason, at the moment 2604:e100:1::2 is unreachable from review03 and from the global internet | 18:57 |
| @clarkb:matrix.org | maybe it is no longer a valid router in their environment? | 18:58 |
| @fungicide:matrix.org | also possible they temporarily took it offline to troubleshoot the networking problem | 19:02 |
| @clarkb:matrix.org | ya could be related to what we've observed too (in a direct way rather than a debugging side effect) | 19:03 |
| @clarkb:matrix.org | fwiw I'm about to figure out lunch. Then I'm hoping to pop out for a bike ride since that didn't work out yesterday with the gerrit stuff. But I'm happy for people to update this change or test things etc | 19:03 |
| -@gerrit:opendev.org- Zuul merged on behalf of Clark Boylan: [opendev/system-config] 1007008: Unquote Registered Users in Gerrit config https://review.opendev.org/c/opendev/system-config/+/1007008 | 22:16 | |
| @clarkb:matrix.org | That change updated the server as expected (but it is a noop until the next restart) | 23:21 |
| @clarkb:matrix.org | the gerrit docs indicate that it will fail to start if the group named there is invalid and we have coverage of starting and checking the running gerrit in CI so I'm fairly confident that it will work as expected when we do restart | 23:21 |
Generated by irclog2html.py 4.1.0 by Marius Gedminas - find it at https://mg.pov.lt/irclog2html/!