Monday, 2026-09-28

-@gerrit:opendev.org- Jeremy Stanley https://matrix.to/#/@fungicide:matrix.org proposed: [opendev/zuul-providers] 1007746: Revert "Disable raxflex sjc3" https://review.opendev.org/c/opendev/zuul-providers/+/100774615:47
@clarkb:matrix.orgI'm doing some quick local system updates and reboots then I'm  going to check on the zuul cluster as it should've reset its podman state automatically. If things look good should I go ahead and approve the revert of the podman state cc corvus 15:48
@jim:acmegating.comthat's looking good15:52
@clarkb:matrix.orgall of the executors have a similar amount of free disk 14-17gb other than ze11 whcih is the brand new server it has 22gb free16:01
@clarkb:matrix.orgI'm trynig to collect some du information on ze01:/var now just to double check there isn't something we're still missing16:03
@fungicide:matrix.orgfewer old container images/layers i guess?16:03
@clarkb:matrix.orgwell we reset the entire podman system so they should all be equivalent there16:04
@clarkb:matrix.orgbut it wouldn't surprise me if we've got more logs or something16:04
@fungicide:matrix.orgoh, good point16:04
@fungicide:matrix.org5-8gb of old rotated debug logs, even compressed, wouldn't surprise me16:05
@jim:acmegating.comalso, i think quite a bit was going into the journal too; should be a bit less now, but could still have remnants in the older executors16:05
@fungicide:matrix.orgyeah, i don't recall what the default journal vacuum period is16:06
@clarkb:matrix.orgI suspect that we can merge the revert16:06
@jim:acmegating.com(i think there was a particular log line that ended up going to stderr a lot and ended up in the journal)16:06
@clarkb:matrix.orgas the cacti graph shows it did drop disk usage implying it was effective16:06
@clarkb:matrix.orghttps://review.opendev.org/c/opendev/system-config/+/1006633 this change is the one I'm talking about16:07
@clarkb:matrix.org/var/log is 1.4gb on ze11 and is 4.9gb on ze0116:10
@clarkb:matrix.orgI think that explains a good chunk of the delta16:10
@clarkb:matrix.orgthe other half is probably in journalctl16:11
@clarkb:matrix.orgso ya I'll approve 1006633 shortly unless there are any other concerns raised16:11
@fungicide:matrix.orgno concerns, +216:19
@clarkb:matrix.orgReminder that we'll start the pre ptg tomorrow https://etherpad.opendev.org/p/opendev-preptg-202609 details are in this document. I'll also send an email at some point today "cancelling" tomorrow's two regularly scheduled meetings in favor of the times blocked out for the pre ptg16:20
@clarkb:matrix.orgplease get any thoughts or ideas for agenda topics on that document now if you've got them16:21
@clarkb:matrix.orgthen sort of related to that I would like to get https://review.opendev.org/c/opendev/system-config/+/1002424 in to update the python version on the gerrit images to 3.14. But I don't want to do that until after the openstack release. I was thinking we could potentially do that as part of the pre ptg on thursday if we run out of other topics (others might find the gerrit restart process informative/useful/interesting)16:22
@fungicide:matrix.orgsure sgtm, or even wednesday since the sensitive parts of openstack release work should be done by the time we start anyway16:23
@fungicide:matrix.org(and if for some reason it's not, i'll be more focused on that than on our pre-ptg)16:24
@clarkb:matrix.orgya thursday seemed safer from a "no more agenda items" and "openstack release should erally be done by now" perspective16:25
@clarkb:matrix.orgZuul seems quite busy for release week too16:25
@clarkb:matrix.orgPS Cloud got back to us and says we can set up an account in their system to get access to the test env16:28
@clarkb:matrix.orgnot sure if anyone was interested in doing the initial account setup. I think Anil Belur and Eric Ball may be interested in helping with the benchmarking that happens after we get set up16:29
@fungicide:matrix.orgzte also reached out to jane who reached out to me about sizing an initial riscv environment to connect to our zuul16:30
@clarkb:matrix.orgmaybe we that would make a good pre ptg topic. I'll make sure we have something on the etherpad so we don't forget16:30
-@gerrit:opendev.org- Zuul merged on behalf of James E. Blair https://matrix.to/#/@jim:acmegating.com: [opendev/system-config] 1006633: Revert "Reset podman data on zuul components" https://review.opendev.org/c/opendev/system-config/+/100663316:32
@scott.little:matrix.orgis there anything broken after the gerrit upgrade?  Over the weekend I observed several 'git review' attemps hang repeatedly, only to pass 5 min later.  As if the service was unstable and rebooting every so often.   Now i have several folks complaining that reviews no longer auto-merge after WF+1.  Instead the are left in a 'Ready to submit' state18:49
@clarkb:matrix.orgscott.little: the issue predates the Gerrit upgrade. The cloud provider has been struggling with ipv6 routing recently which is why you see it hang then work (systems will typically try ipv6 first then fallback to ipv4). Similarly with zuul it will connect via ipv6 at first then when that stops working fallback to ipv4 and work consistently until the next restart (we restart zuul at least weekly)18:51
@clarkb:matrix.orgall that to say I don't believe the new version of Gerrit is to blame. But instead weekend restarts of zuul and general ipv6 connectivity struggles18:51
@clarkb:matrix.orgwe've debated dropping the AAAA record from DNS which may end up being worse for some users if they don't have ipv4 though18:51
@clarkb:matrix.orgcorvus: /dev/mapper/main-mariadb   98G   98G     0 100% /var/mariadb18:54
@clarkb:matrix.orgmordred: ^ I think this explains what you've seen as well18:54
@clarkb:matrix.orgmaybe18:54
@clarkb:matrix.orgcorvus: should we pause result handling in zuul while we figure that out?18:54
@clarkb:matrix.orgI'm thinking we can attach a second bigger volume. Have lvm and the fs expand into that then maybe we even restart mariadb? I don't know how gracefully it will handle running out of disk18:55
@clarkb:matrix.orgmaybe we want to stop mariadb in order to expand the fs anyway?18:56
@clarkb:matrix.org* corvus: /dev/mapper/main-mariadb   98G   98G     0 100% /var/mariadb on zuul-db0118:57
@fungicide:matrix.orgyeah, right now it's using `/dev/xvdb` (ironic name) as the only lvm2 pv, which seems to be a 100gb device18:57
@fungicide:matrix.orgit looks like a cinder volume? i'll have to check the api to be sure18:57
@clarkb:matrix.orgyes I think that is a cindervolume. I think we can theoretically create a new cinder volume that is larger and add it to lvm. Then possibly move the existing lvm content into the new volume (so the old one can be removed) then expand the fs18:58
@fungicide:matrix.orgwe could attach a new larger cinder volume, add it as a pv to the existing vg, then pvmove the extents for the current lv onto it and lvextend it once there and resize the fs18:58
@clarkb:matrix.orgyup that. Do we want to get started on that before corvus has a chance to weigh in?18:59
@fungicide:matrix.orgi'll get to creating the volume in cinder and attaching it18:59
@fungicide:matrix.orgsomeone should #status notice or alert?18:59
@clarkb:matrix.orgthanks. Looks like ext4 can be safely expanded (not shrunk) while online and active18:59
@clarkb:matrix.orgI'll do that18:59
@fungicide:matrix.orgyes, it can be expanded just fine19:00
@clarkb:matrix.orgHow does this look #status notice Zuul is reporting inconsistent results showing failures in Gerrit comments and in progress status in the Zuul UI. This is due to Zuul database issues that we are working to correct.19:01
@fungicide:matrix.orgwfm, thx19:01
@clarkb:matrix.org#status notice Zuul is reporting inconsistent results showing failures in Gerrit comments and in progress status in the Zuul UI. This is due to Zuul database issues that we are working to correct.19:01
@status:opendev.org@clarkb:matrix.org: sending notice19:01
@fungicide:matrix.org200gb? or bigger?19:02
@fungicide:matrix.org`zuul-db01.opendev.org/main01` is ssd type, for the record19:02
@clarkb:matrix.orgI'm thinking bigger like 300 at least just to give us more headroom19:03
@clarkb:matrix.orgI don't think we need to go huge like 1TB (which is where I'd be more concerned with an ssd)19:03
@fungicide:matrix.orgcreated with `openstack --os-cloud=openstackci-rax --os-region-name=DFW volume create --type=SSD --size=300 zuul-db01.opendev.org/main02`19:04
@status:opendev.org@clarkb:matrix.org: finished sending notice19:04
-@status:opendev.org- NOTICE: Zuul is reporting inconsistent results showing failures in Gerrit comments and in progress status in the Zuul UI. This is due to Zuul database issues that we are working to correct.19:04
@clarkb:matrix.orgI am going to pause zuul queues too19:05
@clarkb:matrix.orgso that we stop losing data about results19:05
@fungicide:matrix.org`[Mon Sep 28 19:03:01 2026] blkfront: xvdc: barrier: enabled; persistent grants: disabled; indirect descriptors: enabled; bounce buffer: disabled;` (from dmesg on the server after server add volume)19:06
@fungicide:matrix.org`xvdc             202:32   0  300G  0 disk` (from lsblk)19:06
@fungicide:matrix.orgi'll get it into the main vg19:06
@clarkb:matrix.orgqueues are paused19:06
@clarkb:matrix.orgthat had to pull the zuul-client image after our podman system reset19:07
@fungicide:matrix.orgthe main vg now includes a 100gb `/dev/xvdb1` and 300gb `/dev/xvdc1`19:09
@fungicide:matrix.orgstarting the pvmove now19:09
@clarkb:matrix.orgnote that mariadb is still running but I think we decided that wasn't an issue19:09
@fungicide:matrix.org`pvmove /dev/xvdb1 /dev/xvdc1` is underway in a root screen session on zuul-db01 now19:10
@fungicide:matrix.orgyes, volumes can remain mounted and writable throughout this process19:10
@fungicide:matrix.orgonce it's fully moved all the extents to `/dev/xvdc1` i'll remove `/dev/xvdb1` from the vg and then extent the lv to the size of the remaining extents (which will make it the size of `/dev/xvdc1`)19:11
@fungicide:matrix.orgafter that we can grow the fs19:11
@clarkb:matrix.orgsounds good. I'm around if I can help19:11
@fungicide:matrix.orgnothing beats a monday fire drill19:12
@fungicide:matrix.orgextents are about 40% of the way moved so far19:13
@fungicide:matrix.orgshouldn't take too much longer19:13
@fungicide:matrix.org75%19:16
@fungicide:matrix.orgdone, now removing the old pb19:18
@fungicide:matrix.orger, pv19:18
@fungicide:matrix.orgextended the lv, now lvs reports `mariadb main -wi-ao---- <300.00g`19:20
@fungicide:matrix.orgran resize2fs and now df reports `/dev/mapper/main-mariadb  295G   98G  197G  34% /var/mariadb`19:20
@clarkb:matrix.orgI see the same thing from df. Should I unpause the zuul result queues?19:21
@fungicide:matrix.orgi have not yet detached and deleted the cinder volume, but we should be ready to resume db writes if there's no application level cleanup needed19:21
@clarkb:matrix.orgI don't think there is application level cleanup. Instead I suspect that zuul simply error'd on mysql failures and proceeded along19:21
@fungicide:matrix.orgconvenient that was the only thing on the full fs, at least19:22
@clarkb:matrix.orgin other words we lost the build completion info for those builds19:22
@clarkb:matrix.orgthe results reported to gerrit should be accurate. It will just be difficult (near impossible) to debug reported failuers19:22
@clarkb:matrix.orgyou good with me unpausing zuul now?19:22
@fungicide:matrix.orgi am19:22
@fungicide:matrix.orgplease proceed19:22
@clarkb:matrix.orgdone19:22
@fungicide:matrix.orgi'll get the old cinder volume recycled back into our quota19:23
@fungicide:matrix.orgthe old `zuul-db01.opendev.org/main01` cinder volume has been detached and deleted now19:24
@fungicide:matrix.orgclosing out the screen session now19:25
@clarkb:matrix.orghttps://zuul.opendev.org/t/openstack/build/d3146735cc664df2b18269c543d5111e this build started and completed after I paused/unpaused things19:26
@clarkb:matrix.orgit renders in the UI indicating to me that things are happier moving forward19:27
@fungicide:matrix.orggood, good19:28
@clarkb:matrix.orgI'm going to eat lunch now19:29
@clarkb:matrix.organd then if things remain happy afterwards maybe a bike ride. Then get out the pre ptg is this week along with normal meeting cancellation email sent19:30
@jim:acmegating.comClark: thanks!  that entire thing managed to exactly overlap with my lunch acquisition, sorry19:35
@clarkb:matrix.orgit ended up being relatively straightforward to fix (I think)19:36
@clarkb:matrix.orghopefully we didn't do anything different or wrong compared to what you would've done19:36
@clarkb:matrix.orgok I need to eat now. We might send a notice that builds can be rechecked if missing logs are a problem19:37
@clarkb:matrix.orgbut I don't think there is an easy way to find the log urls for logs without effort19:37
@jim:acmegating.comno definitely not worth that level of effort19:37
@jim:acmegating.comif there's some critical release promote job we need to find, we could probably dig it up19:37
@fungicide:matrix.orgyeah, replacing/enlarging volumes and filesystems is engraved into my finger memory, didn't take long19:37
@jim:acmegating.combut not something we should do for the general case19:38
@jim:acmegating.comwhat's weird is that cacti says that /var/mariadb has been close to capacity for a long time... i think there's probably some weird internal allocation at play that makes it difficult to tell what the actual disk pressure is.19:39
@jim:acmegating.comfungi: thank you!19:42
@fungicide:matrix.orgit was my pleasure, as always19:45
@fungicide:matrix.orgi actually find playing around with kernel-level storage structures fun anyway19:45
@fungicide:matrix.orga fun one i have on my workstation is a pair of nvme drives in external usb type-c housings, i have them both luks2 encrypted and then a layer of luks2 integrity on top of that, then combined as a mdraid mirror, carved up with lvm2 and used to do backups of my important systems with periodic lvm snapshotting19:47
@fungicide:matrix.organd most fun was getting the systemd glue configured so it all comes up in the right order, unlocked with a keyfile on my workstation, auto-mounted on demand if any process reads from or writes to the mount path19:49
@fungicide:matrix.org(decrypted, integrity-checked, raided, activated, and mounted on the fly)19:50
@fungicide:matrix.orgthough the whole process causes the first read or write to block for about 10 seconds while it all spins up19:51
@clarkb:matrix.orgit looks like things are still happy so I'm going to pop out for that bike ride now20:31
@fungicide:matrix.orghave fun!20:31
@mordred:waterwanders.comah! yes, that definitely makes sense. I was starting to consider db/disk issues but didn't have enough data yet to suggest anything. glad it wound up being straightforward \o/21:53
-@gerrit:opendev.org- Steve Baker proposed:21:58
- [openstack/diskimage-builder] 983813: Refactor 02-set-machine-id into an element https://review.opendev.org/c/openstack/diskimage-builder/+/983813
- [openstack/diskimage-builder] 983814: Refactor 03-reset-bls-entries into an element https://review.opendev.org/c/openstack/diskimage-builder/+/983814
- [openstack/diskimage-builder] 983815: Add tarball element for unprivileged container builds https://review.opendev.org/c/openstack/diskimage-builder/+/983815
- [openstack/diskimage-builder] 984486: Skip local loop device creation for no-final-image builds https://review.opendev.org/c/openstack/diskimage-builder/+/984486
- [openstack/diskimage-builder] 1007795: Add DIB_BIND_MOUNTS option for container environments https://review.opendev.org/c/openstack/diskimage-builder/+/1007795
- [openstack/diskimage-builder] 1007796: Add dnf-assert element https://review.opendev.org/c/openstack/diskimage-builder/+/1007796
- [openstack/diskimage-builder] 1007797: epel: allow epel-release to not be installed at all https://review.opendev.org/c/openstack/diskimage-builder/+/1007797
- [openstack/diskimage-builder] 1007798: Check for python3-pyyaml before installing https://review.opendev.org/c/openstack/diskimage-builder/+/1007798
- [openstack/diskimage-builder] 1007799: Handle passwd coming from shadow-utils on rhel-10 https://review.opendev.org/c/openstack/diskimage-builder/+/1007799
- [openstack/diskimage-builder] 1007800: pip-and-virtrualenv don't depend on epel https://review.opendev.org/c/openstack/diskimage-builder/+/1007800
@clarkb:matrix.orgmessage sent about the meeting cancellation and pre ptg schedule22:48

Generated by irclog2html.py 4.1.0 by Marius Gedminas - find it at https://mg.pov.lt/irclog2html/!