| -@gerrit:opendev.org- Jeremy Stanley https://matrix.to/#/@fungicide:matrix.org proposed: [opendev/zuul-providers] 1007746: Revert "Disable raxflex sjc3" https://review.opendev.org/c/opendev/zuul-providers/+/1007746 | 15:47 | |
| @clarkb:matrix.org | I'm doing some quick local system updates and reboots then I'm going to check on the zuul cluster as it should've reset its podman state automatically. If things look good should I go ahead and approve the revert of the podman state cc corvus | 15:48 |
|---|---|---|
| @jim:acmegating.com | that's looking good | 15:52 |
| @clarkb:matrix.org | all of the executors have a similar amount of free disk 14-17gb other than ze11 whcih is the brand new server it has 22gb free | 16:01 |
| @clarkb:matrix.org | I'm trynig to collect some du information on ze01:/var now just to double check there isn't something we're still missing | 16:03 |
| @fungicide:matrix.org | fewer old container images/layers i guess? | 16:03 |
| @clarkb:matrix.org | well we reset the entire podman system so they should all be equivalent there | 16:04 |
| @clarkb:matrix.org | but it wouldn't surprise me if we've got more logs or something | 16:04 |
| @fungicide:matrix.org | oh, good point | 16:04 |
| @fungicide:matrix.org | 5-8gb of old rotated debug logs, even compressed, wouldn't surprise me | 16:05 |
| @jim:acmegating.com | also, i think quite a bit was going into the journal too; should be a bit less now, but could still have remnants in the older executors | 16:05 |
| @fungicide:matrix.org | yeah, i don't recall what the default journal vacuum period is | 16:06 |
| @clarkb:matrix.org | I suspect that we can merge the revert | 16:06 |
| @jim:acmegating.com | (i think there was a particular log line that ended up going to stderr a lot and ended up in the journal) | 16:06 |
| @clarkb:matrix.org | as the cacti graph shows it did drop disk usage implying it was effective | 16:06 |
| @clarkb:matrix.org | https://review.opendev.org/c/opendev/system-config/+/1006633 this change is the one I'm talking about | 16:07 |
| @clarkb:matrix.org | /var/log is 1.4gb on ze11 and is 4.9gb on ze01 | 16:10 |
| @clarkb:matrix.org | I think that explains a good chunk of the delta | 16:10 |
| @clarkb:matrix.org | the other half is probably in journalctl | 16:11 |
| @clarkb:matrix.org | so ya I'll approve 1006633 shortly unless there are any other concerns raised | 16:11 |
| @fungicide:matrix.org | no concerns, +2 | 16:19 |
| @clarkb:matrix.org | Reminder that we'll start the pre ptg tomorrow https://etherpad.opendev.org/p/opendev-preptg-202609 details are in this document. I'll also send an email at some point today "cancelling" tomorrow's two regularly scheduled meetings in favor of the times blocked out for the pre ptg | 16:20 |
| @clarkb:matrix.org | please get any thoughts or ideas for agenda topics on that document now if you've got them | 16:21 |
| @clarkb:matrix.org | then sort of related to that I would like to get https://review.opendev.org/c/opendev/system-config/+/1002424 in to update the python version on the gerrit images to 3.14. But I don't want to do that until after the openstack release. I was thinking we could potentially do that as part of the pre ptg on thursday if we run out of other topics (others might find the gerrit restart process informative/useful/interesting) | 16:22 |
| @fungicide:matrix.org | sure sgtm, or even wednesday since the sensitive parts of openstack release work should be done by the time we start anyway | 16:23 |
| @fungicide:matrix.org | (and if for some reason it's not, i'll be more focused on that than on our pre-ptg) | 16:24 |
| @clarkb:matrix.org | ya thursday seemed safer from a "no more agenda items" and "openstack release should erally be done by now" perspective | 16:25 |
| @clarkb:matrix.org | Zuul seems quite busy for release week too | 16:25 |
| @clarkb:matrix.org | PS Cloud got back to us and says we can set up an account in their system to get access to the test env | 16:28 |
| @clarkb:matrix.org | not sure if anyone was interested in doing the initial account setup. I think Anil Belur and Eric Ball may be interested in helping with the benchmarking that happens after we get set up | 16:29 |
| @fungicide:matrix.org | zte also reached out to jane who reached out to me about sizing an initial riscv environment to connect to our zuul | 16:30 |
| @clarkb:matrix.org | maybe we that would make a good pre ptg topic. I'll make sure we have something on the etherpad so we don't forget | 16:30 |
| -@gerrit:opendev.org- Zuul merged on behalf of James E. Blair https://matrix.to/#/@jim:acmegating.com: [opendev/system-config] 1006633: Revert "Reset podman data on zuul components" https://review.opendev.org/c/opendev/system-config/+/1006633 | 16:32 | |
| @scott.little:matrix.org | is there anything broken after the gerrit upgrade? Over the weekend I observed several 'git review' attemps hang repeatedly, only to pass 5 min later. As if the service was unstable and rebooting every so often. Now i have several folks complaining that reviews no longer auto-merge after WF+1. Instead the are left in a 'Ready to submit' state | 18:49 |
| @clarkb:matrix.org | scott.little: the issue predates the Gerrit upgrade. The cloud provider has been struggling with ipv6 routing recently which is why you see it hang then work (systems will typically try ipv6 first then fallback to ipv4). Similarly with zuul it will connect via ipv6 at first then when that stops working fallback to ipv4 and work consistently until the next restart (we restart zuul at least weekly) | 18:51 |
| @clarkb:matrix.org | all that to say I don't believe the new version of Gerrit is to blame. But instead weekend restarts of zuul and general ipv6 connectivity struggles | 18:51 |
| @clarkb:matrix.org | we've debated dropping the AAAA record from DNS which may end up being worse for some users if they don't have ipv4 though | 18:51 |
| @clarkb:matrix.org | corvus: /dev/mapper/main-mariadb 98G 98G 0 100% /var/mariadb | 18:54 |
| @clarkb:matrix.org | mordred: ^ I think this explains what you've seen as well | 18:54 |
| @clarkb:matrix.org | maybe | 18:54 |
| @clarkb:matrix.org | corvus: should we pause result handling in zuul while we figure that out? | 18:54 |
| @clarkb:matrix.org | I'm thinking we can attach a second bigger volume. Have lvm and the fs expand into that then maybe we even restart mariadb? I don't know how gracefully it will handle running out of disk | 18:55 |
| @clarkb:matrix.org | maybe we want to stop mariadb in order to expand the fs anyway? | 18:56 |
| @clarkb:matrix.org | * corvus: /dev/mapper/main-mariadb 98G 98G 0 100% /var/mariadb on zuul-db01 | 18:57 |
| @fungicide:matrix.org | yeah, right now it's using `/dev/xvdb` (ironic name) as the only lvm2 pv, which seems to be a 100gb device | 18:57 |
| @fungicide:matrix.org | it looks like a cinder volume? i'll have to check the api to be sure | 18:57 |
| @clarkb:matrix.org | yes I think that is a cindervolume. I think we can theoretically create a new cinder volume that is larger and add it to lvm. Then possibly move the existing lvm content into the new volume (so the old one can be removed) then expand the fs | 18:58 |
| @fungicide:matrix.org | we could attach a new larger cinder volume, add it as a pv to the existing vg, then pvmove the extents for the current lv onto it and lvextend it once there and resize the fs | 18:58 |
| @clarkb:matrix.org | yup that. Do we want to get started on that before corvus has a chance to weigh in? | 18:59 |
| @fungicide:matrix.org | i'll get to creating the volume in cinder and attaching it | 18:59 |
| @fungicide:matrix.org | someone should #status notice or alert? | 18:59 |
| @clarkb:matrix.org | thanks. Looks like ext4 can be safely expanded (not shrunk) while online and active | 18:59 |
| @clarkb:matrix.org | I'll do that | 18:59 |
| @fungicide:matrix.org | yes, it can be expanded just fine | 19:00 |
| @clarkb:matrix.org | How does this look #status notice Zuul is reporting inconsistent results showing failures in Gerrit comments and in progress status in the Zuul UI. This is due to Zuul database issues that we are working to correct. | 19:01 |
| @fungicide:matrix.org | wfm, thx | 19:01 |
| @clarkb:matrix.org | #status notice Zuul is reporting inconsistent results showing failures in Gerrit comments and in progress status in the Zuul UI. This is due to Zuul database issues that we are working to correct. | 19:01 |
| @status:opendev.org | @clarkb:matrix.org: sending notice | 19:01 |
| @fungicide:matrix.org | 200gb? or bigger? | 19:02 |
| @fungicide:matrix.org | `zuul-db01.opendev.org/main01` is ssd type, for the record | 19:02 |
| @clarkb:matrix.org | I'm thinking bigger like 300 at least just to give us more headroom | 19:03 |
| @clarkb:matrix.org | I don't think we need to go huge like 1TB (which is where I'd be more concerned with an ssd) | 19:03 |
| @fungicide:matrix.org | created with `openstack --os-cloud=openstackci-rax --os-region-name=DFW volume create --type=SSD --size=300 zuul-db01.opendev.org/main02` | 19:04 |
| @status:opendev.org | @clarkb:matrix.org: finished sending notice | 19:04 |
| -@status:opendev.org- NOTICE: Zuul is reporting inconsistent results showing failures in Gerrit comments and in progress status in the Zuul UI. This is due to Zuul database issues that we are working to correct. | 19:04 | |
| @clarkb:matrix.org | I am going to pause zuul queues too | 19:05 |
| @clarkb:matrix.org | so that we stop losing data about results | 19:05 |
| @fungicide:matrix.org | `[Mon Sep 28 19:03:01 2026] blkfront: xvdc: barrier: enabled; persistent grants: disabled; indirect descriptors: enabled; bounce buffer: disabled;` (from dmesg on the server after server add volume) | 19:06 |
| @fungicide:matrix.org | `xvdc 202:32 0 300G 0 disk` (from lsblk) | 19:06 |
| @fungicide:matrix.org | i'll get it into the main vg | 19:06 |
| @clarkb:matrix.org | queues are paused | 19:06 |
| @clarkb:matrix.org | that had to pull the zuul-client image after our podman system reset | 19:07 |
| @fungicide:matrix.org | the main vg now includes a 100gb `/dev/xvdb1` and 300gb `/dev/xvdc1` | 19:09 |
| @fungicide:matrix.org | starting the pvmove now | 19:09 |
| @clarkb:matrix.org | note that mariadb is still running but I think we decided that wasn't an issue | 19:09 |
| @fungicide:matrix.org | `pvmove /dev/xvdb1 /dev/xvdc1` is underway in a root screen session on zuul-db01 now | 19:10 |
| @fungicide:matrix.org | yes, volumes can remain mounted and writable throughout this process | 19:10 |
| @fungicide:matrix.org | once it's fully moved all the extents to `/dev/xvdc1` i'll remove `/dev/xvdb1` from the vg and then extent the lv to the size of the remaining extents (which will make it the size of `/dev/xvdc1`) | 19:11 |
| @fungicide:matrix.org | after that we can grow the fs | 19:11 |
| @clarkb:matrix.org | sounds good. I'm around if I can help | 19:11 |
| @fungicide:matrix.org | nothing beats a monday fire drill | 19:12 |
| @fungicide:matrix.org | extents are about 40% of the way moved so far | 19:13 |
| @fungicide:matrix.org | shouldn't take too much longer | 19:13 |
| @fungicide:matrix.org | 75% | 19:16 |
| @fungicide:matrix.org | done, now removing the old pb | 19:18 |
| @fungicide:matrix.org | er, pv | 19:18 |
| @fungicide:matrix.org | extended the lv, now lvs reports `mariadb main -wi-ao---- <300.00g` | 19:20 |
| @fungicide:matrix.org | ran resize2fs and now df reports `/dev/mapper/main-mariadb 295G 98G 197G 34% /var/mariadb` | 19:20 |
| @clarkb:matrix.org | I see the same thing from df. Should I unpause the zuul result queues? | 19:21 |
| @fungicide:matrix.org | i have not yet detached and deleted the cinder volume, but we should be ready to resume db writes if there's no application level cleanup needed | 19:21 |
| @clarkb:matrix.org | I don't think there is application level cleanup. Instead I suspect that zuul simply error'd on mysql failures and proceeded along | 19:21 |
| @fungicide:matrix.org | convenient that was the only thing on the full fs, at least | 19:22 |
| @clarkb:matrix.org | in other words we lost the build completion info for those builds | 19:22 |
| @clarkb:matrix.org | the results reported to gerrit should be accurate. It will just be difficult (near impossible) to debug reported failuers | 19:22 |
| @clarkb:matrix.org | you good with me unpausing zuul now? | 19:22 |
| @fungicide:matrix.org | i am | 19:22 |
| @fungicide:matrix.org | please proceed | 19:22 |
| @clarkb:matrix.org | done | 19:22 |
| @fungicide:matrix.org | i'll get the old cinder volume recycled back into our quota | 19:23 |
| @fungicide:matrix.org | the old `zuul-db01.opendev.org/main01` cinder volume has been detached and deleted now | 19:24 |
| @fungicide:matrix.org | closing out the screen session now | 19:25 |
| @clarkb:matrix.org | https://zuul.opendev.org/t/openstack/build/d3146735cc664df2b18269c543d5111e this build started and completed after I paused/unpaused things | 19:26 |
| @clarkb:matrix.org | it renders in the UI indicating to me that things are happier moving forward | 19:27 |
| @fungicide:matrix.org | good, good | 19:28 |
| @clarkb:matrix.org | I'm going to eat lunch now | 19:29 |
| @clarkb:matrix.org | and then if things remain happy afterwards maybe a bike ride. Then get out the pre ptg is this week along with normal meeting cancellation email sent | 19:30 |
| @jim:acmegating.com | Clark: thanks! that entire thing managed to exactly overlap with my lunch acquisition, sorry | 19:35 |
| @clarkb:matrix.org | it ended up being relatively straightforward to fix (I think) | 19:36 |
| @clarkb:matrix.org | hopefully we didn't do anything different or wrong compared to what you would've done | 19:36 |
| @clarkb:matrix.org | ok I need to eat now. We might send a notice that builds can be rechecked if missing logs are a problem | 19:37 |
| @clarkb:matrix.org | but I don't think there is an easy way to find the log urls for logs without effort | 19:37 |
| @jim:acmegating.com | no definitely not worth that level of effort | 19:37 |
| @jim:acmegating.com | if there's some critical release promote job we need to find, we could probably dig it up | 19:37 |
| @fungicide:matrix.org | yeah, replacing/enlarging volumes and filesystems is engraved into my finger memory, didn't take long | 19:37 |
| @jim:acmegating.com | but not something we should do for the general case | 19:38 |
| @jim:acmegating.com | what's weird is that cacti says that /var/mariadb has been close to capacity for a long time... i think there's probably some weird internal allocation at play that makes it difficult to tell what the actual disk pressure is. | 19:39 |
| @jim:acmegating.com | fungi: thank you! | 19:42 |
| @fungicide:matrix.org | it was my pleasure, as always | 19:45 |
| @fungicide:matrix.org | i actually find playing around with kernel-level storage structures fun anyway | 19:45 |
| @fungicide:matrix.org | a fun one i have on my workstation is a pair of nvme drives in external usb type-c housings, i have them both luks2 encrypted and then a layer of luks2 integrity on top of that, then combined as a mdraid mirror, carved up with lvm2 and used to do backups of my important systems with periodic lvm snapshotting | 19:47 |
| @fungicide:matrix.org | and most fun was getting the systemd glue configured so it all comes up in the right order, unlocked with a keyfile on my workstation, auto-mounted on demand if any process reads from or writes to the mount path | 19:49 |
| @fungicide:matrix.org | (decrypted, integrity-checked, raided, activated, and mounted on the fly) | 19:50 |
| @fungicide:matrix.org | though the whole process causes the first read or write to block for about 10 seconds while it all spins up | 19:51 |
| @clarkb:matrix.org | it looks like things are still happy so I'm going to pop out for that bike ride now | 20:31 |
| @fungicide:matrix.org | have fun! | 20:31 |
| @mordred:waterwanders.com | ah! yes, that definitely makes sense. I was starting to consider db/disk issues but didn't have enough data yet to suggest anything. glad it wound up being straightforward \o/ | 21:53 |
| -@gerrit:opendev.org- Steve Baker proposed: | 21:58 | |
| - [openstack/diskimage-builder] 983813: Refactor 02-set-machine-id into an element https://review.opendev.org/c/openstack/diskimage-builder/+/983813 | ||
| - [openstack/diskimage-builder] 983814: Refactor 03-reset-bls-entries into an element https://review.opendev.org/c/openstack/diskimage-builder/+/983814 | ||
| - [openstack/diskimage-builder] 983815: Add tarball element for unprivileged container builds https://review.opendev.org/c/openstack/diskimage-builder/+/983815 | ||
| - [openstack/diskimage-builder] 984486: Skip local loop device creation for no-final-image builds https://review.opendev.org/c/openstack/diskimage-builder/+/984486 | ||
| - [openstack/diskimage-builder] 1007795: Add DIB_BIND_MOUNTS option for container environments https://review.opendev.org/c/openstack/diskimage-builder/+/1007795 | ||
| - [openstack/diskimage-builder] 1007796: Add dnf-assert element https://review.opendev.org/c/openstack/diskimage-builder/+/1007796 | ||
| - [openstack/diskimage-builder] 1007797: epel: allow epel-release to not be installed at all https://review.opendev.org/c/openstack/diskimage-builder/+/1007797 | ||
| - [openstack/diskimage-builder] 1007798: Check for python3-pyyaml before installing https://review.opendev.org/c/openstack/diskimage-builder/+/1007798 | ||
| - [openstack/diskimage-builder] 1007799: Handle passwd coming from shadow-utils on rhel-10 https://review.opendev.org/c/openstack/diskimage-builder/+/1007799 | ||
| - [openstack/diskimage-builder] 1007800: pip-and-virtrualenv don't depend on epel https://review.opendev.org/c/openstack/diskimage-builder/+/1007800 | ||
| @clarkb:matrix.org | message sent about the meeting cancellation and pre ptg schedule | 22:48 |
Generated by irclog2html.py 4.1.0 by Marius Gedminas - find it at https://mg.pov.lt/irclog2html/!