| @mnasiadka:matrix.org | Having a look now, it seems reprepro crashed and there was an old lock - removed it and rerunning the reprepro sync in tmux on mirror-update | 07:00 |
|---|---|---|
| @mnasiadka:matrix.org | Hmm: | 07:01 |
| ``` | ||
| Warning parsing /etc/reprepro/ubuntu/updates, line 21: unknown architecture 'arm64' will be ignored! | ||
| Warning parsing /etc/reprepro/ubuntu/updates, line 29: unknown architecture 'arm64' will be ignored! | ||
| ``` | ||
| @mnasiadka:matrix.org | Ok, the sync plus export plus vos release succeeded - I'll run it again just to check everything is all right. | 08:34 |
| @harbott.osism.tech:regio.chat | mnasiadka: thanks, did you find something in the logs as to why this happened? my own search for that didn't come up with anything | 08:53 |
| @mnasiadka:matrix.org | Jens Harbott: nope, nothing | 08:55 |
| @mnasiadka:matrix.org | And that warning unknown architecture seems to be weird, because ubuntu ports are synced | 09:08 |
| @vurmil:matrix.org | It looks like the mirror is still out of sync. apt-get install sqlite3 fails because sqlite3 expects libsqlite3-0 (= 3.45.1-1ubuntu2.7), but version 3.45.1-1ubuntu2.8 is present: | 09:17 |
| @vurmil:matrix.org | manila: sqlite3 : Depends: libsqlite3-0 (= 3.45.1-1ubuntu2.7) but 3.45.1-1ubuntu2.8 is to be installed | 09:17 |
| @mnasiadka:matrix.org | Vurmil: and when did the job start? | 09:33 |
| @harbott.osism.tech:regio.chat | mnasiadka: the ports mirror is a dedicated cron job, and it has the opposite `Warning parsing /etc/reprepro/ubuntu-ports/updates, line 5: unknown architecture 'amd64' will be ignored!`, so I think this is fine somehow, without knowing all the details | 10:41 |
| @harbott.osism.tech:regio.chat | also, there we still some issues in the latest ubuntu mirror cron run, did you see the same in your manual runs? | 10:41 |
| @harbott.osism.tech:regio.chat | ``` | 10:43 |
| Description-md5 of package 'libreoffice-gnome' does not match the md5 of the description found in the .deb ('c8ebe7be7e73db44a77d5bc22950cfdd' != 'c468d8fd91ddd9ba3bc4401d295be6e1')! | ||
| ... | ||
| files lost their last reference. | 00:07 | |
| (dumpunreferenced lists such files, use deleteunreferenced to delete them.) | ||
| ``` | ||
| @mnasiadka:matrix.org | Yes, a lot of things like: | 10:46 |
| ``` | ||
| Description-md5 of package 'libpipewire-0.3-modules-x11' does not match | ||
| the md5 of the description found in the .deb | ||
| ('bc8ef3ab07b2505388b17ccf79d5a7ee' != 'a533b054e32d10bd276ec8504194efbb')! | ||
| ``` | ||
| @mnasiadka:matrix.org | We should most probably monitor that closely. | 10:46 |
| @mnasiadka:matrix.org | btw, today magically my connections to Gerrit work fine everytime | 10:58 |
| @mnasiadka:matrix.org | So maybe that transit provider has fixed the link | 10:58 |
| @mnasiadka:matrix.org | I think I was too optimistic :) | 11:41 |
| -@gerrit:opendev.org- Armin M proposed: [openstack/diskimage-builder] 1006200: Oracle Linux support for diskimage-builder. https://review.opendev.org/c/openstack/diskimage-builder/+/1006200 | 12:16 | |
| -@gerrit:opendev.org- Zuul merged on behalf of Ivan Anfimov: [openstack/diskimage-builder] 1002749: Replaced usage outdate egrep to grep https://review.opendev.org/c/openstack/diskimage-builder/+/1002749 | 13:53 | |
| -@gerrit:opendev.org- Zuul merged on behalf of Michael Mokricky: [openstack/diskimage-builder] 1001833: Fetch a checkout ref into refs/tags, not a branch https://review.opendev.org/c/openstack/diskimage-builder/+/1001833 | 13:53 | |
| @clarkb:matrix.org | mnasiadka Jens Harbott I think that if we get a short read/write or bit flip of a deb package that reprepro won't replace it automatically when hashes mismatch. Instead it will just complain. I want to say fungi has manually replaced those files in the past to correct this but I'm not certain | 13:59 |
| @clarkb:matrix.org | I am surprised that reprepro succeeds in that case though (I thought it would fail which would prevent vos release of the inconsistent state) | 14:00 |
| @mnasiadka:matrix.org | Clark: Well, second run didn’t have those errors, so it seems it cleared on it’s own. | 15:57 |
| @clarkb:matrix.org | interesting. Maybe we should spot check one of them to be sure but ya if the errors went away then it probably sorted it out on its own | 15:58 |
| @clarkb:matrix.org | Looking at zuul-providers our periodic jobs do continue to fail, but the recent errors are not due to running out of disk space | 16:04 |
| @clarkb:matrix.org | one arm64 image build failed when talking to the local mirror (connection timed out), then the x86 noble and resolute builds failed due to http 500 error responses from flex sjc3 swift when uploading to the intermediate storage location | 16:04 |
| @clarkb:matrix.org | So I think Anil's change has improved reliability, but now we're finding the next set of issues. I haven't looked at the rax dfw specific issue yet as I was trying to get an understanding of the overall situation after Anil's change landed | 16:05 |
| -@gerrit:opendev.org- Zuul merged on behalf of Clark Boylan: [openstack/diskimage-builder] 994017: Revert "Disable centos/10-stream-build-succeeds functest" https://review.opendev.org/c/openstack/diskimage-builder/+/994017 | 16:27 | |
| @clarkb:matrix.org | corvus: I think the launchers are running out of disk, but I'm not sure under which path yet | 16:55 |
| @clarkb:matrix.org | this seems to be causing failed uploads (I don't yet know if that is the source of the rax dfw failure, perhaps it was always at the tail end of the disk usage and failing more consistentyl than others or maybe they are independent issues) | 16:55 |
| @clarkb:matrix.org | `launcher.temp_dir` is set to /var/lib/zuul/tmp which is on a dedicated fs/device. We also bind mount /opt/zuul-launcher-tmp into the container and /opt is on the ephemeral device with its own slightly smaller fs | 16:57 |
| @clarkb:matrix.org | currently /var/lib/zuul/tmp is consuming about 71GB with ~7.6GB free on that fs | 16:58 |
| @clarkb:matrix.org | ok zl01 has a lot more disk free than zl02. But both of them appear to have leaked files from 2025 just under /var/lib/zuul/tmp | 16:59 |
| @clarkb:matrix.org | zl02 has two files at 20GB each and zl01 has two files but the sizes are different | 16:59 |
| @clarkb:matrix.org | I suspect that those files are properly leaked and can just go away (though I'm still investigating and not making any changes) | 17:00 |
| @clarkb:matrix.org | both servers are up for the expected ~6 days so we are not cleaning things up on startup like we do with the executor build dir | 17:01 |
| @clarkb:matrix.org | just pointing that out as it indicates restarting the service is unlikely to help | 17:02 |
| @clarkb:matrix.org | We also have files under `zl02:/var/lib/zuul/tmp/zuul-launcher` like 97ae2f4efbe149de9dd8bfabe084e02c which seems to map to https://zuul.opendev.org/t/opendev/image/ubuntu-noble which has a failed upload | 17:04 |
| @clarkb:matrix.org | I think it less likely that these files under /var/lib/zuul/tmp/zuul-launcher are properly leaked given that evidence but they are also consuming a fair bit of disk | 17:05 |
| @clarkb:matrix.org | so I think there may be two things happening here. The first is properly leaked content from 2025 at /var/lib/zuul/tmp reducing our total usable disk space. Then under /var/lib/zuul/tmp/zuul-launcher we are maybe not as efficient as we could/should be when handling failed uploads? That 97ae... file was from the 14th. Looking in today and yesterday's logs that 97ae... string and the dc4a7703102f46d8932367771ddd2f67 id for the failed upload do not show up implying that we aren't retrying things? So maybe we shouldn't be hanging onto the 24GB vhd file on disk? | 17:07 |
| @clarkb:matrix.org | corvus: I think the "leaking" happening in /var/lib/zuul/tmp/zuul-launcher is due to a bug error handling for image files in the launcher. It looks like we hit an error trying to decompressing the file `zstd -dq` failed with rc 70. There is evidence we try to delete the .zst file which throws and error because the file doesn't exist. I think we failed to decompress so assume the compressed file exists but it actually got far enough to decompress something and remove the original file while still failing. Then our cleanups fail because we try to unlink a file that doesn't exist and don't catch taht error and don't try to remove the decompressed file which results in a leak | 17:14 |
| @clarkb:matrix.org | let me put together a paste log of that. And if that makes sense I think there are two issues to address here. The first is the really old 2025 files if those can be removed lets just remove them. Then separately the launcher probably needs to be more cautious around unlinking things and also try to unlink all the potential file outcome for zstd decompression to avoid leaks | 17:15 |
| @clarkb:matrix.org | Third we may possibly consider startup cleanups to try and keep things pruned like the executor? | 17:15 |
| @clarkb:matrix.org | as a side note state: ready on a failed upload is a bit confusing | 17:19 |
| @clarkb:matrix.org | corvus: https://paste.opendev.org/show/bUMtDbWD42poYsLrCsSe/ here is the log paste so that log rotations don't prevent future debugging | 17:19 |
| @clarkb:matrix.org | also it is entirely possible that the dfw failures are different, but I kinda want to address this iteratively and fix the most obvious errors first and see where that gets us | 17:20 |
| @clarkb:matrix.org | hrm that early "Deleted" log entry is really confusing though. Unfortunately we log "Deleted %s" against path in four different locations so it isn't immediately clear to me why we deleted there. But I suspect what may have happened here is that we deleted the file under zstd possibly? | 17:23 |
| @clarkb:matrix.org | so ztsd is running and it isn't fast then somethign decides to delete the file which causes zstd to error and we fail? Then we don't clean up what was decompressed as we only attempt tp cleanup the compressed file | 17:24 |
| @clarkb:matrix.org | fwiw there is a _cleanTempdir() routine so maybe a service restart would help | 17:32 |
| @clarkb:matrix.org | corvus: reading the log and cross checking against the code it is weird that all of the Deleted logs appear to happen after the _handleCompression routine runs zstd. But in the log we have the Deleted line happening before the zstd decompression related exceptions. The exception to that is the _cleanTempdir routine but it appears to only run on startup and the launcher has been running for 6 days and this occurred 4 days ago | 17:41 |
| @clarkb:matrix.org | `remote: https://review.opendev.org/c/zuul/zuul/+/1006310 Cleanup source and dest file on zstd errors` this is my initial thoughts on how we might mitigate the problem, but it still isn't immediately clear to me why we got the deletion before the exceptions as that may explain the underlying cause of the zstd failure | 17:48 |
| @jim:acmegating.com | Clark: i'll take a look | 18:22 |
| @jim:acmegating.com | Clark: some of the deleted log messages are in except handlers before a raise; that should explain why it shows up before the exception | 18:24 |
| @jim:acmegating.com | Clark: i deleted the 2025 files and approved the zuul change | 18:29 |
| @clarkb:matrix.org | Cool. In theory after the next automated restart the other leaked files will clear out and then we can see if this problem returns. I have a hunch simply fixing the disk space problem may be a big part of it | 18:37 |
| @jim:acmegating.com | Clark: do you think that could be the source of the rax errors? because we were getting some pretty clear server-side errors.... or are you thinking that maybe the server errors were transient and that's just lost in the noise of disk errors? | 18:39 |
| @clarkb:matrix.org | Ya I think either it was a transient cloud thing and now we have new errors. Or this error is masking the cloud side issue and we'll return to that when we fix the disk space problem. But it's hard to debug that other issue right now because we have the disk space issue | 18:40 |
| @clarkb:matrix.org | That noble image did upload to dfw though | 18:40 |
| @jim:acmegating.com | ack | 18:40 |
| @clarkb:matrix.org | So I think dfw generally works and it wouldn't surprise me if this was fallout from some disk space problems. Due to job ordering or similar | 18:41 |
| -@gerrit:opendev.org- Armin M proposed: [openstack/diskimage-builder] 1006200: Oracle Linux support for diskimage-builder. https://review.opendev.org/c/openstack/diskimage-builder/+/1006200 | 20:13 | |
| @clarkb:matrix.org | that zuul launcher fixup is in the gate and still happy. I'm goign to pop out for a bike ride while I can | 21:36 |
| -@gerrit:opendev.org- Armin M proposed: [openstack/diskimage-builder] 1006200: Oracle Linux support for diskimage-builder. https://review.opendev.org/c/openstack/diskimage-builder/+/1006200 | 22:22 | |
Generated by irclog2html.py 4.1.0 by Marius Gedminas - find it at https://mg.pov.lt/irclog2html/!