| -@gerrit:opendev.org- Zuul merged on behalf of Clark Boylan: [opendev/system-config] 1000055: Run apache cache cleanups more often https://review.opendev.org/c/opendev/system-config/+/1000055 | 05:53 | |
| -@gerrit:opendev.org- Anil Shashikumar Belur proposed: | 08:03 | |
| - [opendev/system-config] 999195: Add an opt-in native nftables backend to the iptables role https://review.opendev.org/c/opendev/system-config/+/999195 | ||
| - [opendev/system-config] 1000071: Support both firewall backends in testinfra https://review.opendev.org/c/opendev/system-config/+/1000071 | ||
| @abelur:matrix.org | Clark: respun the nftables change per your feedback ^^ It's two changes now (testinfra first, then the opt-in mechanism with | 08:37 |
|---|---|---|
| the group empty), plus the system-config-run-base-iptables job you've suggested. One corerction in my reply: focal still defaults to | ||
| iptables-legacy, so the nft assertions don't work there - testinfra now probes iptables first. No rush, whenever you get a chance. thank you | ||
| @fungicide:matrix.org | apache proxy cache on mirror03.dfw.rax is staying down around 60gb again now, as intended, so the increased htcacheclean cronjob frequency seems to have helped | 13:34 |
| @fungicide:matrix.org | i'm heading to lunch, but i'll be around after for a gerrit upgrade if folks are up for that (in which case someone should approve https://review.opendev.org/994938 so it will deploy) | 15:00 |
| @clarkb:matrix.org | fungi: I think that rax dfw is disabled still right? | 15:03 |
| @clarkb:matrix.org | So hard to say if the increased cache cleanup interval is enough to keep up with CI. | 15:04 |
| @clarkb:matrix.org | I can approve the Gerrit change in a bit | 15:04 |
| @fungicide:matrix.org | oh good point about it still being disabled | 15:04 |
| @fungicide:matrix.org | so yeah, can't draw any conclusions yet | 15:05 |
| @clarkb:matrix.org | I have approved the gerrit upgrade change and now need to find breakfast | 15:37 |
| @jim:acmegating.com | what's the story on re-enabling dfw? with the more frequent cleanup, are we just waiting for someone to be around to monitor and revert back out if it still fills up? or something else? | 15:38 |
| @clarkb:matrix.org | corvus: yes I think its just that | 15:39 |
| @clarkb:matrix.org | basically put it back and monitor the disk consumption | 15:39 |
| -@gerrit:opendev.org- Clark Boylan proposed: [opendev/zuul-providers] 1000119: Revert "Reapply "Disable rax-dfw again due to continued mirror issues"" https://review.opendev.org/c/opendev/zuul-providers/+/1000119 | 15:43 | |
| @clarkb:matrix.org | if we're ready to do that ^ this shoudl reenable things | 15:43 |
| @jim:acmegating.com | +2 but did not +w because i may need to head out in a bit | 15:59 |
| @jim:acmegating.com | i went to a gerrit meetup yesterday -- luca said 3.14 is much faster due to internal performance improvements | 16:00 |
| -@gerrit:opendev.org- Zuul merged on behalf of Clark Boylan: [opendev/system-config] 994938: Update Gerrit images to 3.13.8 and 3.14.2 https://review.opendev.org/c/opendev/system-config/+/994938 | 16:53 | |
| @dmsimard:matrix.org | Hi, I'll spare you the details but the billing issue has been incredibly complicated :( | 16:56 |
| @fungicide:matrix.org | dmsimard: thanks for looking into it! | 17:01 |
| @dmsimard:matrix.org | In spite of my best efforts (and the billing people involved) we might not be able to credit that pending invoice in time, but we did find out why this was all so terribly complicated | 17:04 |
| @clarkb:matrix.org | fungi: the gerrit change landed and images promoted successfully per the zuul comment. I'm around today if we want to send a notice and then do an upgrade in an hour or so | 17:06 |
| @clarkb:matrix.org | I do want to try and get out on a bike ride. Smoke has dissipated enough today that if I get out before its super hot I should be fine. But we can get this done first | 17:06 |
| @fungicide:matrix.org | Clark: i could announce an 18:00 utc upgrade via statusbot now if that works for you | 17:08 |
| @clarkb:matrix.org | yup I think that works for me | 17:09 |
| @fungicide:matrix.org | #status notice The Gerrit service on review.opendev.openstack.org will be offline momentarily at 18:00 UTC for a patch upgrade (just under an hour from now), but should return within a few minutes. | 17:10 |
| @status:opendev.org | @fungicide:matrix.org: sending notice | 17:10 |
| -@status:opendev.org- NOTICE: The Gerrit service on review.opendev.openstack.org will be offline momentarily at 18:00 UTC for a patch upgrade (just under an hour from now), but should return within a few minutes. | 17:13 | |
| @status:opendev.org | @fungicide:matrix.org: finished sending notice | 17:13 |
| @fungicide:matrix.org | i can't believe i typed review.opendev.openstack.org | 17:14 |
| @fungicide:matrix.org | fingers on autopilot | 17:15 |
| @fungicide:matrix.org | oh well, everyone knows what i meant | 17:15 |
| @mnasiadka:matrix.org | It's friday, maybe upgrading Gerrit is not what you want ;-) | 17:17 |
| @fungicide:matrix.org | maybe i can find a backstreet ripperdoc to upgrade my fingers | 17:18 |
| @fungicide:matrix.org | weekend plans | 17:18 |
| @fungicide:matrix.org | i've done a `docker-compose pull` and it also grabbed a new mariadb image from 3 days ago | 17:26 |
| @fungicide:matrix.org | 10.11 (1a3c5f4fdaa7) | 17:27 |
| @clarkb:matrix.org | the gerrit image seems to match what is on quay at https://quay.io/repository/opendevorg/gerrit/manifest/sha256:36ab1340d326a116f5533ba99953e08c45b2717706961efe3f10f0c9bbfd166f | 17:30 |
| @harbott.osism.tech:regio.chat | do we know when that "in time" will run out and what will be affected? maybe we can shut down some services beforehand, then? | 17:36 |
| @fungicide:matrix.org | Jens Harbott: it may make sense to stop booting nodes there, since that's the tenant that hosts the mirror servers | 17:55 |
| @clarkb:matrix.org | fungi the typical set of caches look large on review03. Though it does seem like their growth is more restrained than with h2v1 | 17:55 |
| @clarkb:matrix.org | so thats good but also we should probably claer them out per usual | 17:56 |
| @dmsimard:matrix.org | I was told it was probably 7 days after the invoice which was created august 1st | 17:57 |
| @clarkb:matrix.org | fungi: we ready ot proceed? I have attached to the screen | 17:59 |
| @fungicide:matrix.org | #status notice The Gerrit service on review.opendev.org is going offline momentarily for a patch upgrade, but should return within a few minutes. | 18:01 |
| @status:opendev.org | @fungicide:matrix.org: sending notice | 18:01 |
| @fungicide:matrix.org | `docker-compose -f /etc/gerrit-compose/docker-compose.yaml down && mv ~gerrit2/review_site/data/replication/ref-updates/waiting ~gerrit2/tmp/waiting_queue_2026-08-07 && rm ~gerrit2/review_site/cache/{gerrit_file_diff,git_file_diff,git_modified_files,modified_files,comment_context}-v2.* && docker-compose -f /etc/gerrit-compose/docker-compose.yaml up -d` | 18:01 |
| @fungicide:matrix.org | i have that queued from a prior restart | 18:01 |
| @dmsimard:matrix.org | * I was told it was probably 7 days after the invoice which was created august 1st, impact would be to virtual machines and object storage | 18:02 |
| @clarkb:matrix.org | that command looks right to me and covers the appropriate caches | 18:03 |
| @fungicide:matrix.org | okay, proceeding | 18:03 |
| @fungicide:matrix.org | the status notice is about done | 18:04 |
| @status:opendev.org | @fungicide:matrix.org: finished sending notice | 18:04 |
| -@status:opendev.org- NOTICE: The Gerrit service on review.opendev.org is going offline momentarily for a patch upgrade, but should return within a few minutes. | 18:04 | |
| @fungicide:matrix.org | stopping took just shy of a minute that time | 18:05 |
| @clarkb:matrix.org | that was a really fast shutdown too | 18:05 |
| @fungicide:matrix.org | definitely better | 18:05 |
| @clarkb:matrix.org | web ui loads for me and shows 3.13.8-dirty | 18:07 |
| @clarkb:matrix.org | https://review.opendev.org/c/opendev/zuul-providers/+/1000119/1/zuul.d/providers.yaml this diff loads for me | 18:07 |
| @fungicide:matrix.org | yeah, looking good so far | 18:08 |
| @clarkb:matrix.org | `[2026-08-07T18:05:34.738Z] [main] INFO com.google.gerrit.pgm.Daemon : Gerrit Code Review 3.13.8-dirty ready` if we need an actual timestamp for later | 18:08 |
| @fungicide:matrix.org | should we retrigger full replication? | 18:08 |
| @clarkb:matrix.org | replication should be fine. Do you mean reindexing? I'm still waiting for a new patchset to check replication though | 18:09 |
| @fungicide:matrix.org | oh is it reindexing that we drop the pending queue for? | 18:10 |
| @clarkb:matrix.org | fungi: no the waiting queue is dropped because there is a bug where it accumulates items to replicate that it cannot replicate due to permissions | 18:10 |
| @clarkb:matrix.org | and instead of marking those as unsatisfiable and moving on it accumulates them as things to do later. Then when we restart it attempts to process that backlog and produces tens of thousands of tracebacks as it does so | 18:10 |
| @clarkb:matrix.org | we reindex changes after restarts because there is a race conditions with creating new changes in gerrit during shutdown where you get things in git but not the index so then a followup patchset can create a second change with the same change id | 18:11 |
| @fungicide:matrix.org | ah okay, so we're not worried about pending replication getting dropped, just changes updated immediately prior to the restart that may have dropped pending index tasks | 18:11 |
| @clarkb:matrix.org | yup | 18:12 |
| @fungicide:matrix.org | yep that, okay | 18:12 |
| @fungicide:matrix.org | it was gitea restarts where we used to be worried about holes in replication | 18:12 |
| @fungicide:matrix.org | okay so run `gerrit index start changes --force` through the ssh api then? | 18:12 |
| @clarkb:matrix.org | https://opendev.org/openstack/sunbeam-charms/commit/173bd19bff542686662ea22c8772240a06e6b150 was replicated from https://review.opendev.org/c/openstack/sunbeam-charms/+/1000139 after the restart | 18:12 |
| @clarkb:matrix.org | gerrit show-queue lgtm (no backed up tasks for replication or cache pruning etc) | 18:13 |
| @clarkb:matrix.org | so I think we can reindex changes now if we like. I think the major things to check after a restart lgtm | 18:13 |
| @fungicide:matrix.org | reindexing in progress now | 18:14 |
| @clarkb:matrix.org | fungi: once gerrit is settled we can reenable rax-dfw with https://review.opendev.org/c/opendev/zuul-providers/+/1000119 if we want to monitor the mirror disk consumption again | 18:19 |
| @fungicide:matrix.org | ah yep that'll be great, i'm around to keep an eye on it for a few hours at least | 18:20 |
| @clarkb:matrix.org | I'm around today too though lunch is coming up and maybe a bike ride | 18:20 |
| @clarkb:matrix.org | but it is quick and easy to disable again if necessary so I'm not too concerned | 18:20 |
| @fungicide:matrix.org | gerrit's task queue is under 3k now | 18:26 |
| @clarkb:matrix.org | and the log file indicates reindexing is about 35% complete | 18:27 |
| @fungicide:matrix.org | 2.5k tasks now. it's a little slower than i remember | 18:51 |
| @clarkb:matrix.org | fungi: its piling up a backlog of things to index that are coming in | 18:53 |
| @clarkb:matrix.org | so we're draining the queue just slightly faster than we are adding to it | 18:53 |
| @clarkb:matrix.org | once the main reindexing completes it should catch up quickly since the total number of changes in the queue to reindex is small after that point | 18:53 |
| @fungicide:matrix.org | i guess we caught openstack at its feature freeziest | 18:55 |
| @clarkb:matrix.org | reindexing is complete with the expected 3 failures | 18:56 |
| @fungicide:matrix.org | yeah that was a weird sudden jump | 18:56 |
| @clarkb:matrix.org | all of the index future for change updates during the reindex completely quickly as expected | 18:56 |
| @fungicide:matrix.org | i'll go ahead and approve the rax-dfw reenablement | 18:57 |
| @clarkb:matrix.org | they represented say 2k things to reindex vs the 1 million we were processing | 18:57 |
| @clarkb:matrix.org | so the task count is not directly comparable | 18:57 |
| -@gerrit:opendev.org- Zuul merged on behalf of Clark Boylan: [opendev/zuul-providers] 1000119: Revert "Reapply "Disable rax-dfw again due to continued mirror issues"" https://review.opendev.org/c/opendev/zuul-providers/+/1000119 | 18:57 | |
| @clarkb:matrix.org | thanks. I'm going to grab lunch now | 18:57 |
| @fungicide:matrix.org | makes sense | 18:57 |
| @fungicide:matrix.org | i'll keep a tab on the cache usage for that mirror | 18:58 |
| @fungicide:matrix.org | closed out the root screen session on review03 now | 18:58 |
| @fungicide:matrix.org | htcacheclean seems to be keeping up so far | 19:31 |
| @fungicide:matrix.org | 56 nodes in use for that region so far according to grafana | 19:38 |
| @fungicide:matrix.org | and as of yet the cache usage is still hovering at 60gb (the clean target) | 19:39 |
| @fungicide:matrix.org | an hour later and the cache size hasn't budged, not even as much as i expected in between htcacheclean runs, so probably whatever jobs were blowing out the cache yesterday aren't running today, or at least not at the same volume | 20:05 |
| @fungicide:matrix.org | could just be because it's already the weekend in apac/emea | 20:05 |
| @dmsimard:matrix.org | Hi, I am coming back with good news, they figured something out and the invoice is credited, everything is good | 20:18 |
| @dmsimard:matrix.org | We will follow up to fix what made this so complicated so that this isn't so troublesome next time, sorry about that | 20:19 |
| @fungicide:matrix.org | thanks again dmsimard! | 20:20 |
| @fungicide:matrix.org | it's been two hours now and i still haven't seen the apache cache on the rax-dfw mirror go above 60gb, not even just before htcacheclean runs, so i think we're unlikely to get a good confirmation until activity picks back up next week | 21:00 |
| @fungicide:matrix.org | also for reference, htcacheclean is taking approximately 7 minutes to complete a pass there at present | 21:01 |
| @fungicide:matrix.org | now we're at three hours and the cache there is still fine. i'm going to call it an evening, but i set myself a reminder to check it again first thing monday morning when we'll hopefully have a better idea of how it's faring under load | 22:02 |
| @clarkb:matrix.org | thanks and have a good weekend | 22:04 |
| @fungicide:matrix.org | you too! | 22:05 |
Generated by irclog2html.py 4.1.0 by Marius Gedminas - find it at https://mg.pov.lt/irclog2html/!