Friday, 2026-08-07

-@gerrit:opendev.org- Zuul merged on behalf of Clark Boylan: [opendev/system-config] 1000055: Run apache cache cleanups more often https://review.opendev.org/c/opendev/system-config/+/100005505:53
-@gerrit:opendev.org- Anil Shashikumar Belur proposed:08:03
- [opendev/system-config] 999195: Add an opt-in native nftables backend to the iptables role https://review.opendev.org/c/opendev/system-config/+/999195
- [opendev/system-config] 1000071: Support both firewall backends in testinfra https://review.opendev.org/c/opendev/system-config/+/1000071
@abelur:matrix.orgClark: respun the nftables change per your feedback ^^ It's two changes now (testinfra first, then the opt-in mechanism with08:37
the group empty), plus the system-config-run-base-iptables job you've suggested. One corerction in my reply: focal still defaults to
iptables-legacy, so the nft assertions don't work there - testinfra now probes iptables first. No rush, whenever you get a chance. thank you
@fungicide:matrix.orgapache proxy cache on mirror03.dfw.rax is staying down around 60gb again now, as intended, so the increased htcacheclean cronjob frequency seems to have helped13:34
@fungicide:matrix.orgi'm heading to lunch, but i'll be around after for a gerrit upgrade if folks are up for that (in which case someone should approve https://review.opendev.org/994938 so it will deploy)15:00
@clarkb:matrix.orgfungi: I think that rax dfw is disabled still right?15:03
@clarkb:matrix.orgSo hard to say if the increased cache cleanup interval is enough to keep up with CI.15:04
@clarkb:matrix.orgI can approve the Gerrit change in a bit15:04
@fungicide:matrix.orgoh good point about it still being disabled15:04
@fungicide:matrix.orgso yeah, can't draw any conclusions yet15:05
@clarkb:matrix.orgI have approved the gerrit upgrade change and now need to find breakfast15:37
@jim:acmegating.comwhat's the story on re-enabling dfw?  with the more frequent cleanup, are we just waiting for someone to be around to monitor and revert back out if it still fills up?  or something else?15:38
@clarkb:matrix.orgcorvus: yes I think its just that15:39
@clarkb:matrix.orgbasically put it back and monitor the disk consumption15:39
-@gerrit:opendev.org- Clark Boylan proposed: [opendev/zuul-providers] 1000119: Revert "Reapply "Disable rax-dfw again due to continued mirror issues"" https://review.opendev.org/c/opendev/zuul-providers/+/100011915:43
@clarkb:matrix.orgif we're ready to do that ^ this shoudl reenable things15:43
@jim:acmegating.com+2 but did not +w because i may need to head out in a bit15:59
@jim:acmegating.comi went to a gerrit meetup yesterday -- luca said 3.14 is much faster due to internal performance improvements16:00
-@gerrit:opendev.org- Zuul merged on behalf of Clark Boylan: [opendev/system-config] 994938: Update Gerrit images to 3.13.8 and 3.14.2 https://review.opendev.org/c/opendev/system-config/+/99493816:53
@dmsimard:matrix.orgHi, I'll spare you the details but the billing issue has been incredibly complicated :(16:56
@fungicide:matrix.orgdmsimard: thanks for looking into it!17:01
@dmsimard:matrix.orgIn spite of my best efforts (and the billing people involved) we might not be able to credit that pending invoice in time, but we did find out why this was all so terribly complicated17:04
@clarkb:matrix.orgfungi: the gerrit change landed and images promoted successfully per the zuul comment. I'm around today if we want to send a notice and then do an upgrade in an hour or so17:06
@clarkb:matrix.orgI do want to try and get out on a bike ride. Smoke has dissipated enough today that if I get out before its super hot I should be fine. But we can get this done first17:06
@fungicide:matrix.orgClark: i could announce an 18:00 utc upgrade via statusbot now if that works for you17:08
@clarkb:matrix.orgyup I think that works for me17:09
@fungicide:matrix.org#status notice The Gerrit service on review.opendev.openstack.org will be offline momentarily at 18:00 UTC for a patch upgrade (just under an hour from now), but should return within a few minutes.17:10
@status:opendev.org@fungicide:matrix.org: sending notice17:10
-@status:opendev.org- NOTICE: The Gerrit service on review.opendev.openstack.org will be offline momentarily at 18:00 UTC for a patch upgrade (just under an hour from now), but should return within a few minutes.17:13
@status:opendev.org@fungicide:matrix.org: finished sending notice17:13
@fungicide:matrix.orgi can't believe i typed review.opendev.openstack.org17:14
@fungicide:matrix.orgfingers on autopilot17:15
@fungicide:matrix.orgoh well, everyone knows what i meant17:15
@mnasiadka:matrix.orgIt's friday, maybe upgrading Gerrit is not what you want ;-)17:17
@fungicide:matrix.orgmaybe i can find a backstreet ripperdoc to upgrade my fingers17:18
@fungicide:matrix.orgweekend plans17:18
@fungicide:matrix.orgi've done a `docker-compose pull` and it also grabbed a new mariadb image from 3 days ago17:26
@fungicide:matrix.org10.11 (1a3c5f4fdaa7)17:27
@clarkb:matrix.orgthe gerrit image seems to match what is on quay at https://quay.io/repository/opendevorg/gerrit/manifest/sha256:36ab1340d326a116f5533ba99953e08c45b2717706961efe3f10f0c9bbfd166f17:30
@harbott.osism.tech:regio.chatdo we know when that "in time" will run out and what will be affected? maybe we can shut down some services beforehand, then?17:36
@fungicide:matrix.orgJens Harbott: it may make sense to stop booting nodes there, since that's the tenant that hosts the mirror servers17:55
@clarkb:matrix.orgfungi the typical set of caches look large on review03. Though it does seem like their growth is more restrained than with h2v117:55
@clarkb:matrix.orgso thats good but also we should probably claer them out per usual17:56
@dmsimard:matrix.orgI was told it was probably 7 days after the invoice which was created august 1st17:57
@clarkb:matrix.orgfungi: we ready ot proceed? I have attached to the screen17:59
@fungicide:matrix.org#status notice The Gerrit service on review.opendev.org is going offline momentarily for a patch upgrade, but should return within a few minutes.18:01
@status:opendev.org@fungicide:matrix.org: sending notice18:01
@fungicide:matrix.org`docker-compose -f /etc/gerrit-compose/docker-compose.yaml down && mv ~gerrit2/review_site/data/replication/ref-updates/waiting ~gerrit2/tmp/waiting_queue_2026-08-07 && rm ~gerrit2/review_site/cache/{gerrit_file_diff,git_file_diff,git_modified_files,modified_files,comment_context}-v2.* && docker-compose -f /etc/gerrit-compose/docker-compose.yaml up -d`18:01
@fungicide:matrix.orgi have that queued from a prior restart18:01
@dmsimard:matrix.org* I was told it was probably 7 days after the invoice which was created august 1st, impact would be to virtual machines and object storage18:02
@clarkb:matrix.orgthat command looks right to me and covers the appropriate caches18:03
@fungicide:matrix.orgokay, proceeding18:03
@fungicide:matrix.orgthe status notice is about done18:04
@status:opendev.org@fungicide:matrix.org: finished sending notice18:04
-@status:opendev.org- NOTICE: The Gerrit service on review.opendev.org is going offline momentarily for a patch upgrade, but should return within a few minutes.18:04
@fungicide:matrix.orgstopping took just shy of a minute that time18:05
@clarkb:matrix.orgthat was a really fast shutdown too18:05
@fungicide:matrix.orgdefinitely better18:05
@clarkb:matrix.orgweb ui loads for me and shows 3.13.8-dirty18:07
@clarkb:matrix.orghttps://review.opendev.org/c/opendev/zuul-providers/+/1000119/1/zuul.d/providers.yaml this diff loads for me18:07
@fungicide:matrix.orgyeah, looking good so far18:08
@clarkb:matrix.org`[2026-08-07T18:05:34.738Z] [main] INFO  com.google.gerrit.pgm.Daemon : Gerrit Code Review 3.13.8-dirty ready` if we need an actual timestamp for later18:08
@fungicide:matrix.orgshould we retrigger full replication?18:08
@clarkb:matrix.orgreplication should be fine. Do you mean reindexing? I'm still waiting for a new patchset to check replication though18:09
@fungicide:matrix.orgoh is it reindexing that we drop the pending queue for?18:10
@clarkb:matrix.orgfungi: no the waiting queue is dropped because there is a bug where it accumulates items to replicate that it cannot replicate due to permissions18:10
@clarkb:matrix.organd instead of marking those as unsatisfiable and moving on it accumulates them as things to do later. Then when we restart it attempts to process that backlog and produces tens of thousands of tracebacks as it does so18:10
@clarkb:matrix.orgwe reindex changes after restarts because there is a race conditions with creating new changes in gerrit during shutdown where you get things in git but not the index so then a followup patchset can create a second change with the same change id18:11
@fungicide:matrix.orgah okay, so we're not worried about pending replication getting dropped, just changes updated immediately prior to the restart that may have dropped pending index tasks18:11
@clarkb:matrix.orgyup18:12
@fungicide:matrix.orgyep that, okay18:12
@fungicide:matrix.orgit was gitea restarts where we used to be worried about holes in replication18:12
@fungicide:matrix.orgokay so run `gerrit index start changes --force` through the ssh api then?18:12
@clarkb:matrix.orghttps://opendev.org/openstack/sunbeam-charms/commit/173bd19bff542686662ea22c8772240a06e6b150 was replicated from https://review.opendev.org/c/openstack/sunbeam-charms/+/1000139 after the restart18:12
@clarkb:matrix.orggerrit show-queue lgtm (no backed up tasks for replication or cache pruning etc)18:13
@clarkb:matrix.orgso I think we can reindex changes now if we like. I think the major things to check after a restart lgtm18:13
@fungicide:matrix.orgreindexing in progress now18:14
@clarkb:matrix.orgfungi: once gerrit is settled we can reenable rax-dfw with https://review.opendev.org/c/opendev/zuul-providers/+/1000119 if we want to monitor the mirror disk consumption again18:19
@fungicide:matrix.orgah yep that'll be great, i'm around to keep an eye on it for a few hours at least18:20
@clarkb:matrix.orgI'm around today too though lunch is coming up and maybe a bike ride18:20
@clarkb:matrix.orgbut it is quick and easy to disable again if necessary so I'm not too concerned18:20
@fungicide:matrix.orggerrit's task queue is under 3k now18:26
@clarkb:matrix.organd the log file indicates reindexing is about 35% complete18:27
@fungicide:matrix.org2.5k tasks now. it's a little slower than i remember18:51
@clarkb:matrix.orgfungi: its piling up a backlog of things to index that are coming in18:53
@clarkb:matrix.orgso we're draining the queue just slightly faster than we are adding to it18:53
@clarkb:matrix.orgonce the main reindexing completes it should catch up quickly since the total number of changes in the queue to reindex is small after that point18:53
@fungicide:matrix.orgi guess we caught openstack at its feature freeziest18:55
@clarkb:matrix.orgreindexing is complete with the expected 3 failures18:56
@fungicide:matrix.orgyeah that was a weird sudden jump18:56
@clarkb:matrix.orgall of the index future for change updates during the reindex completely quickly as expected18:56
@fungicide:matrix.orgi'll go ahead and approve the rax-dfw reenablement18:57
@clarkb:matrix.orgthey represented say 2k things to reindex vs the 1 million we were processing18:57
@clarkb:matrix.orgso the task count is not directly comparable18:57
-@gerrit:opendev.org- Zuul merged on behalf of Clark Boylan: [opendev/zuul-providers] 1000119: Revert "Reapply "Disable rax-dfw again due to continued mirror issues"" https://review.opendev.org/c/opendev/zuul-providers/+/100011918:57
@clarkb:matrix.orgthanks. I'm going to grab lunch now18:57
@fungicide:matrix.orgmakes sense18:57
@fungicide:matrix.orgi'll keep a tab on the cache usage for that mirror18:58
@fungicide:matrix.orgclosed out the root screen session on review03 now18:58
@fungicide:matrix.orghtcacheclean seems to be keeping up so far19:31
@fungicide:matrix.org56 nodes in use for that region so far according to grafana19:38
@fungicide:matrix.organd as of yet the cache usage is still hovering at 60gb (the clean target)19:39
@fungicide:matrix.organ hour later and the cache size hasn't budged, not even as much as i expected in between htcacheclean runs, so probably whatever jobs were blowing out the cache yesterday aren't running today, or at least not at the same volume20:05
@fungicide:matrix.orgcould just be because it's already the weekend in apac/emea20:05
@dmsimard:matrix.orgHi, I am coming back with good news, they figured something out and the invoice is credited, everything is good20:18
@dmsimard:matrix.orgWe will follow up to fix what made this so complicated so that this isn't so troublesome next time, sorry about that20:19
@fungicide:matrix.orgthanks again dmsimard!20:20
@fungicide:matrix.orgit's been two hours now and i still haven't seen the apache cache on the rax-dfw mirror go above 60gb, not even just before htcacheclean runs, so i think we're unlikely to get a good confirmation until activity picks back up next week21:00
@fungicide:matrix.orgalso for reference, htcacheclean is taking approximately 7 minutes to complete a pass there at present21:01
@fungicide:matrix.orgnow we're at three hours and the cache there is still fine. i'm going to call it an evening, but i set myself a reminder to check it again first thing monday morning when we'll hopefully have a better idea of how it's faring under load22:02
@clarkb:matrix.orgthanks and have a good weekend22:04
@fungicide:matrix.orgyou too!22:05

Generated by irclog2html.py 4.1.0 by Marius Gedminas - find it at https://mg.pov.lt/irclog2html/!