Thursday, 2026-10-08

@jim:acmegating.comyeah, test failed with the same error, so i'm going to set an autohold00:12
@jim:acmegating.comwill be easier to inspect stuff like that on the hosts00:12
-@gerrit:opendev.org- James E. Blair https://matrix.to/#/@jim:acmegating.com proposed:14:43
- [opendev/system-config] 1009351: Use private ipv4 for zookeeper quorum configuration in testing https://review.opendev.org/c/opendev/system-config/+/1009351
- [opendev/system-config] 1009279: WIP: Test 2-node zk cluster https://review.opendev.org/c/opendev/system-config/+/1009279
- [opendev/system-config] 1009387: Add SAN to certs generated by the opendev CA https://review.opendev.org/c/opendev/system-config/+/1009387
@clarkb:matrix.orgWas the previous version writing the SAN then removing it because that is an extension?14:52
@jim:acmegating.comyep!  from my held node i learned all 3 of those had errors:14:53
1) the nested inventory thing is a mapping, we can't use a template in there, we have to just use a variable name.
2) minor error in an unrelated test
3) the ca has to be configured to let the extension through
@jim:acmegating.comunfortunately, the held nodes were on rax flex, so the actual error was the private_ipv4, but i went ahead and investigated the contents of the certs anyway, and then did local testing to verify the fix.14:54
@jim:acmegating.comi'm optimistic all 3 are fixed, but i also set another autohold :)14:54
@clarkb:matrix.orgCool I'll be at my desk soon and can properly review them shortly14:55
@fungicide:matrix.orgseems like zuul is getting spammed with reconfigurations today, every time i look there's an event pileup and a reconfiguration a few minutes prior15:02
@fungicide:matrix.org(in the openstack tenant)15:03
@jim:acmegating.comi suspect the actual cause is a couple of steps away -- reconfigurations are quite fast now.  but they only happen at the end of pipeline processing, so if we're processing a huge check pipeline (because someone uploaded a bunch of changes at once), we'll wait for that to complete.  when someone uploads a bunch of changes at once, we start a bunch of jobs at once, and when we get the build started event, we query the database for the estimated time.  the time it takes to query the db is slowly increasing, it's long enough to be noticeable now.  that makes pipeline processing take longer.15:07
@fungicide:matrix.orggot it15:09
@fungicide:matrix.orgso the event queue backlog isn't because event processing is paused for reconfiguration, there just happens to be another reconfiguration triggered by something every time the queue gets burned back down at the moment15:10
@jim:acmegating.comthat's my theory, but i think i may be wrong15:11
@jim:acmegating.com2026-10-08 15:10:15,759 INFO zuul.Scheduler: [e: ca1a423c97a94d6da0360ccb472251d1] Tenant reconfiguration complete for openstack (duration: 198.401 seconds)15:11
@jim:acmegating.comthat took an unexpectedly long time15:12
@fungicide:matrix.orgyeah, i've been waiting for a promote pipeline item to get enqueued for over 10 minutes since the corresponding change merged, longer than usual15:12
@fungicide:matrix.orgblam, just appeared, so that was approximately 15 minutes to act on an internal trigger event15:15
@fungicide:matrix.orger, not an internal trigger, gerrit `change-merged` event i guess15:16
@fungicide:matrix.orgaha, starlingx dumped 70+ tags into gerrit all at once15:17
@jim:acmegating.comwell starlingx is tagging a lot of stuff right now it looks like, so this likely represents a workload change, not a systemic problem.  it's still curious why it's taking so long, but i suspect we'll get through it15:17
@jim:acmegating.comall at once is good!  they'll get combined into one reconfiguration event15:17
@fungicide:matrix.orgyeah, just seems probably more loaded than usual and trying to publish a security advisory at a scheduled time collided with unexpected delays from that. it happens15:18
@jim:acmegating.comhttps://zuul.opendev.org/t/openstack/system-events15:18
@jim:acmegating.comyou can see that's already happening15:18
@jim:acmegating.coma couple of reconfigs for standalone events, then more and more are combined15:18
@fungicide:matrix.orgnice batching, yep15:19
-@gerrit:opendev.org- Clark Boylan proposed: [opendev/system-config] 1009604: Trigger node exporter deployments when groups are modified https://review.opendev.org/c/opendev/system-config/+/100960415:45
@clarkb:matrix.orginfra-root ^ if anyone has time for that it should be a straightforward review. I'd like to get that in first then approve the change ot add codesearch to the node exporter group15:46
@jim:acmegating.comfungi: i may have left some extra debugging/profiling enabled which may have slowed things a bit... i'm correcting that, and will check to see if reconfigs are faster.15:49
@fungicide:matrix.orgcorvus: aha, thanks!15:49
@jim:acmegating.comyeah that's looking a lot better, sorry about that.  now we're back to where i thought we were: reconfigs are pretty fast, and db time queries are a little slow.15:54
@fungicide:matrix.orgno worries, glad it turned out to be something simple and not a performance regression in the code16:02
@clarkb:matrix.orgnow we are just starved for nodes16:39
@jim:acmegating.comi'm working on some query optimizations; i think we may actually be able to make an improvement here17:31
-@gerrit:opendev.org- Julia Kreger proposed: [openstack/diskimage-builder] 1009353: epel: partial revert of Ic6f6c92b28b80c4f0b93e4663129e73134f5b447 https://review.opendev.org/c/openstack/diskimage-builder/+/100935317:33
@fungicide:matrix.orgi don't think it's purely node starvation, i was pinged about changes sitting in the openstack tenant check pipeline for long periods of time with unprocessed result events for their builds, and https://zuul.opendev.org/t/openstack/status/pipeline/check is still showing the results queue count spiking up (though i think not as bad as it was earlier)17:35
@fungicide:matrix.orgthe backlog on the pending node requests graph was comparatively tame17:36
@jim:acmegating.comit's batchy due to the slow queue processing17:37
@jim:acmegating.combut in general, we have had more jobs than nodes for the last 4 hours or so17:39
@fungicide:matrix.orgyeah, and the executors are all hitting backoffs pretty frequently at the moment too17:40
@fungicide:matrix.orglooks like mainly due to the starting builds governor17:41
-@gerrit:opendev.org- Julia Kreger proposed: [openstack/diskimage-builder] 1009353: epel: partial revert of Ic6f6c92b28b80c4f0b93e4663129e73134f5b447 https://review.opendev.org/c/openstack/diskimage-builder/+/100935317:41
-@gerrit:opendev.org- Zuul merged on behalf of Clark Boylan: [opendev/system-config] 1009604: Trigger node exporter deployments when groups are modified https://review.opendev.org/c/opendev/system-config/+/100960417:43
@jim:acmegating.comokay, on the zookeeper cert front, the latest error is now: No subject alternative names matching IP address 213.32.78.235 found18:04
@jim:acmegating.comhttps://zuul.opendev.org/t/openstack/build/75c918b26c184083b713736ba5d4c615/log/zk99.opendev.org/docker/zookeeper-compose-zk-1.txt18:05
@jim:acmegating.com```18:06
213.32.78.235 zk98.opendev.org
213.32.78.235 zk98
213.32.79.194 zk99.opendev.org
213.32.79.194 zk99
```
@jim:acmegating.comthat's in /etc/hosts18:06
@jim:acmegating.com```18:06
server.98=213.32.78.235:2888:3888
server.99=213.32.79.194:2888:3888
```
@jim:acmegating.comthat's in zoo.cfg18:07
@jim:acmegating.comso i guess since we're telling it to connect by ip, it wants the ip in the cert18:07
@jim:acmegating.comi guess we could switch to hostnames?18:07
@clarkb:matrix.orgI think the reason we don't use hostnames is to avoid dns problems impacting the cluster?18:07
@jim:acmegating.comi suppose we only ever put the ip addrs in there for resili..yeah that18:08
@clarkb:matrix.orgWhether or not that is a good idea I don't know18:08
@jim:acmegating.comon its own, i like it.  but it's now adding complexity....18:08
@clarkb:matrix.orgmaybe we put IP addrs and names in the SAN list for completeness?18:08
@jim:acmegating.comwe can do that too, just need to be aware of that18:08
@jim:acmegating.comi don't think we ever have or plan to migrate a zk to a new ip without also moving to a new hostname, so as long as that holds, adding ips should be fine18:09
@clarkb:matrix.orghistorically I've done the zk replacements and always added new servers with new names so the cluster can stay up18:10
@clarkb:matrix.orgso right now we have zk01-03. To replace servers we would have 04, then 05 then 0618:10
@clarkb:matrix.orgso I think this would be safe18:10
@jim:acmegating.comhere's another hypothesis though: what if we don't need SAN if we switch to hostnames.  basically, maybe zookeeper just wants to validate whatever we put in the config file, and it was only looking for SAN because CN didn't match the ip18:11
@jim:acmegating.comif that hypo holds, then we could drop the san stuff and just switch zoo.cfg to dns18:11
@clarkb:matrix.orgah that would make sense18:13
@clarkb:matrix.orgrather than specifically wanting the SAN it wants to validate the cert against the config18:13
@jim:acmegating.comi'll try that real quick on these held nodes18:13
@jim:acmegating.comoh, but would we run into the ip address binding error on fip nodes in that case?18:17
@jim:acmegating.com(because i imagine we would put the fip in /etc/hosts)18:17
@clarkb:matrix.orgya I suspect that /etc/hosts has the public ips not the private ips in it so the name lookup would resolve to the fip then bind would fail?18:19
@fungicide:matrix.orgthat seems highly likely18:19
@clarkb:matrix.orgmordred: not sure if this si the sort of thing that interests you or if you have time to drop in but OPenStack will be talking about LLM contributions and code review related stuff at the PTG Monday: https://lists.openstack.org/archives/list/openstack-discuss@lists.openstack.org/thread/BULT6DYJQU7C7FGIK7HOMFSIFKU4FZB7/ I am unlikely to make that time slot myself18:34
@jim:acmegating.comi changed zoo.cfg to use a hostname and it decided to bind to 127.0.0.1 because that was the first match in /etc/hosts18:36
@jim:acmegating.comi guess we want to stick to ip addresses in zoo.cfg18:36
@fungicide:matrix.orghah, good point18:37
@jim:acmegating.comor we could set quorumListenOnAllIPs maybe?18:38
@jim:acmegating.comi'll try that18:38
@clarkb:matrix.orgwow thats a fun behavior18:38
@mordred:waterwanders.comClark: yeah, thanks! it's on my calendar, ildikov pointed me at it. I haven't gotten my notes together or organized yet :)18:39
@fungicide:matrix.orgi planned to be in there, should be interesting18:40
-@gerrit:opendev.org- Zuul merged on behalf of Clark Boylan: [opendev/system-config] 1009102: Add codesearch nodes to the node exporter group https://review.opendev.org/c/opendev/system-config/+/100910218:42
@clarkb:matrix.orghttps://grafana.opendev.org/d/49af61ae98/node-exporter-full?orgId=1&from=now-15m&to=now&timezone=utc&var-job=node&var-nodename=codesearch02&var-node=codesearch02.opendev.org&refresh=1m ok it sees /opt but seems to assign rootfs properties to it. Maybe because /opt is mouted under / which we then mount to /host/rootfs or whaetver so it flattens things out. But then it is able to look at something else to know /opt is a different fs?18:58
@clarkb:matrix.orgAnyway this is good data and confirms we likely haev some work to do to make other mount points report properly18:58
-@gerrit:opendev.org- James E. Blair https://matrix.to/#/@jim:acmegating.com proposed: [opendev/system-config] 1009387: Add SAN to certs generated by the opendev CA https://review.opendev.org/c/opendev/system-config/+/100938719:02
@fungicide:matrix.orgi wonder if we need to create monitoring-specific bindmounts like `/real/root` and `/real/opt` and `/real/var-lib-mailman` then collect stats on each individual item in `/real` without any nesting19:03
@fungicide:matrix.orgthough it does start to seem more and more like we're fighting against the chroot-oriented nature of containers for something that needs access to unfiltered host-level resources instead19:04
@clarkb:matrix.orgso apparently the rslave thing we have on the rootfs bind mount is there so that the mounts under / are preserved19:10
@clarkb:matrix.orgthat is what gives it some of this visibility. I'm wondering if we are just not exposing enough sysfs/procfs to lookup the details from there? Some of the examples show people bind mounting those in too19:10
@clarkb:matrix.orgI'll dig in a bit more after lunch19:13
@clarkb:matrix.orgIt seems like this is really close and the internet implies this should be workable without extra effort19:14
@clarkb:matrix.orgI wonder if this is podman specific behavior19:15
@clarkb:matrix.orgsince we are doing what all the examples do with docker but our backend is actually podman19:16
-@gerrit:opendev.org- Julia Kreger proposed: [openstack/diskimage-builder] 1009353: epel: partial revert of Ic6f6c92b28b80c4f0b93e4663129e73134f5b447 https://review.opendev.org/c/openstack/diskimage-builder/+/100935319:52
-@gerrit:opendev.org- Julia Kreger proposed: [openstack/diskimage-builder] 1009353: epel: partial revert of Ic6f6c92b28b80c4f0b93e4663129e73134f5b447 https://review.opendev.org/c/openstack/diskimage-builder/+/100935319:53
@clarkb:matrix.orgok we bind mount / to /host as ro,rslave. We also use pid:host. That means /proc in the container is the host proc which should give the container the info it needs to lookup disks and mounts and things. The rslave flag should mean (I thought anyway) that mounts like /opt are propogated to /host/opt. But ls -l /opt on the host side and in the container have different listings19:59
@clarkb:matrix.orgI suspect that the problem here is that the host pid is sufficient for node exporter to discover the mount but then when it looks at /host/opt statfs it returns the same data as /host because they are the same fs in the container despite the rslave flag20:00
@clarkb:matrix.organd now I'm reading the mount manpage. Podman docs indicates we may need to do something like mount --make-rshared /opt for this to work20:04
@clarkb:matrix.orgbut I'm not sure I fully understand the implications there. I'm guessing with docker this just works implicitly but podman doesn't do that?20:04
@fungicide:matrix.orginteresting, my mount(8) manpage mentions rshared but doesn't explain it20:07
@clarkb:matrix.orgya the documentation here is pretty basic and lacking. But basically some googling around and finding podman docs and issues seems to imply that they will not rslave/rshare unless the host side is already rshared/rslaved too20:08
@clarkb:matrix.orgmy hunch is that is what leads to /host/opt not being the /opt on the host side20:08
@clarkb:matrix.orgwhich then breaks statfs. One workaround does indeed seem to be manually mounting /opt to /host/opt (since that is where the mount tables will look for it we need to keep that mapping correct under the /host rootfs prefix20:09
@fungicide:matrix.orgaha, it's in kernel docs20:09
@clarkb:matrix.orgso maybe I should work on a change that takes an optional ansible var list of bind mounts for node-exporter and then we can set that in group vars for each group or something. Feels clunky but I think it should work20:09
@fungicide:matrix.orghttps://www.kernel.org/doc/Documentation/filesystems/sharedsubtree.txt20:10
@clarkb:matrix.orgthat also doesn't mention fstab so not sure if we can rely on that either20:11
@fungicide:matrix.org`--make-rshared` apparently indicates that the mount can be shared by other namespaces20:11
@clarkb:matrix.orgif we wanted to go with the more automated route we'd probably want fstab to take care of this for us20:11
@clarkb:matrix.orgin this case I think we want --make-rslave so that node-exporter sees changes made upstream of it but it can't make changes to the upstream20:11
@clarkb:matrix.orgbut that is the only difference with rshared aiui20:12
@clarkb:matrix.organd I guess with docker they take the presence of that flag to mean yes please fix the mounts for me and make it work20:13
@clarkb:matrix.orgat least that is my assumption based on the internet making it sound like this works reliably with docker compose which for most people means docker itself20:13
@clarkb:matrix.orgthen when mount is rewritten in rust we'll haev a whole new set of expectations and behaviors :P20:14
@fungicide:matrix.organd random crashes, apparently, if uutils is any indication20:15
@clarkb:matrix.orgI'll work on a patch for codesearch that tries the explicit bind mount route and see how ugly that turns out to be20:16
@clarkb:matrix.orgI think that is better and less magical etc20:16
@clarkb:matrix.organd if we hate it maybe we start considering running node-exporter on the host itself and not in a container20:16
@jim:acmegating.comor using snmp?20:17
@jim:acmegating.comhttps://github.com/prometheus/snmp_exporter20:17
@clarkb:matrix.orggood point there are other ways to get this sort of info out of linux machines that prometheus supports. I think there is some other system similar to node exporter too if we like http for some reason (maybe it handles this better)20:18
@jim:acmegating.comyeah.  snmp has the advantage of being a near universal standard for "run tiny thing on any host ever created and get info out of it in a consistent way".20:19
@jim:acmegating.comalso, near universal standard for "remote root exploit" but that's only if you use traps and we know not to do that.  i mean it's right there in the name.20:20
@jim:acmegating.comalso we have heard of firewalls20:20
@jim:acmegating.comanyway, i use snmp exporter at home for some stuff; it's been uneventful20:21
@jim:acmegating.comremote:   https://review.opendev.org/c/zuul/zuul/+/1009636 Add an optimized build time index [NEW]        20:23
remote: https://review.opendev.org/c/zuul/zuul/+/1009637 Reduce "group by" usage in sql queries [NEW]
i'm currently importing a copy of the opendev zuul database locally so that i can validate that these address the issues we're seeing. i went ahead and pushed them up in case anyone has early feedback while i wait for the import to finish
@jim:acmegating.com(in my initial testing, they made a huge improvement, but i want to repeat with a better test process and up-to-date data)20:24
@jim:acmegating.comalso, i'll check to see what impact a maridb upgrade has20:24
@clarkb:matrix.orgsounds good thanks20:27
-@gerrit:opendev.org- Clark Boylan proposed: [opendev/system-config] 1009639: Add explicit filesystem bind mounts to node exporter docker compose file https://review.opendev.org/c/opendev/system-config/+/100963920:36
@clarkb:matrix.orgIt actually doesn't seem too ugly other than it is another thing to account for20:37
@clarkb:matrix.orgbut we don't change mount paths often so maybe this is ok if it works?20:37
-@gerrit:opendev.org- Clark Boylan proposed on behalf of James E. Blair https://matrix.to/#/@jim:acmegating.com: [opendev/system-config] 1009387: Add SAN to certs generated by the opendev CA https://review.opendev.org/c/opendev/system-config/+/100938720:39
@clarkb:matrix.orgcorvus: ^ I went ahead and quickly fixed that linter failure since you're busy with sql stuff20:39
@clarkb:matrix.orgthough I wonder do we need to use the private IP in ci?20:40
@jim:acmegating.comoh thanks!  i have a lot of windows open right now.  what do you mean?20:40
@jim:acmegating.comoh20:40
@clarkb:matrix.orgcorvus: the update to the CA uses the public_v4 value here: https://review.opendev.org/c/opendev/system-config/+/1009387/6/playbooks/roles/opendev-ca/defaults/main.yaml but if that isn't a bindable address we're using the private ip20:41
@clarkb:matrix.orgI haven't looked at the other test failures yet but wonder if that is what went wrong20:41
@jim:acmegating.comyou mean do we need to do the ternary dance, yeah i think we do20:41
@clarkb:matrix.orgya that20:41
@jim:acmegating.comi can make that update20:42
@clarkb:matrix.orgk20:43
-@gerrit:opendev.org- James E. Blair https://matrix.to/#/@jim:acmegating.com proposed: [opendev/system-config] 1009387: Add SAN to certs generated by the opendev CA https://review.opendev.org/c/opendev/system-config/+/100938720:44
@clarkb:matrix.orglooks like they do publish binaries for node exporter that may be easy to install with a systemd unit and avoid a container all together: https://github.com/prometheus/node_exporter/releases/tag/v1.12.120:44
@clarkb:matrix.org(Just trying to look at options though the bind mount thing didn't end up being that terrible)20:45
@jim:acmegating.comthe sql changes are looking really good so far.  the build time query is instantaneous (0.010 seconds with no cache).  i did a build runtime graph of system-config-run-zookeeper for the past 2 years, and that took no time at all.  i did run-gitea for the last 2 years, and that took about a second with no cache (instant after that).  and that one has a lot of results.21:07
@jim:acmegating.comall of the build page queries i've done so far look good too.21:07
@fungicide:matrix.orgthat's amazing performance21:08
@clarkb:matrix.orgI guess I should review those changes now21:08
@clarkb:matrix.orgthen I'm going to try and pop out for a bike ride21:08
@clarkb:matrix.orgcorvus: job_name is defined as sa.Column(sa.String(SQL_MAX_STRING_LENGTH)) how can we haev an index on something that large (iirc the limits are quite small for index byte lengths)21:15
@clarkb:matrix.orgresult is also defined the same way. final and end_time are relatively small21:16
@jim:acmegating.comaiui, typical modern max is 3072 bytes.  with 4 byte encoding those two strings would be 2k.21:21
@clarkb:matrix.orgok 255 * 4 * 2 ~= 2k < 307221:24
@clarkb:matrix.orgthat would explain why it is working and not exploding for you. Do we think we need to worry about any zuul deployments with the old 768 byte limit? (we'd be well over that)21:24
@jim:acmegating.com255 with utf8 is already too much for the old style 768 byte limit; so if it's working for them, then 255+255<76821:30
@jim:acmegating.comso probably not?21:30
@clarkb:matrix.orgdatetime is 5-8 bytes and boolean is 1 byte21:30
@clarkb:matrix.orgah ok21:31
@clarkb:matrix.organyway I 2'd after following up to my comment with some of the math above21:31
@clarkb:matrix.orgI did not approve either change as I wasn't sure if you were still testing or if those are ready to go21:31
@jim:acmegating.comi'm working on testing pgsql locally to make sure there aren't any obvious performance regressions there21:31
@clarkb:matrix.orgI guess 3072 / 4 = 768 so they basically scaled up so the same number of characters would fit in an index using modern string encoding21:35
@clarkb:matrix.orgok I'm going to pop out now. Bcak in a bit21:45
-@gerrit:opendev.org- Jay Faulkner proposed: [openstack/diskimage-builder] 1009643: Move gentoo tests to experimental https://review.opendev.org/c/openstack/diskimage-builder/+/100964321:53
@jim:acmegating.comoh, the schema change took 4.5 minutes locally.22:01
-@gerrit:opendev.org- Zuul merged on behalf of Julia Kreger: [openstack/diskimage-builder] 1009353: epel: partial revert of Ic6f6c92b28b80c4f0b93e4663129e73134f5b447 https://review.opendev.org/c/openstack/diskimage-builder/+/100935322:32
-@gerrit:opendev.org- James E. Blair https://matrix.to/#/@jim:acmegating.com proposed: [opendev/system-config] 1009387: Add SAN to certs generated by the opendev CA https://review.opendev.org/c/opendev/system-config/+/100938722:37
@clarkb:matrix.orgcorvus: that seems not great but within the realm of reasonable for the benefit?23:33
@jim:acmegating.comClark: we have 31 million build records, that seems just dandy to me, and definitely worth the benefit.  :)23:34
@jim:acmegating.comalso, we have 31 million build records, nice :)23:37
@jim:acmegating.com5.2 million buildsets23:38
@jim:acmegating.com60 million build artifacts23:38
@jim:acmegating.comthe postgres data load finally finished and performance checks out there too.  i'll clean up the changes so they're ready to merge.23:40
@jim:acmegating.comClark: mordred https://review.opendev.org/1009636 could use a re-review; i fixed the test and a lint23:50
@mordred:waterwanders.comcorvus: doh23:51
@jim:acmegating.comwith the local testing i've done, i'm now as confident as i can be with a change like that23:51
@jim:acmegating.comi think it would be fine to let it do the schema upgrade with this weekends restarts23:52
@clarkb:matrix.orgyes we probably will take longer in production as I'm guessing our db isn't as zoomy as your test setup but still 15 minutes maybe? that should be fine23:57
@clarkb:matrix.orgcorvus: I approved it given your statement just above that you're as confident as you can be and letting that run in the weekend restarts is fine23:58

Generated by irclog2html.py 4.1.0 by Marius Gedminas - find it at https://mg.pov.lt/irclog2html/!