Wednesday, 2026-10-07

@harbott.osism.tech:regio.chatinfra-root: something seems wrong with zuul, I see a lot of periodic jobs still running which is unusual at this time of day. also some recheck triggers don't seem to work, e.g. on https://review.opendev.org/c/openstack/tempest/+/995934 and I couldn't find any related event in scheduler logs07:39
@harbott.osism.tech:regio.chatlots of graphs on the zuul status dashboard have flatlined since about 02:47 https://grafana.opendev.org/d/21a6e53ea4/zuul-status?orgId=1&from=now-6h&to=now&timezone=utc07:43
@noonedeadpunk:matrix.orghey folks. There seems to be over an hour queue to the zuul reading from gerrit?07:45
@noonedeadpunk:matrix.orgAs recheck made for https://review.opendev.org/c/openstack/openstack-ansible/+/1009082 never appeared in zuul after an hour07:45
@harbott.osism.tech:regio.chatyes, see my above comments, something is broken, not sure what yet07:46
@harbott.osism.tech:regio.chatoh, seems zookeeper is down07:47
@noonedeadpunk:matrix.orgI somehow see only messages from yesterday.... Will switch the client I guess07:47
@harbott.osism.tech:regio.chatso zookeeper containers seem to all have been rebuilt 5h ago. and zuul cannot connect to it since then07:49
```
08718085f4af zookeeper:3.9 "zkServer.sh start-f…" 5 hours ago Up 5 hours zookeeper-compose-zk-1
```
@noonedeadpunk:matrix.orgCould there be some issue with tls?07:51
@noonedeadpunk:matrix.orgunlikely though...07:52
@harbott.osism.tech:regio.chatwell I do see some tls related messages in the log, do you have some more information or was that just a wild guess?07:56
@harbott.osism.tech:regio.chatsadly our upgrade tooling involves immediately pruning the old container and it also seems to be no longer available on dockerhub07:58
-@status:opendev.org- NOTICE: zuul processing is broken since about 03:00 UTC, investigation is in progress, please be patient08:04
@noonedeadpunk:matrix.orgwell, it's not that I had something specific, there were just issues I saw with zookeeper before related to tls, including regressions in newer versions08:07
@noonedeadpunk:matrix.orgyou use upstream docker images for zookeeper?08:12
@noonedeadpunk:matrix.orgJust checking that you don't use https://opendev.org/openstack/ansible-role-zookeeper for instance to build them... As we've landed a change yesterday related to TLS08:13
@mnasiadka:matrix.orgWe're using the ones from docker hub08:13
@harbott.osism.tech:regio.chatyes, upstream zookeeper:3.908:14
@harbott.osism.tech:regio.chatthis is what I see in the logs since the restart (within some larger traceback):08:14
```
2026-10-07T02:47:58.024982+00:00 zk03 zookeeper-compose-zk-1[218792]: Caused by: java.security.cert.CertificateException: No subject alternative names present
```
@mnasiadka:matrix.orgOk, I disabled hostname verification in quorum ssl/mtls and it seems it helped08:21
@mnasiadka:matrix.orglet me raise a patch for this08:21
@mnasiadka:matrix.organd we can think of regenerating the certs later on - seems those don't include hostname properly08:21
-@gerrit:opendev.org- Michal Nasiadka proposed: [opendev/system-config] 1009160: zookeeper: Disable quorum TLS hostname verification https://review.opendev.org/c/opendev/system-config/+/100916008:26
@mnasiadka:matrix.orgJens Harbott: ^^08:30
@noonedeadpunk:matrix.orgI would guess it would need infra-root +Verified?:)08:33
@harbott.osism.tech:regio.chatit looks like zuul is slowly recovering, since the fix is manually applied, I'd say let's give it a bit of time08:34
@harbott.osism.tech:regio.chatalso periodic reminder that I've muted this channel, please ping me on IRC for urgent issues08:35
@harbott.osism.tech:regio.chatseems things are still slow with the backlog from tonight. so I think we should wait some more before we send the "all good, go ahead and recheck"09:49
@harbott.osism.tech:regio.chatI've added zk01-03 to the emergency list for now, just to be sure10:10
@harbott.osism.tech:regio.chatpushing monster stacks seems to be the new normal, now it is swifts turn13:15
@harbott.osism.tech:regio.chatbut otherwise the CI looks mostly fine to me. not sure how to word an "all clear" notice, though. as mentioned elsewhere it might be good to also include a word of caution regarding the new pbr release. suggestions welcome13:17
@fungicide:matrix.orgthanks mnasiadka and Jens Harbott for getting zk back on track! i'm actually a little surprised we auto-upgrade zk containers outside our normal zuul updating schedule13:41
@jim:acmegating.comi'm looking into a more permanent fix13:43
@noonedeadpunk:matrix.orgI wonder if still jobs need to be dropped... As I see bunch of jobs looking like stuck in queued state?14:02
@noonedeadpunk:matrix.orgAs this one https://review.opendev.org/c/openstack/openstack-ansible/+/1003930 was r4echeck after ZK was back14:03
@noonedeadpunk:matrix.orgthere're for sure more extreme cases for 30+ hours which was before zk died I guess14:04
@jim:acmegating.comlooks like that, and several other items, are waiting on nodes from rax-flex-dfw314:05
@clarkb:matrix.orgfungi: we pin to the stable version so we should in theory only get bugfix updates. Upgrading to the next stable release is something we trigger manually14:14
@jim:acmegating.comthe last few stable point releases have had netty/tls related security fixes; i'm guessing this was related14:15
@fungicide:matrix.orggot it, and we probably need to start including the individual server names in the cert's san field14:16
@jim:acmegating.comyeah, i'm working on a set of patches for that, but paused to look into the rax-flex-dfw3 issue14:17
@jim:acmegating.comhttps://grafana.opendev.org/d/0172e0bb72/zuul-launcher3a-rackspace-flex?orgId=1&from=now-2d&to=now&timezone=utc&var-region=raxflex-dfw314:17
@jim:acmegating.comsomething happened there a while ago, not related to zk14:18
@clarkb:matrix.orgDan With: may be interested in whatever we find14:18
@clarkb:matrix.orgfungi: did you see my notes about the Debian source package cleanup? I did have to run clearvanished and deleteunreferenced but I think the actual number of bytes freed was minimal. So we are on to the next cleanup in afs unfortunately 14:20
@harbott.osism.tech:regio.chatcorvus: looking at grafana: could this be a large rush of multinode requests which all have been partially filled, but are blocking all available quota?14:22
@jim:acmegating.comit's possible, i haven't been able to prove it yet14:23
@clarkb:matrix.orgI have some held nodes for Gerrit testing and maybe for node exporter debugging. They can be dropped if they are in dfw3 and we think that might help the scheduling14:24
@harbott.osism.tech:regio.chatthere are some kolla stacks in that backlog that we also could (step by step) dequeue to see if that helps, but I don't want to disturb the investigation. like e.g. starting at https://review.opendev.org/c/openstack/kolla-ansible/+/967801/4714:26
@fungicide:matrix.orgClark: yeah, comparing to e.g. http://deb.debian.org/debian/pool/main/o/openssh/ i don't see any tar or dsc files in our copy14:31
@jim:acmegating.comyes, it does look like almost every ready node in dfw3 is part of an incomplete multi-node request14:32
@jim:acmegating.comit looks like we have one node in-use; i want to know why it's not on the oldest request14:34
@jim:acmegating.comwe're hitting floating ip quota exceeded in flex iad314:39
@fungicide:matrix.orgso almost certainly a leak of some sort14:39
@fungicide:matrix.orgi need to step away to run some errands but should be back within a couple of hours and can help with fip cleanup at that point if needed14:40
@jim:acmegating.comthe one node in use on dfw3 is 8gb, all the incomplete requests are for 3x 16gb nodes14:52
@jim:acmegating.comkind of looks like our resource consumption escalated very quickly once 16gb nodes became available14:53
@jim:acmegating.comClark: there are 3 held 8gb nodes in dfw; i think if we release them, there may be enough quota to start to slowly work through the backlog, and that process should ramp up as each request is returned14:54
@harbott.osism.tech:regio.chatok, I'll dequeue some of the kolla things, too. those won't get reviewed soon anyway15:00
@jim:acmegating.comsounds good.  i think our culprit here is:15:02
1) lots of changes uploaded simultaneously (zuul quota view is not quite real-time)
2) those changes run lots of multi-node jobs (opportunity for partial fulfillment)
3) those nodes are large (cuts our quota in half compared to "normal" node size)
-@gerrit:opendev.org- James E. Blair https://matrix.to/#/@jim:acmegating.com proposed: [opendev/system-config] 1009279: WIP: Test 2-node zk cluster https://review.opendev.org/c/opendev/system-config/+/100927915:06
@harbott.osism.tech:regio.chatthere's also quite a number of recent builds getting RETRY results and no logs available. not sure whether or how that might be related to the other issues https://zuul.opendev.org/t/openstack/builds?result=RETRY&skip=015:07
@harbott.osism.tech:regio.chatlike https://zuul.opendev.org/t/openstack/build/de76ba4d0a48443dbe1c5136252c4377 for a concrete example15:08
@jim:acmegating.cominfra-root: ^ i'd like to see if we can have our tests show the zk error we're having, then fix it by adding a SAN.  it would be helpful to avoid merging mnasiadka 's fix for now while we see if this approach works15:08
@jim:acmegating.comthat would be due to the zookeeper connection loss15:15
@jim:acmegating.com2026-10-07 14:52:14,316 ERROR zuul.zk.ZooKeeper:   kazoo.exceptions.ConnectionLoss15:15
@jim:acmegating.comlooks like that was a transient error from only ze0415:21
@jim:acmegating.comi don't see any obvious cause15:23
@harbott.osism.tech:regio.chatthis seems to match. so either a network glitch or the executor was busy somehow?15:23
```
2026-10-07T14:52:09.714356+00:00 zk02 zookeeper-compose-zk-1[213907]: 2026-10-07 14:52:09,713 [myid:] - INFO [SessionTracker:o.a.z.s.ZooKeeperServer@730] - Expiring session 0x3068d34266e000e, timeout of 40000ms exceeded
```
@jim:acmegating.comyep15:24
@harbott.osism.tech:regio.chatok, I've dequeued all changes > 24h old and the next ones seem all to be proceeding now15:26
@harbott.osism.tech:regio.chatdoes anyone still want to send a status notice? I think it would be fine for people to recheck where needed by now?15:28
@harbott.osism.tech:regio.chatjust when you think that should be all for today, there seems also to be some issue with github mirror tasks, Internal Server Error on their end, so not much we can do about it https://zuul.opendev.org/t/openstack/build/e3c85e15f15c4fd0874aa56f84d956cd15:35
@clarkb:matrix.orgJens Harbott: corvus: to catch up did you end up deleting any held nodes or was clearing out the old queued builds for kolla sufficient? Also would it help if I looked into floating ip leak possibilities now?16:28
@jim:acmegating.comleak is low-priority, it's just a nuisance error in iad3 right now (but could get worse?)16:30
@jim:acmegating.comlooks like everything is cleaned out after the dequeue; i don't think any holds were deleted16:31
@clarkb:matrix.orgits usually pretty easy to check that. Do a listing then another listing in 15 minutes and any unattached IPs that exist in both listings can be removed. I can work on that shortly16:31
@fungicide:matrix.orgnote i've also still got https://review.opendev.org/c/opendev/zuul-providers/+/1007746 (Revert "Disable raxflex sjc3") waiting to increase our capacity somewhat16:38
@clarkb:matrix.org+2 from me on that one. We can always roll it back if that region has problems again16:39
-@gerrit:opendev.org- Zuul merged on behalf of Jeremy Stanley https://matrix.to/#/@fungicide:matrix.org: [opendev/zuul-providers] 1007746: Revert "Disable raxflex sjc3" https://review.opendev.org/c/opendev/zuul-providers/+/100774616:41
@clarkb:matrix.orgcorvus: we are using ~31 FIPs and they all have a fixed IP address assigned. This seems to roughly match a count of ~32 instances at the moment. Our FIP limit shows as -1 too. So I think the problem there may be that the cloud ran out of IPs?16:46
@clarkb:matrix.orgin any case I don't see any obvious signs of leaks16:46
@clarkb:matrix.orgthis was in raxflex iad3 to be specific16:46
@jim:acmegating.comoh so probably we're just not observing fip quota, i think that's not implemented yet; have we limited instances to 31 to match?  should we?16:47
@clarkb:matrix.orgwell FIP quota is at -1 which i think means unlimited?16:48
@jim:acmegating.comoh, i wonder if we ever got more than 31?16:48
@fungicide:matrix.orgthat's something Dan With may be able to double-check/adjust16:49
@clarkb:matrix.orgso we probably need someone like Dan With to weigh in on whether or not we need to adjust either the quota or our max server limit to make that region happy. It could just be a temporary contention for resources and things may auto resolve as the resource usage shifts16:49
@jim:acmegating.com++16:49
@fungicide:matrix.orgalso just as a reminder, we wouldn't need as much ipv4 fip quota if ipv6 is finally working in flex16:50
@fungicide:matrix.org(less a reminder to us, more a reminder to rackspace folks possibly seeing this)16:50
-@gerrit:opendev.org- James E. Blair https://matrix.to/#/@jim:acmegating.com proposed: [opendev/system-config] 1009279: WIP: Test 2-node zk cluster https://review.opendev.org/c/opendev/system-config/+/100927916:56
@jim:acmegating.comhttps://zuul.opendev.org/t/openstack/build/f558111cef31420a9fc3e2b5c2ca65aa/log/zk99.opendev.org/docker/zookeeper-compose-zk-1.txt#11817:30
@jim:acmegating.comit looks like we're not getting to the point of the expected error because the test zk servers can't listen on their assigned addresses; that's surprising to me.  any ideas why?17:31
@fungicide:matrix.orgthis is running in the context of a container we're not building ourselves, right?17:32
@jim:acmegating.comright17:32
@fungicide:matrix.orgin fact, jvm in a container17:33
-@gerrit:opendev.org- Clark Boylan proposed: [openstack/project-config] 1009325: Undefine zuul_site_mirror_fqdn on non x86_64 hosts https://review.opendev.org/c/openstack/project-config/+/100932517:33
-@gerrit:opendev.org- Clark Boylan proposed: [opendev/zuul-providers] 1009326: Use upstream package mirrors in arm64 image builds https://review.opendev.org/c/opendev/zuul-providers/+/100932617:33
@clarkb:matrix.orgcorvus: are they floating IPs so the host doesn't know about that address?17:33
@fungicide:matrix.orgthe port isn't in the traditional privileged range, but it's possible it needs some special permission to bind to any listening port instead of, like, a unix socket fd?17:34
@fungicide:matrix.orgoh...17:34
@fungicide:matrix.org`cloud: raxflex`17:34
@fungicide:matrix.orggood call17:34
@clarkb:matrix.orginfra-root I am not confident in 1009325 and it may cause every singe job we run to fail... But I think something like that as well as 1009326 is the next step in trying to drop arm64 pacakges if we go down that path17:34
@jim:acmegating.comyep i think that's it, i don't see any real ips there17:34
@jim:acmegating.comdo we have any other hosts with this problem?17:35
@fungicide:matrix.orgthings are probably still not quite right with how zuul conveys addresses in flex, because the inventory does state `interface_ip: 146.20.60.58` which isn't the actual address bound to the interface17:36
@clarkb:matrix.orgcorvus: no raxflex is the only cloud provider we currently use that has floating IPs17:36
@jim:acmegating.comfungi: well, that's the address we use to connect to it17:36
@clarkb:matrix.orgright from zuuls persepctive that is the IP17:36
@jim:acmegating.comwe can't connect to a node from the executor on a private ip17:36
@clarkb:matrix.orgif you look at the ansible facts you'll get the private "real" ips17:36
@jim:acmegating.comi'm wondering if we've solved this for any system-config-run jobs17:37
@fungicide:matrix.orgmaybe it's a terminology thing. it does have `public_ipv4: 146.20.60.58` which is what i would expect outside connections to rely on, to me `interface_ip` would be the address bound locally on the interface17:37
@jim:acmegating.combasically, is there any other service we run ansible playbook tests for that also needs to bind to a specific ip address?17:37
@clarkb:matrix.orgcorvus: I want to say we have handling for  that in the gitea haproxy rules17:37
@clarkb:matrix.org`address: "{{ (hostvars['gitea99.opendev.org'] | default({})).get('nodepool', {}).get('public_ipv4', '') }}:3080"` hrm maybe we're pointing at the floating ips there17:38
@jim:acmegating.comyeah, that's more or less what we're doing for zk17:39
@jim:acmegating.com`server.{{ host | regex_replace('^zk(\\d+)\\.open.*\\.org$', '\\1') | int }}={{ (hostvars[host].public_v4) }}:2888:3888`17:39
@fungicide:matrix.orgi take it zk/java needs an explicit binding address and can't just use `0.0.0.0` or `::`17:40
@clarkb:matrix.orgI think it is more zookeeper in this instance?17:40
@clarkb:matrix.orgya I'm not finding any evidence of solving this for something else in the system-config/playbooks/zuul dir17:41
@jim:acmegating.comit's a little surprising that it's using that as the binding address17:42
@fungicide:matrix.orgi wonder if it's relying on something like `gethostbyname()`17:42
@fungicide:matrix.orgin which case maybe overriding entries in `/etc/hosts` would suffice17:43
@jim:acmegating.comi'm not following; that's going to just have an ip in the config file17:44
@clarkb:matrix.orgno its the literal ip addresses to set up the cluster members17:44
@clarkb:matrix.orgGoogle seems to think that it should bind to 0.0.0.0 by default but I think it must override that when we set the cluster membership like this17:44
@clarkb:matrix.orghttps://oneuptime.com/blog/post/2026-03-20-zookeeper-bind-ipv4-kafka/view#binding-quorum-ports-to-specific-ips maybe this is 3.7+ "new" behavior?17:45
@jim:acmegating.commaybe fungi is suggesting to change to a hostname?  that could be a solution... we'd probably have to answer a lot of questions to confirm the right behavior in both testing and prod.17:45
@fungicide:matrix.orgi mean whatever is deciding to tell the ListenerHandler to explicitly bind on `146.20.63.157:3888` may be doing a lookup by name, but yeah probably not since it's not going to be in dns for a test node and odds are we don't inject the global address into `/etc/hosts` ourselves anyway17:46
@clarkb:matrix.orgit can't find on that IP if the IP isn't configured on the host though?17:46
@clarkb:matrix.orgI'm still not sure I follow the name lookup path17:46
@clarkb:matrix.org* it can't bind on that IP if the IP isn't configured on the host though?17:47
@fungicide:matrix.orgi doubt the test node itself has any awareness of its relationship with 146.20.63.157 so presumably something somewhere in ansible is plumbing that through to configuration17:47
@clarkb:matrix.orgfungi: yes our inventory17:47
@jim:acmegating.comit's in the zookeeper config17:47
@fungicide:matrix.orgi mean something in the ansible playbooks/roles is telling it to use that inventory value17:48
@jim:acmegating.comprobably the line i pasted17:48
@fungicide:matrix.orgrather than a literal `0.0.0.0` or `::` or empty string17:48
@clarkb:matrix.orgyup its this https://opendev.org/opendev/system-config/src/branch/master/playbooks/roles/zookeeper/templates/zoo.cfg.j2#L3117:48
@clarkb:matrix.orgthe reason is that you ahve to set up the cluster this way. It is how zookeeper knows who its peers are17:48
@fungicide:matrix.orgso my earlier question was can we just not specify an address (if it will bind to all addresses by default), or specify a wildcard address?17:49
@clarkb:matrix.orgI don't think so because then you don't have a cluster17:49
@fungicide:matrix.orgoh, so it's both the address the process binds to *and* the address other cluster members expect to reach it at? those aren't separate configuration values?17:49
@jim:acmegating.comit's certainly behaving that way17:50
@jim:acmegating.comthe docs are decidedly unclear on this17:50
@clarkb:matrix.orgthe blog post I linked above implies this is how it works too17:50
@fungicide:matrix.orgso literally can't be made to work through address translation at all, if so17:51
@jim:acmegating.comoh wait, better docs here: https://zookeeper.apache.org/doc/r3.9.6/zookeeperAdmin.html17:51
@jim:acmegating.comokay that new multi address thing is scary17:52
@jim:acmegating.comand i'm not sure that would actually solve it for us, if one of ips  listed is the fip and it still can't bind to it17:54
@clarkb:matrix.orgI have some hacky ideas for working around this in ansible: essentially update run-base.yaml to set private_v4 next to public_v4 if present. Then in the zoo.cfg.j2 file use private_v4 if present else use public_v4. that should mean prod doesn't change and allow us to use the private known address on the host in CI. I don't know if that means we also need to update zuul config to prefer private over public if set when talking to zk17:54
@clarkb:matrix.orghttps://opendev.org/opendev/system-config/src/branch/master/playbooks/zuul/run-base.yaml#L81-L83 this is where run-base.yaml would be updated in that scenario17:55
@jim:acmegating.comClark: yeah, i think we may need something like that.  i think we may only use one zk host with the zuul tests though, so maybe we don't need to worry about that?17:56
@clarkb:matrix.orgcorvus: if those jobs share the same playbook though it will bind on the private ip and zuul will try to connect to it via public ip if we don't also update zuul's config?17:56
@clarkb:matrix.orgits possible that it will work though17:56
@jim:acmegating.comClark: lol it would have worked 4 monhs ago:  https://opendev.org/opendev/system-config/commit/6d1490b20e3f443149eae3f6c72bf26cc323f13717:56
@clarkb:matrix.orgoh yes this is the thing we wanted to change to make it less confusing :)17:57
@clarkb:matrix.orgwe traded one version of confusion for another :/17:57
@clarkb:matrix.orgthis is still probably the better outcome because its less magic. We'd have to explicitly use the private addr where we know we might need it18:00
@clarkb:matrix.orgbut still18:00
@fungicide:matrix.orgthe previous treat-the-private-address-as-public hack was also breaking things in the other direction, to be fair18:03
@jim:acmegating.comit could be argued that the old setup let us have production playbooks that worked without alteration in test environments18:03
@jim:acmegating.comwhat was it breaking?18:03
@jim:acmegating.com(basically, we reframed our test environment to make it look more like production, but now we will need to update our production playbooks to understand whether they are running in test or not)18:04
@clarkb:matrix.orgcorvus: it broke mixed cloud nodesets since the private IPs couldn't be routed between cloud regions18:04
@clarkb:matrix.orgwhich we're using to test arm64 stuff with an x86 bridge to better match reality18:05
@clarkb:matrix.orgthere was also some issue with having an arm64 test bridge I don't rmember the details of18:05
@fungicide:matrix.org(once those became a thing that zuul supported, and we occasionally got them as a fallback after launch failures)18:05
@clarkb:matrix.orgI think there was a package ansible needed without a wheel maybe18:05
@fungicide:matrix.orgoh, right, the mixed-arch testing was the more consistent problem18:05
@fungicide:matrix.orgyes, i don't recall the details now, but there was some ubuntu upgrade problem which led us to want to switch to mixed-node18:07
@fungicide:matrix.orgpretty sure it was related to migrating bridge to newer ubuntu18:07
@jim:acmegating.comokay, so we need to:18:07
1) add private_ipv4 to inventory if it exists
2) use private_ipv4 when writing zoo.cfg if it exists and is not empty, otherwise public_ipv4
3) do the same in zuul.conf
that about it?
@clarkb:matrix.orgcorvus: I think that should do it18:07
@fungicide:matrix.orgthat does seem like it would solve the immediate error18:08
@clarkb:matrix.orgas a side note it really is weird that knowing who your cluster members are and where you bind your socket are not separate concerns18:09
@clarkb:matrix.orgcorvus: supposedly `clientPortAddress=0.0.0.0` may cause it to bind on 0.0.0.0?18:10
@clarkb:matrix.orgmaybe it is worth testing that?18:10
@jim:acmegating.combut this is the quorum binding18:10
@clarkb:matrix.orgoh right its a different port18:10
@clarkb:matrix.orgnevermind I think this won't solve it18:10
@clarkb:matrix.orgbut also maybe that maens zuul doesn't need to know?18:10
@clarkb:matrix.orgif the client bind is 0.0.0.0 then zuul should be able to connct to the floating ip properly18:11
@jim:acmegating.comi'm with you... i still have a terminal open grepping through zookeeper to try to find where it's doing that18:11
@clarkb:matrix.orgthat would explain why this worked until you added a second server18:13
@jim:acmegating.comClark: oh yeah i think the default for clientPortAddress is any, so i think we don't need to worry about zuul already18:13
@clarkb:matrix.orgmaybe18:13
@clarkb:matrix.org++18:13
@jim:acmegating.comi should have that change as soon as i spend the next 15 minutes testing ansible udefined vars18:15
@jim:acmegating.comif we ever add "private_ipv4" to our public inventory, we're going to immediately break this since it'll be backwards18:18
@clarkb:matrix.orgcorvus: should it be private_v4 to match public_v4?18:19
@jim:acmegating.comsure whatever it's called :)18:19
@clarkb:matrix.orgbut yes good call. Maybe we even add a comment to the inventory entries for zookeeper nodes about it?18:19
@jim:acmegating.commaybe we should call it "silly_fake_private_v4_for_use_in_testing"18:19
@clarkb:matrix.orgI'm fine with something along those lines18:19
@clarkb:matrix.orgci_specific_private_v4 might be less silly18:20
@clarkb:matrix.orgbut I'm good with silly too18:20
@fungicide:matrix.org`slightly_less_silly_ci_specific_private_v4` :P18:21
@fungicide:matrix.orgi approve on behalf of the ministry of silly variable names18:21
-@gerrit:opendev.org- James E. Blair https://matrix.to/#/@jim:acmegating.com proposed:19:18
- [opendev/system-config] 1009279: WIP: Test 2-node zk cluster https://review.opendev.org/c/opendev/system-config/+/1009279
- [opendev/system-config] 1009351: Use private ipv4 for zookeeper quorum configuration in testing https://review.opendev.org/c/opendev/system-config/+/1009351
@clarkb:matrix.orgcorvus: the ternary thing should work fine but would private_ipv4 | default (public_ipv4) also work? Or maybe there is something subtle I'm missing19:45
@clarkb:matrix.orgMaybe default only works for unset values rather than untruthy values?19:45
@fungicide:matrix.orgi would expect it to only work for undefined values19:46
@fungicide:matrix.orgbut i could be expecting wrong19:46
@fungicide:matrix.orgi've always assumed it's more like `.get()` methods in python19:47
@fungicide:matrix.orgrather than if/then/else19:47
@jim:acmegating.comthat's the assumption i wrote it with; i want to handle private_ipv4=null19:47
@clarkb:matrix.orgGot it19:49
-@gerrit:opendev.org- Julia Kreger proposed: [openstack/diskimage-builder] 1009353: epel: exclude yum-utils from auto-installing on non-rh style distros https://review.opendev.org/c/openstack/diskimage-builder/+/100935320:01
@jim:acmegating.comthere is a lot of red on that change, but i *think* it's all ubuntu archive network errors20:25
@clarkb:matrix.orgya the gitea change hit ubuntu archive network errors too20:25
@clarkb:matrix.orgwe setup our system-config jobs to not use our mirrors because our normal prod service hosts don't use our mirrors. But that makes them vulnerable to upstream mirror issues20:26
-@gerrit:opendev.org- Steve Baker proposed on behalf of Julia Kreger: [openstack/diskimage-builder] 1009353: epel: Use sed for disabling epel repo https://review.opendev.org/c/openstack/diskimage-builder/+/100935321:56
@jim:acmegating.comall right!  https://zuul.opendev.org/t/openstack/build/9a2c8593870c40ffa551c82e290db5e5 is failing as expected22:47
@jim:acmegating.comtestinfra is failing the test that the cluster is up, and the cluster is not up because of no SAN22:48
@clarkb:matrix.orgsuccess!ful failure22:48
@jim:acmegating.comso now we can build on that and fix the san thing22:48
@clarkb:matrix.organy idea if this is just a problem for the quorum members or if clients will hit it too yet?22:50
@jim:acmegating.com`ssl.clientHostnameVerification and ssl.quorum.clientHostnameVerification : (Java system properties: zookeeper.ssl.clientHostnameVerification and zookeeper.ssl.quorum.clientHostnameVerification) New in 3.9.4: Specifies whether the client's hostname verification is enabled in client and quorum TLS negotiation process. This option requires the corresponding hostnameVerification option to be true, or it will be ignored. Default: true for quorum, false for clients`22:54
@jim:acmegating.comhttps://zookeeper.apache.org/doc/r3.9.5/zookeeperAdmin.html22:54
@jim:acmegating.comthat looks like a "not a problem for clients" to me22:55
@clarkb:matrix.orgyup seems to be that way22:56
-@gerrit:opendev.org- James E. Blair https://matrix.to/#/@jim:acmegating.com proposed:22:59
- [opendev/system-config] 1009279: WIP: Test 2-node zk cluster https://review.opendev.org/c/opendev/system-config/+/1009279
- [opendev/system-config] 1009387: Add SAN to certs generated by the opendev CA https://review.opendev.org/c/opendev/system-config/+/1009387
@jim:acmegating.comif 387 works we can squash it into 279 for a mergeable change;23:00
-@gerrit:opendev.org- Zuul merged on behalf of Clark Boylan: [opendev/system-config] 1009126: Upgrade gitea to 28.1.0 https://review.opendev.org/c/opendev/system-config/+/100912623:01
@clarkb:matrix.orgthat got in ahead of the hourly jobs23:02
@clarkb:matrix.orgit is about halfway through the cluster now. When it is done I'll get general browseability and git clone and then also double check replication has succeeded23:09
-@gerrit:opendev.org- James E. Blair https://matrix.to/#/@jim:acmegating.com proposed: [opendev/system-config] 1009387: Add SAN to certs generated by the opendev CA https://review.opendev.org/c/opendev/system-config/+/100938723:14
@jim:acmegating.comapparently opendev-ca changes didn't trigger the zookeeper job23:14
@clarkb:matrix.orgcluster is upgraded as of 20 seconds ago23:15
@jim:acmegating.commaybe we should run zuul on those changes too23:15
@clarkb:matrix.org++23:15
-@gerrit:opendev.org- James E. Blair https://matrix.to/#/@jim:acmegating.com proposed: [opendev/system-config] 1009387: Add SAN to certs generated by the opendev CA https://review.opendev.org/c/opendev/system-config/+/100938723:16
@clarkb:matrix.orggit clone works. I can browse system config, look at files on HEAD, look at file commit history and view the diff of a file on a commit23:17
@clarkb:matrix.orgso it seems to generally work23:17
@clarkb:matrix.organd 1009387 replicated: https://opendev.org/opendev/system-config/commit/5ab12e2f7de0b493e268cebeea33b756c73cf15023:17
@clarkb:matrix.orgso gitea 28.1.0 is looking good to me23:17
@clarkb:matrix.orgcorvus: in theory our multinode test setup updates /etc/hosts so that each nodes knows about all of the others via name lookups. So I think this should work. But if things fail then actually validating the name:ip mappings is probably where I would look next23:19

Generated by irclog2html.py 4.1.0 by Marius Gedminas - find it at https://mg.pov.lt/irclog2html/!