| @harbott.osism.tech:regio.chat | infra-root: something seems wrong with zuul, I see a lot of periodic jobs still running which is unusual at this time of day. also some recheck triggers don't seem to work, e.g. on https://review.opendev.org/c/openstack/tempest/+/995934 and I couldn't find any related event in scheduler logs | 07:39 |
|---|---|---|
| @harbott.osism.tech:regio.chat | lots of graphs on the zuul status dashboard have flatlined since about 02:47 https://grafana.opendev.org/d/21a6e53ea4/zuul-status?orgId=1&from=now-6h&to=now&timezone=utc | 07:43 |
| @noonedeadpunk:matrix.org | hey folks. There seems to be over an hour queue to the zuul reading from gerrit? | 07:45 |
| @noonedeadpunk:matrix.org | As recheck made for https://review.opendev.org/c/openstack/openstack-ansible/+/1009082 never appeared in zuul after an hour | 07:45 |
| @harbott.osism.tech:regio.chat | yes, see my above comments, something is broken, not sure what yet | 07:46 |
| @harbott.osism.tech:regio.chat | oh, seems zookeeper is down | 07:47 |
| @noonedeadpunk:matrix.org | I somehow see only messages from yesterday.... Will switch the client I guess | 07:47 |
| @harbott.osism.tech:regio.chat | so zookeeper containers seem to all have been rebuilt 5h ago. and zuul cannot connect to it since then | 07:49 |
| ``` | ||
| 08718085f4af zookeeper:3.9 "zkServer.sh start-f…" 5 hours ago Up 5 hours zookeeper-compose-zk-1 | ||
| ``` | ||
| @noonedeadpunk:matrix.org | Could there be some issue with tls? | 07:51 |
| @noonedeadpunk:matrix.org | unlikely though... | 07:52 |
| @harbott.osism.tech:regio.chat | well I do see some tls related messages in the log, do you have some more information or was that just a wild guess? | 07:56 |
| @harbott.osism.tech:regio.chat | sadly our upgrade tooling involves immediately pruning the old container and it also seems to be no longer available on dockerhub | 07:58 |
| -@status:opendev.org- NOTICE: zuul processing is broken since about 03:00 UTC, investigation is in progress, please be patient | 08:04 | |
| @noonedeadpunk:matrix.org | well, it's not that I had something specific, there were just issues I saw with zookeeper before related to tls, including regressions in newer versions | 08:07 |
| @noonedeadpunk:matrix.org | you use upstream docker images for zookeeper? | 08:12 |
| @noonedeadpunk:matrix.org | Just checking that you don't use https://opendev.org/openstack/ansible-role-zookeeper for instance to build them... As we've landed a change yesterday related to TLS | 08:13 |
| @mnasiadka:matrix.org | We're using the ones from docker hub | 08:13 |
| @harbott.osism.tech:regio.chat | yes, upstream zookeeper:3.9 | 08:14 |
| @harbott.osism.tech:regio.chat | this is what I see in the logs since the restart (within some larger traceback): | 08:14 |
| ``` | ||
| 2026-10-07T02:47:58.024982+00:00 zk03 zookeeper-compose-zk-1[218792]: Caused by: java.security.cert.CertificateException: No subject alternative names present | ||
| ``` | ||
| @mnasiadka:matrix.org | Ok, I disabled hostname verification in quorum ssl/mtls and it seems it helped | 08:21 |
| @mnasiadka:matrix.org | let me raise a patch for this | 08:21 |
| @mnasiadka:matrix.org | and we can think of regenerating the certs later on - seems those don't include hostname properly | 08:21 |
| -@gerrit:opendev.org- Michal Nasiadka proposed: [opendev/system-config] 1009160: zookeeper: Disable quorum TLS hostname verification https://review.opendev.org/c/opendev/system-config/+/1009160 | 08:26 | |
| @mnasiadka:matrix.org | Jens Harbott: ^^ | 08:30 |
| @noonedeadpunk:matrix.org | I would guess it would need infra-root +Verified?:) | 08:33 |
| @harbott.osism.tech:regio.chat | it looks like zuul is slowly recovering, since the fix is manually applied, I'd say let's give it a bit of time | 08:34 |
| @harbott.osism.tech:regio.chat | also periodic reminder that I've muted this channel, please ping me on IRC for urgent issues | 08:35 |
| @harbott.osism.tech:regio.chat | seems things are still slow with the backlog from tonight. so I think we should wait some more before we send the "all good, go ahead and recheck" | 09:49 |
| @harbott.osism.tech:regio.chat | I've added zk01-03 to the emergency list for now, just to be sure | 10:10 |
| @harbott.osism.tech:regio.chat | pushing monster stacks seems to be the new normal, now it is swifts turn | 13:15 |
| @harbott.osism.tech:regio.chat | but otherwise the CI looks mostly fine to me. not sure how to word an "all clear" notice, though. as mentioned elsewhere it might be good to also include a word of caution regarding the new pbr release. suggestions welcome | 13:17 |
| @fungicide:matrix.org | thanks mnasiadka and Jens Harbott for getting zk back on track! i'm actually a little surprised we auto-upgrade zk containers outside our normal zuul updating schedule | 13:41 |
| @jim:acmegating.com | i'm looking into a more permanent fix | 13:43 |
| @noonedeadpunk:matrix.org | I wonder if still jobs need to be dropped... As I see bunch of jobs looking like stuck in queued state? | 14:02 |
| @noonedeadpunk:matrix.org | As this one https://review.opendev.org/c/openstack/openstack-ansible/+/1003930 was r4echeck after ZK was back | 14:03 |
| @noonedeadpunk:matrix.org | there're for sure more extreme cases for 30+ hours which was before zk died I guess | 14:04 |
| @jim:acmegating.com | looks like that, and several other items, are waiting on nodes from rax-flex-dfw3 | 14:05 |
| @clarkb:matrix.org | fungi: we pin to the stable version so we should in theory only get bugfix updates. Upgrading to the next stable release is something we trigger manually | 14:14 |
| @jim:acmegating.com | the last few stable point releases have had netty/tls related security fixes; i'm guessing this was related | 14:15 |
| @fungicide:matrix.org | got it, and we probably need to start including the individual server names in the cert's san field | 14:16 |
| @jim:acmegating.com | yeah, i'm working on a set of patches for that, but paused to look into the rax-flex-dfw3 issue | 14:17 |
| @jim:acmegating.com | https://grafana.opendev.org/d/0172e0bb72/zuul-launcher3a-rackspace-flex?orgId=1&from=now-2d&to=now&timezone=utc&var-region=raxflex-dfw3 | 14:17 |
| @jim:acmegating.com | something happened there a while ago, not related to zk | 14:18 |
| @clarkb:matrix.org | Dan With: may be interested in whatever we find | 14:18 |
| @clarkb:matrix.org | fungi: did you see my notes about the Debian source package cleanup? I did have to run clearvanished and deleteunreferenced but I think the actual number of bytes freed was minimal. So we are on to the next cleanup in afs unfortunately | 14:20 |
| @harbott.osism.tech:regio.chat | corvus: looking at grafana: could this be a large rush of multinode requests which all have been partially filled, but are blocking all available quota? | 14:22 |
| @jim:acmegating.com | it's possible, i haven't been able to prove it yet | 14:23 |
| @clarkb:matrix.org | I have some held nodes for Gerrit testing and maybe for node exporter debugging. They can be dropped if they are in dfw3 and we think that might help the scheduling | 14:24 |
| @harbott.osism.tech:regio.chat | there are some kolla stacks in that backlog that we also could (step by step) dequeue to see if that helps, but I don't want to disturb the investigation. like e.g. starting at https://review.opendev.org/c/openstack/kolla-ansible/+/967801/47 | 14:26 |
| @fungicide:matrix.org | Clark: yeah, comparing to e.g. http://deb.debian.org/debian/pool/main/o/openssh/ i don't see any tar or dsc files in our copy | 14:31 |
| @jim:acmegating.com | yes, it does look like almost every ready node in dfw3 is part of an incomplete multi-node request | 14:32 |
| @jim:acmegating.com | it looks like we have one node in-use; i want to know why it's not on the oldest request | 14:34 |
| @jim:acmegating.com | we're hitting floating ip quota exceeded in flex iad3 | 14:39 |
| @fungicide:matrix.org | so almost certainly a leak of some sort | 14:39 |
| @fungicide:matrix.org | i need to step away to run some errands but should be back within a couple of hours and can help with fip cleanup at that point if needed | 14:40 |
| @jim:acmegating.com | the one node in use on dfw3 is 8gb, all the incomplete requests are for 3x 16gb nodes | 14:52 |
| @jim:acmegating.com | kind of looks like our resource consumption escalated very quickly once 16gb nodes became available | 14:53 |
| @jim:acmegating.com | Clark: there are 3 held 8gb nodes in dfw; i think if we release them, there may be enough quota to start to slowly work through the backlog, and that process should ramp up as each request is returned | 14:54 |
| @harbott.osism.tech:regio.chat | ok, I'll dequeue some of the kolla things, too. those won't get reviewed soon anyway | 15:00 |
| @jim:acmegating.com | sounds good. i think our culprit here is: | 15:02 |
| 1) lots of changes uploaded simultaneously (zuul quota view is not quite real-time) | ||
| 2) those changes run lots of multi-node jobs (opportunity for partial fulfillment) | ||
| 3) those nodes are large (cuts our quota in half compared to "normal" node size) | ||
| -@gerrit:opendev.org- James E. Blair https://matrix.to/#/@jim:acmegating.com proposed: [opendev/system-config] 1009279: WIP: Test 2-node zk cluster https://review.opendev.org/c/opendev/system-config/+/1009279 | 15:06 | |
| @harbott.osism.tech:regio.chat | there's also quite a number of recent builds getting RETRY results and no logs available. not sure whether or how that might be related to the other issues https://zuul.opendev.org/t/openstack/builds?result=RETRY&skip=0 | 15:07 |
| @harbott.osism.tech:regio.chat | like https://zuul.opendev.org/t/openstack/build/de76ba4d0a48443dbe1c5136252c4377 for a concrete example | 15:08 |
| @jim:acmegating.com | infra-root: ^ i'd like to see if we can have our tests show the zk error we're having, then fix it by adding a SAN. it would be helpful to avoid merging mnasiadka 's fix for now while we see if this approach works | 15:08 |
| @jim:acmegating.com | that would be due to the zookeeper connection loss | 15:15 |
| @jim:acmegating.com | 2026-10-07 14:52:14,316 ERROR zuul.zk.ZooKeeper: kazoo.exceptions.ConnectionLoss | 15:15 |
| @jim:acmegating.com | looks like that was a transient error from only ze04 | 15:21 |
| @jim:acmegating.com | i don't see any obvious cause | 15:23 |
| @harbott.osism.tech:regio.chat | this seems to match. so either a network glitch or the executor was busy somehow? | 15:23 |
| ``` | ||
| 2026-10-07T14:52:09.714356+00:00 zk02 zookeeper-compose-zk-1[213907]: 2026-10-07 14:52:09,713 [myid:] - INFO [SessionTracker:o.a.z.s.ZooKeeperServer@730] - Expiring session 0x3068d34266e000e, timeout of 40000ms exceeded | ||
| ``` | ||
| @jim:acmegating.com | yep | 15:24 |
| @harbott.osism.tech:regio.chat | ok, I've dequeued all changes > 24h old and the next ones seem all to be proceeding now | 15:26 |
| @harbott.osism.tech:regio.chat | does anyone still want to send a status notice? I think it would be fine for people to recheck where needed by now? | 15:28 |
| @harbott.osism.tech:regio.chat | just when you think that should be all for today, there seems also to be some issue with github mirror tasks, Internal Server Error on their end, so not much we can do about it https://zuul.opendev.org/t/openstack/build/e3c85e15f15c4fd0874aa56f84d956cd | 15:35 |
| @clarkb:matrix.org | Jens Harbott: corvus: to catch up did you end up deleting any held nodes or was clearing out the old queued builds for kolla sufficient? Also would it help if I looked into floating ip leak possibilities now? | 16:28 |
| @jim:acmegating.com | leak is low-priority, it's just a nuisance error in iad3 right now (but could get worse?) | 16:30 |
| @jim:acmegating.com | looks like everything is cleaned out after the dequeue; i don't think any holds were deleted | 16:31 |
| @clarkb:matrix.org | its usually pretty easy to check that. Do a listing then another listing in 15 minutes and any unattached IPs that exist in both listings can be removed. I can work on that shortly | 16:31 |
| @fungicide:matrix.org | note i've also still got https://review.opendev.org/c/opendev/zuul-providers/+/1007746 (Revert "Disable raxflex sjc3") waiting to increase our capacity somewhat | 16:38 |
| @clarkb:matrix.org | +2 from me on that one. We can always roll it back if that region has problems again | 16:39 |
| -@gerrit:opendev.org- Zuul merged on behalf of Jeremy Stanley https://matrix.to/#/@fungicide:matrix.org: [opendev/zuul-providers] 1007746: Revert "Disable raxflex sjc3" https://review.opendev.org/c/opendev/zuul-providers/+/1007746 | 16:41 | |
| @clarkb:matrix.org | corvus: we are using ~31 FIPs and they all have a fixed IP address assigned. This seems to roughly match a count of ~32 instances at the moment. Our FIP limit shows as -1 too. So I think the problem there may be that the cloud ran out of IPs? | 16:46 |
| @clarkb:matrix.org | in any case I don't see any obvious signs of leaks | 16:46 |
| @clarkb:matrix.org | this was in raxflex iad3 to be specific | 16:46 |
| @jim:acmegating.com | oh so probably we're just not observing fip quota, i think that's not implemented yet; have we limited instances to 31 to match? should we? | 16:47 |
| @clarkb:matrix.org | well FIP quota is at -1 which i think means unlimited? | 16:48 |
| @jim:acmegating.com | oh, i wonder if we ever got more than 31? | 16:48 |
| @fungicide:matrix.org | that's something Dan With may be able to double-check/adjust | 16:49 |
| @clarkb:matrix.org | so we probably need someone like Dan With to weigh in on whether or not we need to adjust either the quota or our max server limit to make that region happy. It could just be a temporary contention for resources and things may auto resolve as the resource usage shifts | 16:49 |
| @jim:acmegating.com | ++ | 16:49 |
| @fungicide:matrix.org | also just as a reminder, we wouldn't need as much ipv4 fip quota if ipv6 is finally working in flex | 16:50 |
| @fungicide:matrix.org | (less a reminder to us, more a reminder to rackspace folks possibly seeing this) | 16:50 |
| -@gerrit:opendev.org- James E. Blair https://matrix.to/#/@jim:acmegating.com proposed: [opendev/system-config] 1009279: WIP: Test 2-node zk cluster https://review.opendev.org/c/opendev/system-config/+/1009279 | 16:56 | |
| @jim:acmegating.com | https://zuul.opendev.org/t/openstack/build/f558111cef31420a9fc3e2b5c2ca65aa/log/zk99.opendev.org/docker/zookeeper-compose-zk-1.txt#118 | 17:30 |
| @jim:acmegating.com | it looks like we're not getting to the point of the expected error because the test zk servers can't listen on their assigned addresses; that's surprising to me. any ideas why? | 17:31 |
| @fungicide:matrix.org | this is running in the context of a container we're not building ourselves, right? | 17:32 |
| @jim:acmegating.com | right | 17:32 |
| @fungicide:matrix.org | in fact, jvm in a container | 17:33 |
| -@gerrit:opendev.org- Clark Boylan proposed: [openstack/project-config] 1009325: Undefine zuul_site_mirror_fqdn on non x86_64 hosts https://review.opendev.org/c/openstack/project-config/+/1009325 | 17:33 | |
| -@gerrit:opendev.org- Clark Boylan proposed: [opendev/zuul-providers] 1009326: Use upstream package mirrors in arm64 image builds https://review.opendev.org/c/opendev/zuul-providers/+/1009326 | 17:33 | |
| @clarkb:matrix.org | corvus: are they floating IPs so the host doesn't know about that address? | 17:33 |
| @fungicide:matrix.org | the port isn't in the traditional privileged range, but it's possible it needs some special permission to bind to any listening port instead of, like, a unix socket fd? | 17:34 |
| @fungicide:matrix.org | oh... | 17:34 |
| @fungicide:matrix.org | `cloud: raxflex` | 17:34 |
| @fungicide:matrix.org | good call | 17:34 |
| @clarkb:matrix.org | infra-root I am not confident in 1009325 and it may cause every singe job we run to fail... But I think something like that as well as 1009326 is the next step in trying to drop arm64 pacakges if we go down that path | 17:34 |
| @jim:acmegating.com | yep i think that's it, i don't see any real ips there | 17:34 |
| @jim:acmegating.com | do we have any other hosts with this problem? | 17:35 |
| @fungicide:matrix.org | things are probably still not quite right with how zuul conveys addresses in flex, because the inventory does state `interface_ip: 146.20.60.58` which isn't the actual address bound to the interface | 17:36 |
| @clarkb:matrix.org | corvus: no raxflex is the only cloud provider we currently use that has floating IPs | 17:36 |
| @jim:acmegating.com | fungi: well, that's the address we use to connect to it | 17:36 |
| @clarkb:matrix.org | right from zuuls persepctive that is the IP | 17:36 |
| @jim:acmegating.com | we can't connect to a node from the executor on a private ip | 17:36 |
| @clarkb:matrix.org | if you look at the ansible facts you'll get the private "real" ips | 17:36 |
| @jim:acmegating.com | i'm wondering if we've solved this for any system-config-run jobs | 17:37 |
| @fungicide:matrix.org | maybe it's a terminology thing. it does have `public_ipv4: 146.20.60.58` which is what i would expect outside connections to rely on, to me `interface_ip` would be the address bound locally on the interface | 17:37 |
| @jim:acmegating.com | basically, is there any other service we run ansible playbook tests for that also needs to bind to a specific ip address? | 17:37 |
| @clarkb:matrix.org | corvus: I want to say we have handling for that in the gitea haproxy rules | 17:37 |
| @clarkb:matrix.org | `address: "{{ (hostvars['gitea99.opendev.org'] | default({})).get('nodepool', {}).get('public_ipv4', '') }}:3080"` hrm maybe we're pointing at the floating ips there | 17:38 |
| @jim:acmegating.com | yeah, that's more or less what we're doing for zk | 17:39 |
| @jim:acmegating.com | `server.{{ host | regex_replace('^zk(\\d+)\\.open.*\\.org$', '\\1') | int }}={{ (hostvars[host].public_v4) }}:2888:3888` | 17:39 |
| @fungicide:matrix.org | i take it zk/java needs an explicit binding address and can't just use `0.0.0.0` or `::` | 17:40 |
| @clarkb:matrix.org | I think it is more zookeeper in this instance? | 17:40 |
| @clarkb:matrix.org | ya I'm not finding any evidence of solving this for something else in the system-config/playbooks/zuul dir | 17:41 |
| @jim:acmegating.com | it's a little surprising that it's using that as the binding address | 17:42 |
| @fungicide:matrix.org | i wonder if it's relying on something like `gethostbyname()` | 17:42 |
| @fungicide:matrix.org | in which case maybe overriding entries in `/etc/hosts` would suffice | 17:43 |
| @jim:acmegating.com | i'm not following; that's going to just have an ip in the config file | 17:44 |
| @clarkb:matrix.org | no its the literal ip addresses to set up the cluster members | 17:44 |
| @clarkb:matrix.org | Google seems to think that it should bind to 0.0.0.0 by default but I think it must override that when we set the cluster membership like this | 17:44 |
| @clarkb:matrix.org | https://oneuptime.com/blog/post/2026-03-20-zookeeper-bind-ipv4-kafka/view#binding-quorum-ports-to-specific-ips maybe this is 3.7+ "new" behavior? | 17:45 |
| @jim:acmegating.com | maybe fungi is suggesting to change to a hostname? that could be a solution... we'd probably have to answer a lot of questions to confirm the right behavior in both testing and prod. | 17:45 |
| @fungicide:matrix.org | i mean whatever is deciding to tell the ListenerHandler to explicitly bind on `146.20.63.157:3888` may be doing a lookup by name, but yeah probably not since it's not going to be in dns for a test node and odds are we don't inject the global address into `/etc/hosts` ourselves anyway | 17:46 |
| @clarkb:matrix.org | it can't find on that IP if the IP isn't configured on the host though? | 17:46 |
| @clarkb:matrix.org | I'm still not sure I follow the name lookup path | 17:46 |
| @clarkb:matrix.org | * it can't bind on that IP if the IP isn't configured on the host though? | 17:47 |
| @fungicide:matrix.org | i doubt the test node itself has any awareness of its relationship with 146.20.63.157 so presumably something somewhere in ansible is plumbing that through to configuration | 17:47 |
| @clarkb:matrix.org | fungi: yes our inventory | 17:47 |
| @jim:acmegating.com | it's in the zookeeper config | 17:47 |
| @fungicide:matrix.org | i mean something in the ansible playbooks/roles is telling it to use that inventory value | 17:48 |
| @jim:acmegating.com | probably the line i pasted | 17:48 |
| @fungicide:matrix.org | rather than a literal `0.0.0.0` or `::` or empty string | 17:48 |
| @clarkb:matrix.org | yup its this https://opendev.org/opendev/system-config/src/branch/master/playbooks/roles/zookeeper/templates/zoo.cfg.j2#L31 | 17:48 |
| @clarkb:matrix.org | the reason is that you ahve to set up the cluster this way. It is how zookeeper knows who its peers are | 17:48 |
| @fungicide:matrix.org | so my earlier question was can we just not specify an address (if it will bind to all addresses by default), or specify a wildcard address? | 17:49 |
| @clarkb:matrix.org | I don't think so because then you don't have a cluster | 17:49 |
| @fungicide:matrix.org | oh, so it's both the address the process binds to *and* the address other cluster members expect to reach it at? those aren't separate configuration values? | 17:49 |
| @jim:acmegating.com | it's certainly behaving that way | 17:50 |
| @jim:acmegating.com | the docs are decidedly unclear on this | 17:50 |
| @clarkb:matrix.org | the blog post I linked above implies this is how it works too | 17:50 |
| @fungicide:matrix.org | so literally can't be made to work through address translation at all, if so | 17:51 |
| @jim:acmegating.com | oh wait, better docs here: https://zookeeper.apache.org/doc/r3.9.6/zookeeperAdmin.html | 17:51 |
| @jim:acmegating.com | okay that new multi address thing is scary | 17:52 |
| @jim:acmegating.com | and i'm not sure that would actually solve it for us, if one of ips listed is the fip and it still can't bind to it | 17:54 |
| @clarkb:matrix.org | I have some hacky ideas for working around this in ansible: essentially update run-base.yaml to set private_v4 next to public_v4 if present. Then in the zoo.cfg.j2 file use private_v4 if present else use public_v4. that should mean prod doesn't change and allow us to use the private known address on the host in CI. I don't know if that means we also need to update zuul config to prefer private over public if set when talking to zk | 17:54 |
| @clarkb:matrix.org | https://opendev.org/opendev/system-config/src/branch/master/playbooks/zuul/run-base.yaml#L81-L83 this is where run-base.yaml would be updated in that scenario | 17:55 |
| @jim:acmegating.com | Clark: yeah, i think we may need something like that. i think we may only use one zk host with the zuul tests though, so maybe we don't need to worry about that? | 17:56 |
| @clarkb:matrix.org | corvus: if those jobs share the same playbook though it will bind on the private ip and zuul will try to connect to it via public ip if we don't also update zuul's config? | 17:56 |
| @clarkb:matrix.org | its possible that it will work though | 17:56 |
| @jim:acmegating.com | Clark: lol it would have worked 4 monhs ago: https://opendev.org/opendev/system-config/commit/6d1490b20e3f443149eae3f6c72bf26cc323f137 | 17:56 |
| @clarkb:matrix.org | oh yes this is the thing we wanted to change to make it less confusing :) | 17:57 |
| @clarkb:matrix.org | we traded one version of confusion for another :/ | 17:57 |
| @clarkb:matrix.org | this is still probably the better outcome because its less magic. We'd have to explicitly use the private addr where we know we might need it | 18:00 |
| @clarkb:matrix.org | but still | 18:00 |
| @fungicide:matrix.org | the previous treat-the-private-address-as-public hack was also breaking things in the other direction, to be fair | 18:03 |
| @jim:acmegating.com | it could be argued that the old setup let us have production playbooks that worked without alteration in test environments | 18:03 |
| @jim:acmegating.com | what was it breaking? | 18:03 |
| @jim:acmegating.com | (basically, we reframed our test environment to make it look more like production, but now we will need to update our production playbooks to understand whether they are running in test or not) | 18:04 |
| @clarkb:matrix.org | corvus: it broke mixed cloud nodesets since the private IPs couldn't be routed between cloud regions | 18:04 |
| @clarkb:matrix.org | which we're using to test arm64 stuff with an x86 bridge to better match reality | 18:05 |
| @clarkb:matrix.org | there was also some issue with having an arm64 test bridge I don't rmember the details of | 18:05 |
| @fungicide:matrix.org | (once those became a thing that zuul supported, and we occasionally got them as a fallback after launch failures) | 18:05 |
| @clarkb:matrix.org | I think there was a package ansible needed without a wheel maybe | 18:05 |
| @fungicide:matrix.org | oh, right, the mixed-arch testing was the more consistent problem | 18:05 |
| @fungicide:matrix.org | yes, i don't recall the details now, but there was some ubuntu upgrade problem which led us to want to switch to mixed-node | 18:07 |
| @fungicide:matrix.org | pretty sure it was related to migrating bridge to newer ubuntu | 18:07 |
| @jim:acmegating.com | okay, so we need to: | 18:07 |
| 1) add private_ipv4 to inventory if it exists | ||
| 2) use private_ipv4 when writing zoo.cfg if it exists and is not empty, otherwise public_ipv4 | ||
| 3) do the same in zuul.conf | ||
| that about it? | ||
| @clarkb:matrix.org | corvus: I think that should do it | 18:07 |
| @fungicide:matrix.org | that does seem like it would solve the immediate error | 18:08 |
| @clarkb:matrix.org | as a side note it really is weird that knowing who your cluster members are and where you bind your socket are not separate concerns | 18:09 |
| @clarkb:matrix.org | corvus: supposedly `clientPortAddress=0.0.0.0` may cause it to bind on 0.0.0.0? | 18:10 |
| @clarkb:matrix.org | maybe it is worth testing that? | 18:10 |
| @jim:acmegating.com | but this is the quorum binding | 18:10 |
| @clarkb:matrix.org | oh right its a different port | 18:10 |
| @clarkb:matrix.org | nevermind I think this won't solve it | 18:10 |
| @clarkb:matrix.org | but also maybe that maens zuul doesn't need to know? | 18:10 |
| @clarkb:matrix.org | if the client bind is 0.0.0.0 then zuul should be able to connct to the floating ip properly | 18:11 |
| @jim:acmegating.com | i'm with you... i still have a terminal open grepping through zookeeper to try to find where it's doing that | 18:11 |
| @clarkb:matrix.org | that would explain why this worked until you added a second server | 18:13 |
| @jim:acmegating.com | Clark: oh yeah i think the default for clientPortAddress is any, so i think we don't need to worry about zuul already | 18:13 |
| @clarkb:matrix.org | maybe | 18:13 |
| @clarkb:matrix.org | ++ | 18:13 |
| @jim:acmegating.com | i should have that change as soon as i spend the next 15 minutes testing ansible udefined vars | 18:15 |
| @jim:acmegating.com | if we ever add "private_ipv4" to our public inventory, we're going to immediately break this since it'll be backwards | 18:18 |
| @clarkb:matrix.org | corvus: should it be private_v4 to match public_v4? | 18:19 |
| @jim:acmegating.com | sure whatever it's called :) | 18:19 |
| @clarkb:matrix.org | but yes good call. Maybe we even add a comment to the inventory entries for zookeeper nodes about it? | 18:19 |
| @jim:acmegating.com | maybe we should call it "silly_fake_private_v4_for_use_in_testing" | 18:19 |
| @clarkb:matrix.org | I'm fine with something along those lines | 18:19 |
| @clarkb:matrix.org | ci_specific_private_v4 might be less silly | 18:20 |
| @clarkb:matrix.org | but I'm good with silly too | 18:20 |
| @fungicide:matrix.org | `slightly_less_silly_ci_specific_private_v4` :P | 18:21 |
| @fungicide:matrix.org | i approve on behalf of the ministry of silly variable names | 18:21 |
| -@gerrit:opendev.org- James E. Blair https://matrix.to/#/@jim:acmegating.com proposed: | 19:18 | |
| - [opendev/system-config] 1009279: WIP: Test 2-node zk cluster https://review.opendev.org/c/opendev/system-config/+/1009279 | ||
| - [opendev/system-config] 1009351: Use private ipv4 for zookeeper quorum configuration in testing https://review.opendev.org/c/opendev/system-config/+/1009351 | ||
| @clarkb:matrix.org | corvus: the ternary thing should work fine but would private_ipv4 | default (public_ipv4) also work? Or maybe there is something subtle I'm missing | 19:45 |
| @clarkb:matrix.org | Maybe default only works for unset values rather than untruthy values? | 19:45 |
| @fungicide:matrix.org | i would expect it to only work for undefined values | 19:46 |
| @fungicide:matrix.org | but i could be expecting wrong | 19:46 |
| @fungicide:matrix.org | i've always assumed it's more like `.get()` methods in python | 19:47 |
| @fungicide:matrix.org | rather than if/then/else | 19:47 |
| @jim:acmegating.com | that's the assumption i wrote it with; i want to handle private_ipv4=null | 19:47 |
| @clarkb:matrix.org | Got it | 19:49 |
| -@gerrit:opendev.org- Julia Kreger proposed: [openstack/diskimage-builder] 1009353: epel: exclude yum-utils from auto-installing on non-rh style distros https://review.opendev.org/c/openstack/diskimage-builder/+/1009353 | 20:01 | |
| @jim:acmegating.com | there is a lot of red on that change, but i *think* it's all ubuntu archive network errors | 20:25 |
| @clarkb:matrix.org | ya the gitea change hit ubuntu archive network errors too | 20:25 |
| @clarkb:matrix.org | we setup our system-config jobs to not use our mirrors because our normal prod service hosts don't use our mirrors. But that makes them vulnerable to upstream mirror issues | 20:26 |
| -@gerrit:opendev.org- Steve Baker proposed on behalf of Julia Kreger: [openstack/diskimage-builder] 1009353: epel: Use sed for disabling epel repo https://review.opendev.org/c/openstack/diskimage-builder/+/1009353 | 21:56 | |
| @jim:acmegating.com | all right! https://zuul.opendev.org/t/openstack/build/9a2c8593870c40ffa551c82e290db5e5 is failing as expected | 22:47 |
| @jim:acmegating.com | testinfra is failing the test that the cluster is up, and the cluster is not up because of no SAN | 22:48 |
| @clarkb:matrix.org | success!ful failure | 22:48 |
| @jim:acmegating.com | so now we can build on that and fix the san thing | 22:48 |
| @clarkb:matrix.org | any idea if this is just a problem for the quorum members or if clients will hit it too yet? | 22:50 |
| @jim:acmegating.com | `ssl.clientHostnameVerification and ssl.quorum.clientHostnameVerification : (Java system properties: zookeeper.ssl.clientHostnameVerification and zookeeper.ssl.quorum.clientHostnameVerification) New in 3.9.4: Specifies whether the client's hostname verification is enabled in client and quorum TLS negotiation process. This option requires the corresponding hostnameVerification option to be true, or it will be ignored. Default: true for quorum, false for clients` | 22:54 |
| @jim:acmegating.com | https://zookeeper.apache.org/doc/r3.9.5/zookeeperAdmin.html | 22:54 |
| @jim:acmegating.com | that looks like a "not a problem for clients" to me | 22:55 |
| @clarkb:matrix.org | yup seems to be that way | 22:56 |
| -@gerrit:opendev.org- James E. Blair https://matrix.to/#/@jim:acmegating.com proposed: | 22:59 | |
| - [opendev/system-config] 1009279: WIP: Test 2-node zk cluster https://review.opendev.org/c/opendev/system-config/+/1009279 | ||
| - [opendev/system-config] 1009387: Add SAN to certs generated by the opendev CA https://review.opendev.org/c/opendev/system-config/+/1009387 | ||
| @jim:acmegating.com | if 387 works we can squash it into 279 for a mergeable change; | 23:00 |
| -@gerrit:opendev.org- Zuul merged on behalf of Clark Boylan: [opendev/system-config] 1009126: Upgrade gitea to 28.1.0 https://review.opendev.org/c/opendev/system-config/+/1009126 | 23:01 | |
| @clarkb:matrix.org | that got in ahead of the hourly jobs | 23:02 |
| @clarkb:matrix.org | it is about halfway through the cluster now. When it is done I'll get general browseability and git clone and then also double check replication has succeeded | 23:09 |
| -@gerrit:opendev.org- James E. Blair https://matrix.to/#/@jim:acmegating.com proposed: [opendev/system-config] 1009387: Add SAN to certs generated by the opendev CA https://review.opendev.org/c/opendev/system-config/+/1009387 | 23:14 | |
| @jim:acmegating.com | apparently opendev-ca changes didn't trigger the zookeeper job | 23:14 |
| @clarkb:matrix.org | cluster is upgraded as of 20 seconds ago | 23:15 |
| @jim:acmegating.com | maybe we should run zuul on those changes too | 23:15 |
| @clarkb:matrix.org | ++ | 23:15 |
| -@gerrit:opendev.org- James E. Blair https://matrix.to/#/@jim:acmegating.com proposed: [opendev/system-config] 1009387: Add SAN to certs generated by the opendev CA https://review.opendev.org/c/opendev/system-config/+/1009387 | 23:16 | |
| @clarkb:matrix.org | git clone works. I can browse system config, look at files on HEAD, look at file commit history and view the diff of a file on a commit | 23:17 |
| @clarkb:matrix.org | so it seems to generally work | 23:17 |
| @clarkb:matrix.org | and 1009387 replicated: https://opendev.org/opendev/system-config/commit/5ab12e2f7de0b493e268cebeea33b756c73cf150 | 23:17 |
| @clarkb:matrix.org | so gitea 28.1.0 is looking good to me | 23:17 |
| @clarkb:matrix.org | corvus: in theory our multinode test setup updates /etc/hosts so that each nodes knows about all of the others via name lookups. So I think this should work. But if things fail then actually validating the name:ip mappings is probably where I would look next | 23:19 |
Generated by irclog2html.py 4.1.0 by Marius Gedminas - find it at https://mg.pov.lt/irclog2html/!