Thursday, 2026-09-03

@jim:acmegating.comi believe the zuul jobs are actually affected by this; i will restart the executors on the revert image00:33
@jim:acmegating.com#status log restarted all executors on the logreceiver revert patch (in order to allow the logreceiver fix to merge)00:52
@status:opendev.org@jim:acmegating.com: finished logging00:52
-@gerrit:opendev.org- Zuul merged on behalf of Takashi Kajinami: [openstack/diskimage-builder] 1001676: Declare Python 3.13 support https://review.opendev.org/c/openstack/diskimage-builder/+/100167601:09
-@gerrit:opendev.org- Zuul merged on behalf of Clark Boylan: [opendev/system-config] 988993: Set kernel.yama.ptrace_scope to 2 on executors https://review.opendev.org/c/opendev/system-config/+/98899301:46
@gthiemonge:matrix.orgHi Folks, I'm troubleshooting CI failures with the octavia centos 10 stream jobs. It looks like they always fail when they run on the vexxhost provider, the route to the rdoproject repos doesn't work:07:39
Connecting to trunk.rdoproject.org (trunk.rdoproject.org)|38.129.56.180|:443... failed: No route to host.
https://35f57d723fb280f5d9a6-c5982ef1d1780edfccf5471d12724c1b.ssl.cf2.rackcdn.com/openstack/311f30a8482b4729be4addc0a0bb8b41/job-output.txt
I searched on opensearch, the connection to the repo works from all regions except vexxhost-ca-ymq-1
Is it an issue that we can fix here?
@harbott.osism.tech:regio.chatlooks like trunk.rdoproject.org itself is also hosted by vexxhost, so maybe they have some internal routing issue? mnaser might be able to check or someone can check from our mirror node11:51
@harbott.osism.tech:regio.chatcorvus: not sure if you saw that, 1002828,2 seems to also have been stuck in gate until you did the revert on the executors. given that we are very close to the end of the OpenStack release cycle, I would vote to keep zuul in this state for the next four weeks unless any further issues are detected of course. maybe we'd even need to disable the weekly restart cron to achieve that?11:57
@jim:acmegating.comJens Harbott: i don't think we should freeze zuul for 4 weeks, that's an exceptionally long time.  i also think we're mostly finished with the debugging of this particular change.  i'm not worried about it having significant impacts.13:17
@jim:acmegating.com@status log restarted all zuul executors on zuul master to pick up latest logreceiver fixes13:21
@jim:acmegating.com#status log restarted all zuul executors on master to pick up latest logreceiver fixes13:22
@status:opendev.org@jim:acmegating.com: finished logging13:22
-@gerrit:opendev.org- Stephen Finucane proposed: [openstack/project-config] 988094: Add nova-reviewers to openstack/nova ACL https://review.opendev.org/c/openstack/project-config/+/98809415:22
@clarkb:matrix.orghttps://github.com/orgs/community/discussions/206581 this is the upstream github discussion around requiring authentication. This has affected Gerrit CI. I am not aware of us having hit issues but it wouldn't surprise me if people have qusetions about it15:22
@clarkb:matrix.orggthiemonge: Jens Harbott I can reach trunk.rdoproject.org from mirror.ca-ymq-1.vexxhost.opendev.org. Makes me wonder if the issue is/was temporary or the routing trouble may be more subtle and impact specific hypervisors ?15:24
@clarkb:matrix.orgre github `a subset of unauthenticated clone or fetch requests may now be asked to authenticate as part of GitHub’s protections against abusive traffic. If you receive a 401, update your application or script to use GitHub credentials.`15:28
@gthiemonge:matrix.orgClark: in opensearch, if I search "Connecting to trunk.rdoproject.org" and filter with hosts_region = "vexxhost-ca-ymq-1", all the attempts have failied since Aug 2615:28
@clarkb:matrix.orggthiemonge: are there any successful jobs in that region too? Note that particular log message may only appear when it fails to connect (I'm not sure)15:30
@clarkb:matrix.orgjust wondering if we can narrow it down further15:30
@clarkb:matrix.orgit is also theoretically possible that the firewall(s) (if any) in front of trunk.rdoproject.org are blockign our requests15:31
@gthiemonge:matrix.orgClark: only failures after Aug 26, I see that it worked before ("Connecting to trunk.rdoproject.org (trunk.rdoproject.org)|38.129.56.180|:443... connected")15:32
@clarkb:matrix.orgack thank you for confirming15:33
@clarkb:matrix.orgso likely something changes around August 26 either in the cloud itself (routing, networks, neutron, etc) or in the configuration for trunk.rdoproject.org (firewalls, bot mitigation, etc)15:34
@clarkb:matrix.organd that has consistently prevented test nodes launched in the same cloud region from accessing the resources but not random curls issued from mirror.ca-ymq-1.vexxhost.opendev.org which is hosted in the same region15:34
-@gerrit:opendev.org- Stephen Finucane proposed:15:44
- [openstack/project-config] 1003838: Initiate retirement of molteniron https://review.opendev.org/c/openstack/project-config/+/1003838
- [openstack/project-config] 1003839: Retire molteniron https://review.opendev.org/c/openstack/project-config/+/1003839
- [openstack/project-config] 1003840: Remove references to ironic-core https://review.opendev.org/c/openstack/project-config/+/1003840
-@gerrit:opendev.org- Stephen Finucane proposed:15:45
- [openstack/project-config] 1003839: Retire molteniron https://review.opendev.org/c/openstack/project-config/+/1003839
- [openstack/project-config] 1003840: Remove references to ironic-core https://review.opendev.org/c/openstack/project-config/+/1003840
-@gerrit:opendev.org- Stephen Finucane proposed: [openstack/project-config] 1003840: Remove references to ironic-core https://review.opendev.org/c/openstack/project-config/+/100384015:51
-@gerrit:opendev.org- Stephen Finucane proposed: [openstack/project-config] 1003840: Remove references to ironic-core https://review.opendev.org/c/openstack/project-config/+/100384015:55
-@gerrit:opendev.org- Stephen Finucane proposed: [openstack/project-config] 1003848: Add nova-approvers to openstack/nova ACL https://review.opendev.org/c/openstack/project-config/+/100384816:00
@stephenfin:matrix.orgHow does one get a new Gerrit group created? Context being https://lists.openstack.org/archives/list/openstack-discuss@lists.openstack.org/thread/EI44ZQG6UHIBNH3O26KHUVHIE576SCOC/16:11
@stephenfin:matrix.orgAnd additionally, how does one get one *deleted* (assuming such a thing is possible/a good idea). Context for that being https://review.opendev.org/c/openstack/project-config/+/100384016:12
@clarkb:matrix.orgstephenfin: when the change (988094) merges there is a deployment step to update the acl in gerrit. That step will create any missing groups before applying acls that use the group16:12
@stephenfin:matrix.orgTIL16:12
@clarkb:matrix.orgthe group will be created empty so we will need to add an initial seed member who can fill in everyone else (usually that is given to the PTL of the project)16:12
@clarkb:matrix.orgdeleting groups is something we haven't done historically because it wasn't well supported. You just stop using the group in the acls and remove everyone from them. I think in like the last gerrit release or two we may actually be able to delete them safely now (it checks they aren't used etc) but we haven't done any work around that yet16:13
@stephenfin:matrix.orgOkay, on deletion, I'll tell the ironic folks just to remove themselves16:14
@stephenfin:matrix.orgOn creation, Uggla is PTL but not in nova-core (thus I had to make the changes I announced earlier today). I suspect the best thing to do is include nova-core? That's what ironic has done in ironic-reviewers https://review.opendev.org/admin/groups/cfda7dc8666aec43e442acd9c7a586c8d2c93895,members16:14
@clarkb:matrix.orgya I think that would work16:15
@stephenfin:matrix.orgI will ask him if he can +1 the project-config change though16:15
@clarkb:matrix.org++16:15
-@gerrit:opendev.org- James E. Blair https://matrix.to/#/@jim:acmegating.com proposed: [zuul/zuul-jobs] 1003852: Add buildset registry password to redactions https://review.opendev.org/c/zuul/zuul-jobs/+/100385216:18
-@gerrit:opendev.org- Stephen Finucane proposed: [openstack/project-config] 1003839: Retire molteniron https://review.opendev.org/c/openstack/project-config/+/100383916:23
-@gerrit:opendev.org- Stephen Finucane proposed: [openstack/project-config] 1003840: Remove references to ironic-core https://review.opendev.org/c/openstack/project-config/+/100384016:23
-@gerrit:opendev.org- Zuul merged on behalf of Roja Eswaran: [openstack/diskimage-builder] 999989: debootstrap: add DIB_MMDEBSTRAP_EXTRA_ARGS support https://review.opendev.org/c/openstack/diskimage-builder/+/99998916:51
@fungicide:matrix.orgstephenfin: an alternative is that if an openstack project declares tact sig liaisons we add those initially instead17:27
@fungicide:matrix.orgon the expectations that they're the ones delegated by the ptl or under dpl to handle adding other members17:27
@harbott.osism.tech:regio.chathttps://review.opendev.org/c/opendev/zuul-providers/+/1000965 keeps failing in gate with swift upload failures :(18:23
@fungicide:matrix.orgi think most zuul-providers changes that end up uploads to all providers are failing on enough random providers that they can never merge. i have one i've been rechecking off and on for months18:25
@harbott.osism.tech:regio.chatmaybe we need to force-merge and then have the regular rebuilds roll out the changes? or make all those jobs non-voting in gate?18:26
@fungicide:matrix.orgyeah, https://review.opendev.org/c/opendev/zuul-providers/+/989136 was proposed in may and still hasn't merged18:26
@harbott.osism.tech:regio.chathmm, seems like V-2 doesn't trigger the attention set, so I missed that one18:27
@fungicide:matrix.orgboth options sound reasonable if we can't figure out why uploads are so unreliable18:27
@jim:acmegating.comin gate, we're only uploading to a single location: the new rax-flex swift in dfw18:41
@jim:acmegating.comwe could choose a different location if we think it's more reliable, or we could do some fallback stuff18:41
@jim:acmegating.comi think it's worth keeping the uploads in gate though; it's been very beneficial to have changes roll out immediately18:42
@jim:acmegating.comhttps://review.opendev.org/c/opendev/zuul-providers/+/989136?tab=change-view-tab-header-zuul-results-summary -- that change had no upload failures, those were all build failures18:44
@jim:acmegating.comfor https://review.opendev.org/c/opendev/zuul-providers/+/1000965?tab=change-view-tab-header-zuul-results-summary -- 2 of those were build failures 4 were upload failures18:45
@jim:acmegating.comso here's my proposal for action for the upload failures:18:51
1) if someone wants to talk with the rax flex folks about those 500 errors, that could be beneficial
2) if we suspect those errors may be transient (the other successful uploads suggest they may be), someone could propose a change to wrap that in a retry loop. i would be happy to advise and review that
3) if someone has a suggestion for a different location to upload to, i would be happy to set up the swift container and write the change to switch to that location
@jim:acmegating.commeanwhile, identifying the cause of those build failures (looks like there may be some disk space issues) is another thing that folks could do to improve things18:51
@clarkb:matrix.orgfungi: do you still have a line to James?18:53
@clarkb:matrix.org(maybe we should ask if he/they have a preferred escalation method eg using the ticket system or contacting directly etc)18:53
@fungicide:matrix.orgyeah, i think we tipped over the reliability threshhold to completely blocked when the disk space errors started18:53
@jim:acmegating.comhere's an example swift upload 500 error: https://zuul.opendev.org/t/opendev/build/7ac0e6e2fd0e4350ae708a5895e379c518:54
@jim:acmegating.combetter link: https://zuul.opendev.org/t/opendev/build/7ac0e6e2fd0e4350ae708a5895e379c5/console#4/0/40/ubuntu-noble18:55
@fungicide:matrix.orgi was talking to james denton via irc privmsg but it seems like he only appeared when someone else (doug?) prodded, and then stopped responding again once the immediate problem with the mirror servers in error mode was resolved18:55
@jim:acmegating.comthat actually has our container name in it18:55
@clarkb:matrix.orgmaybe we should file a couple of tickets then? One for the auth thing and another for this swift upload thing18:55
@clarkb:matrix.orgnot to distract from this conversation, but it has been ~24 hours since I shutdown the old backup02 server. I don't see any complaints in the infra root inbox. We good with deleting that backup02 server now?18:56
@fungicide:matrix.orgi'll go ahead and try resetting our api token for the control plane account first and see if that solves the flex keystone auth problem, then open a support ticket if it doesn't18:58
@fungicide:matrix.orgmainly because i suspect that's the first thing they'll ask us to do before they look into their end18:58
@clarkb:matrix.orgsounds good18:59
@fungicide:matrix.orgit will invalidate the auth for that tenant in all classic and flex regions until we get clouds.yaml files updated, fair warning18:59
@clarkb:matrix.orgit is the control plane side so shouldn't affect zuul though18:59
@clarkb:matrix.orgthe impact is likely to be small if any19:00
@fungicide:matrix.orgso i want to make sure that if anyone is in the middle of trying to bring up new servers anywhere in our control plane in rackspace i don't step on toes19:00
@clarkb:matrix.orgfungi: you didn't see any backup complaints in the root inbox did you? Just to make sure my email filtering foo isn't missing anything obvious there19:00
@fungicide:matrix.orglemme check19:00
@fungicide:matrix.orgi'm a bit behind on e-mail this week, it's been busy19:01
@clarkb:matrix.organd I am in the middle of talking to vexxhost not rax so I don't expect any issues with the plan to change the token19:01
@fungicide:matrix.orgthe ze11 replacement or the backup server removal?19:01
@clarkb:matrix.orgbackup server removal is vexxhost19:02
@clarkb:matrix.orgI haven't looked at ze11 replacement. That would be rax classic if sticking to where the other executors are though19:02
@clarkb:matrix.org(which would be affected)19:02
@fungicide:matrix.orgokay yeah ze11 work would get impacted if i'm doing this at the same time19:02
@jim:acmegating.comi'm planning on doing ze11 work right after lunch actually, but don't worry, if it breaks, that's no big deal19:03
@fungicide:matrix.orgi see https://review.opendev.org/c/opendev/system-config/+/988993 deploy failed on infra-prod-service-zuul early this morning utc, i'm assuming that's already known19:04
@clarkb:matrix.orgnope. But we run the zuul deployment hourly so let me check if it caught up later19:05
@clarkb:matrix.orghttps://zuul.opendev.org/t/openstack/builds?job_name=infra-prod-service-zuul&skip=0 it did catch up19:05
@fungicide:matrix.orggood enough19:05
@jim:acmegating.comi'll look into that anyway; there's a good chance i'm responsible for whatever broke it19:06
@fungicide:matrix.orgtiming seems like it might have coincided with something getting restarted19:06
@clarkb:matrix.orglooks like docker compose pull failed on ze0319:07
@clarkb:matrix.org`service-zuul.yaml.log.2026-09-03T03:43:34` is the log file on bridge I think19:07
@clarkb:matrix.orgcorvus: `no space left on device`19:08
@clarkb:matrix.orgcorvus: do we need to prune docker stuff there after the work you'ev done on it?19:08
@jim:acmegating.comyeah, also, we seem to be emitting a lot of log lines to docker; i think something weird is going on there19:08
@clarkb:matrix.orgdf says we have a little headroom there right now but not a ton19:08
@jim:acmegating.comare those logs going to the journal?19:10
@clarkb:matrix.orgwe don't appear to have any special journal logging config in the docker compose file on ze0319:10
@clarkb:matrix.orgso I think anything going to stdout/stderr would be captured by the container runtime and not forwarded further19:11
@clarkb:matrix.orgI suspect we did that because we configure zuul to log to files directly and don't expect anything out stdout/stderr19:11
@jim:acmegating.com            "LogConfig": {19:12
"Type": "journald",
"Config": null
},
@jim:acmegating.comthat's from docker inspect19:12
@clarkb:matrix.orgoh maybe its a default? and without our extra config it just doesn't go to /var/log/containers because it doesn't get the extra tagging. So it goes into the main journal?19:12
@jim:acmegating.comi think so, that's probably why it recovered19:13
@jim:acmegating.commaybe the main thing to do here is to track down why we're getting so many messages to the container log; i conisdered that low priority, but i'll bump that to the top of the list.19:13
@clarkb:matrix.orgya I am surprised we get anything to stdout/stderr since we expect things in /var/log/zuul/ files via the python logging config. But maybe we've missed something in that config or its giong to stdout too?19:14
@jim:acmegating.com(i mean, i know the mechanism, i just haven't collected all the info to propose a solution yet)19:14
@clarkb:matrix.orgack19:15
@jim:acmegating.comi'll do that first after lunch, then ze11 :)19:15
@clarkb:matrix.orgok last call on backup02 deletion19:16
@clarkb:matrix.orghas anyone seen any raeson we should delay that server delete?19:16
@fungicide:matrix.orgstill checking infra-root e-mail19:17
@clarkb:matrix.orggot it I'll wait for your report before proceeding19:17
@fungicide:matrix.orgclearing out the noise takes time19:17
@fungicide:matrix.orgmainly my mua is waiting on gmail's imap to respond to the request to process 22k deletes based on a pattern match19:18
@fungicide:matrix.orgthat's several days worth19:18
@fungicide:matrix.orgusually takes a few minutes19:18
@fungicide:matrix.orgwith batches this large it frequently times out and eventually has to be retried/continued19:20
@fungicide:matrix.orgdone. checking the infra-root inbox (after deleting a few thousand backscatter from the weekend) i see no notifications about backup failures. also checked the spam folder since they sometimes land there (after deleting tens of thousands of noise messages) and nothing there either19:21
@fungicide:matrix.orgshould be safe to proceed!19:21
@clarkb:matrix.orgthanks I am proceeding now19:24
@clarkb:matrix.org#status log Deleted backup02.ca-ymq-1.vexxhost.opendev.org as it has been replaced by backup03.ca-ymq-1.vexxhost.opendev.org19:31
@status:opendev.org@clarkb:matrix.org: finished logging19:31
@jim:acmegating.comremote:   https://review.opendev.org/c/zuul/zuul/+/1003883 Reduce "Ansible output exceeds max. line size" log entries [NEW]        21:02
@jim:acmegating.comClark: fungi i think that our log config is fine and we need to change zuul ^21:03
@fungicide:matrix.orgneat21:08
@clarkb:matrix.orgI guess this is a side effect of the changes then too and not an outstanding issue21:16
@jim:acmegating.comClark: actually no!21:16
@jim:acmegating.comi really dug into it, and this has been around for a long time21:17
@fungicide:matrix.orgjust think of the rust we spun unnecessarily all those years21:18
@jim:acmegating.commy only guess as to why we noticed it on ze03: my temporary log files from debugging, more restarts means more jobs and more logs21:18
@clarkb:matrix.orgah got it21:23
@jim:acmegating.comdo we expect the "Ubuntu Noble OpenDev 20250110" image in rax-dfw to work?21:39
@jim:acmegating.comi ask because it looks like the launch script is failing at the ssh connection stage21:39
@jim:acmegating.comoh hey there's "Ubuntu 24.04 LTS (Cloud)"21:40
@jim:acmegating.comi thought i searched for that21:40
@fungicide:matrix.orgi would expect it to work, but probably for the same reasons you expected it to work21:40
@jim:acmegating.comi'll try that next :)21:40
@clarkb:matrix.orgcorvus: yes our image works, but you have to tell it to use config drive21:41
@jim:acmegating.comooh21:41
@clarkb:matrix.orgit only works with config drive because rax doesn't do regular metadata service and the image we upload uses normal config drive not their patched thing21:41
@jim:acmegating.comClark: would you recommend i use ours or try theirs?21:41
@jim:acmegating.comnot sure what we did for the other executors21:41
@clarkb:matrix.org* it only works with config drive because rax doesn't do regular metadata service and the image we upload uses normal cloud init not their patched thing21:41
@jim:acmegating.com(but would probably be good to match that)21:42
@clarkb:matrix.orgcorvus: I think most of our stuff initially was based on our image because they didn't haev a noble image available for some time21:42
@clarkb:matrix.orgI agree that matching what the others are is probably best for consistency21:42
@clarkb:matrix.orgopenstack server show ze01.opendev.org should tell you what the image was21:42
@jim:acmegating.comours21:42
@clarkb:matrix.orglaunch node script has a --config-drive flag to enable config drive usage21:43
@jim:acmegating.comyep, running `/usr/launcher-venv/bin/launch-node $FQDN --flavor "$FLAVOR" --cloud=$OS_CLOUD --region=$OS_REGION_NAME --image 8e6a5ba1-0447-4652-bae3-04d016746a9f --config-drive` now21:44
@jim:acmegating.comcouple more times i might remember this21:44
-@gerrit:opendev.org- James E. Blair https://matrix.to/#/@jim:acmegating.com proposed: [opendev/zone-opendev.org] 1003887: Add new ze11 https://review.opendev.org/c/opendev/zone-opendev.org/+/100388721:57
-@gerrit:opendev.org- James E. Blair https://matrix.to/#/@jim:acmegating.com proposed: [opendev/system-config] 1003888: Add ze11 https://review.opendev.org/c/opendev/system-config/+/100388821:58
-@gerrit:opendev.org- James E. Blair https://matrix.to/#/@jim:acmegating.com proposed: [opendev/system-config] 1003889: Add ze11 to cacti https://review.opendev.org/c/opendev/system-config/+/100388921:59
-@gerrit:opendev.org- Zuul merged on behalf of James E. Blair https://matrix.to/#/@jim:acmegating.com: [opendev/zone-opendev.org] 1003887: Add new ze11 https://review.opendev.org/c/opendev/zone-opendev.org/+/100388722:15
-@gerrit:opendev.org- Zuul merged on behalf of James E. Blair https://matrix.to/#/@jim:acmegating.com:22:53
- [opendev/system-config] 1003888: Add ze11 https://review.opendev.org/c/opendev/system-config/+/1003888
- [opendev/system-config] 1003889: Add ze11 to cacti https://review.opendev.org/c/opendev/system-config/+/1003889
@clarkb:matrix.orgI think ze11 should be deployed now23:41

Generated by irclog2html.py 4.1.0 by Marius Gedminas - find it at https://mg.pov.lt/irclog2html/!